MIP for all channels in 4:4:4 chroma format and single tree case
By applying a unified MIP mode for intra-prediction blocks of multiple color components, the decoder/encoder addresses inefficiencies in current video encoding, reducing bitstream size and complexity while maintaining encoding efficiency.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
- Filing Date
- 2026-01-14
- Publication Date
- 2026-04-10
AI Technical Summary
Current video encoding technologies do not efficiently utilize matrix-based intra-prediction (MIP) for all color components of a picture, leading to increased bitstream size and encoding complexity due to the need for different block shapes between luma and chroma components.
Implementing a block-based decoder/encoder that uses the same MIP mode for intra-prediction blocks of multiple color components sharing the same location, reducing the bitstream size and complexity by leveraging the strong correlation across color components.
This approach reduces the bitstream size and encoding complexity by using a unified MIP mode for color components with equal spatial resolution, thereby optimizing signal transmission costs and decoder/encoder efficiency.
Smart Images

Figure 2026063155000211 
Figure 2026063155000212 
Figure 2026063155000213
Abstract
Description
[Technical Field]
[0001] Embodiments of the present invention relate to an apparatus and method for encoding or decoding a picture or video using matrix-based intra-prediction (MIP) for all channels in a 4:4:4 chroma format and single-tree case. [Background technology]
[0002] In current VTMs, MIP is used only for the luma component [3]. If the intra-mode of a chroma intra-block is direct mode (DM) and the intra-mode of a luma block in the same location is MIP mode, the chroma block must use planar mode to generate the intra-prediction signal. The main reason for treating the DM mode in this way in the case of MIP is that, in the case of 4:2:0 or dual trees, a luma block in the same location may have a different shape than the color block. Therefore, since the MIP mode is not applicable to all block shapes, the MIP mode of a luma block in the same location may not be applicable to the chroma block in that case.
[0003] Therefore, to support matrix-based intra-prediction, it is desirable to provide a concept for rendering picture and / or video encodings more efficiently. In addition, or instead, it is desirable to reduce the bitstream size and thus reduce the signal transmission cost.
[0004] This is achieved by the subject matter of the independent claims of this application.
[0005] Further embodiments according to the invention are defined by the subject matter of the dependent claims of this application. [Overview of the Initiative]
[0006] According to a first aspect of the present invention, one problem encountered when attempting to use a matrix-based intra-prediction mode (MIP mode) to predict samples of a given block of a picture stems from the fact that it is currently not possible to use MIP for all color components of a picture. According to a first aspect of the present application, this difficulty is overcome by using the same MIP mode for intra-prediction blocks of two color components that share the same location in a picture, i.e., intra-prediction blocks that are in the same location, for example, if two color components are sampled equally and divided equally into blocks. The inventors have found that it is advantageous to use the same MIP mode for intra-prediction blocks of two or more color components of a picture that are in the same location. This is based on the idea that a strong correlation of the prediction signals across all color components of a picture is beneficial, and that such correlation can be increased when the same intra-prediction mode is used for two or more color components of intra-prediction blocks of a picture. By using the same MIP mode for blocks of different color components of the same picture that are in the same location, the bitstream and therefore the signal transmission cost can be reduced. Furthermore, the complexity of the encoding can be reduced.
[0007] Accordingly, according to a first aspect of this application, a block-based decoder / encoder is configured to divide a picture of multiple color components and color sampling formats using a division scheme in which the picture is divided equally with respect to each color component, such as all color components of the picture having the same spatial resolution, for example, the first color component of the picture is divided in the same way as all other color components of the picture, for example, all color components have the same spatial resolution. The picture may consist of a luminous component and two chroma components in, for example, the YUV, YPbPr and / or YCbCr color space, where the luminous component and the two chroma components represent multiple color components. For example, in the RGB color space, the picture may consist of red, green, and blue components representing multiple color components. The picture may consist of, for example, a first color component, a second color component, and an optional third color component. It is clear that the block-based decoder / encoder described can also be used for photographs containing other color components or for photographs with different numbers of color components. For example, a picture is divided evenly for each color component by dividing the first color component into a first color component block, the second color component into a second color component block, and optionally the third color component into a third color component block. The decision of whether the color components of a picture are inter-predicted, i.e., inter-encoded, or intra-predicted, i.e., intra-encoded, can be made at a granular level or at the color component block level. The block-based decoder / encoder is configured to encode / decode the first color components of a picture to and from a data stream at the block level, i.e., at the first color component block level, by selecting one of a first set of intra-prediction modes for each of the intra-predicted first color component blocks of the picture, i.e., for each of the first color component blocks associated with intra-prediction. The first color component blocks associated with inter-prediction, i.e., inter-predicted first color component blocks, are treated differently.A first set of intra-prediction modes includes matrix-based intra-prediction modes by which the interior of a block is predicted by each mode, which involves deriving a sample-value vector from a reference sample, adjacency within a block, calculating the matrix-vector product between the sample-value vector and the prediction matrix associated with each matrix-based intra-prediction mode to obtain a prediction vector, and predicting the samples within the block based on the prediction vector. Furthermore, a block-based decoder / encoder is configured to encode / decode a second color component of a picture on a block-by-block basis by intra-predicting a given second color component block of the picture using a selected matrix-based intra-prediction mode for the intra-predicted first color component block located in the same place. The first color component of the picture may be the lumens component of the picture, and the second color component of the picture may be the chromens component of the picture.
[0008] According to one embodiment, the number of components in the prediction vector is less than the number of samples within the block, and the block-based decoder / encoder is configured to predict samples within the block based on the prediction vector by interpolating samples based on the components of the prediction vector allocated to support the sample positions within the block. This is based on the idea that such a prediction vector can be obtained by reducing the total number of multiplications required to compute the matrix-vector product, which may reduce the complexity of the decoder / encoder and the signal transmission cost.
[0009] According to one embodiment, a block-based decoder / encoder is configured to select one of a first option and a second option for each intra-predicted second color component block of a picture, i.e., for each second color component block associated with an intra-prediction. In the first option, the intra-prediction mode for each intra-predicted second color component block is derived based on the intra-prediction mode selected for the intra-predicted first color component block at the same location, such that the intra-prediction mode for each intra-predicted second color component block is equal to the intra-prediction mode selected for the intra-predicted first color component block at the same location, if the intra-prediction mode selected for the intra-predicted first color component block at the same location is one of the matrix-based intra-prediction modes. In the second option, the intra-prediction mode for each intra-predicted second color component block is selected based on the intra-mode index present / notified in the data stream for each intra-predicted second color component block. For example, a decoder / encoder may be configured to select the first option if either a direct mode or a residual coding color transformation mode is indicated for the intra-predicted second color component block. Otherwise, for example, by selecting the second option, the mode indicated in the data stream is used to predict the intra-predicted second color component block.
[0010] According to one embodiment, a block-based decoder / encoder is configured such that a matrix-based intra-prediction mode, consisting of a first set of intra-prediction modes, is selected according to the block dimension of each intra-predicted first color component block, such that the matrix-based intra-prediction mode is selected from a set of coprime subsets of matrix-based intra-prediction modes. For example, each subset of matrix-based intra-prediction modes may be associated with a particular block dimension. For example, only the matrix-based intra-prediction modes of the selected subset constitute the first set of intra-prediction modes. The first set of intra-prediction modes from which the intra-prediction mode for each intra-predicted first color component block is selected may differ for different block dimensions. Therefore, the decoder / encoder is configured to pre-select an intra-prediction mode suitable for each intra-predicted first color component block based on the block dimension of the intra-predicted first color component block. This can reduce the complexity of the decoder / encoder and the cost of signal transmission.
[0011] According to one embodiment, a prediction matrix associated with a set of coprime subsets of matrix-based intra-prediction modes is machine-learned, where one subset of matrix-based intra-prediction modes has a prediction matrix of equal size, and two subsets of matrix-based intra-prediction modes selected for different block sizes have prediction matrices of different sizes. Each subset may group multiple matrix-based intra-prediction modes associated with prediction matrices of the same dimension.
[0012] According to one embodiment, a block-based decoder / encoder is configured such that prediction matrices associated with a matrix-based intra-prediction mode, which consists of a first set of intra-prediction modes, are equal in size to each other and are machine-learned.
[0013] According to one embodiment, a block-based decoder / encoder is configured such that an intra-prediction mode consisting of a first set of intra-prediction modes, other than a matrix-based intra-prediction mode, includes a DC mode, a planar mode, and a directional mode. For example, the first set of intra-prediction modes includes a matrix-based intra-prediction mode, plus a DC mode and / or a planar mode and / or a directional mode.
[0014] According to one embodiment, a block-based decoder / encoder is configured to select a partitioning scheme from a set of partitioning schemes. The set of partitioning schemes includes a further partitioning scheme in which a picture is partitioned for a first color component using a first partitioning information present / notified in a data stream, and a second partitioning scheme in which the picture is partitioned for a second color component using a second partitioning information present / notified in a different data stream from the first partitioning information. Thus, the set of partitioning schemes includes, for example, a partitioning scheme in which the picture is partitioned equally for each color component, and a further partitioning scheme in which the picture is partitioned for the first color component in a way that is different for the second color component.
[0015] According to one embodiment, a block-based decoder / encoder is configured to select one of the first and second options, as described above, for each intra-predicted second color component block of a picture. In the first option, the intra-prediction mode for each intra-predicted second color component block is derived based on the intra-prediction mode selected for the intra-predicted first color component block at the same location, such that the intra-prediction mode for each intra-predicted second color component block is equal to the intra-prediction mode selected for the intra-predicted first color component block at the same location, if the intra-prediction mode selected for the intra-predicted first color component block at the same location is one of the matrix-based intra-prediction modes. In the second option, the intra-prediction mode for each intra-predicted second color component block is selected based on the intra-mode index present / notified in the data stream for each intra-predicted second color component block. Furthermore, the block-based decoder / encoder is configured to divide the further picture into multiple color components and color sampling formats, such that each color component is sampled evenly into further blocks using a further division scheme, i.e., the further picture is divided with respect to the first color component differently with respect to the second color component, as described above. The first color component of the further picture may be divided into further blocks of the first color component, and the second color component of the further picture may be divided into further blocks of the second color component. The decoder / encoder is configured to encode / decode the first color component of the further picture in units of further blocks, for example, by selecting one from a first set of intra-prediction modes for each of the further blocks of the intra-predicted first color component of the further picture. Furthermore, the decoder / encoder may be configured to select one from a first option and a second option for each of the further blocks of the further picture of the intra-predicted second color component.In the first option, the intra-prediction mode for each intra-predicted second color component further block is derived based on the intra-prediction mode selected for the same location further block of the first color component, such that the intra-prediction mode for each intra-predicted second color component further block is equivalent to a planar intra-prediction mode if the intra-prediction mode selected for the same location further block of the first color component is one of the matrix-based intra-prediction modes. In the second option, the intra-prediction mode for each intra-predicted second color component further block is selected based on the intra-mode index present / notified in the data stream for each intra-predicted second color component further block. As already outlined above, the first option shows, for example, that the prediction mode for an intra-predicted second color component block is directly derived from the prediction mode for an intra-predicted first color component block at the same location. This derivation may depend on the partitioning scheme selected for the picture. Therefore, the implementation of the first option may differ for the picture compared with further pictures. This is because the picture is divided using a division scheme that divides the picture equally with respect to each color component, further pictures are divided using a further division scheme that divides the picture for a first color component using a first division information present / notified in the data stream, and the picture is divided for a second color component using a second division information present / notified in a different data stream from the first division information.
[0016] According to one embodiment, a block-based decoder / encoder is configured to select one of the first and second options, as described above, for each intra-predicted second color component block of a picture. In the first option, the intra-prediction mode for each intra-predicted second color component block is derived based on the intra-prediction mode selected for the intra-predicted first color component block at the same location, such that the intra-prediction mode for each intra-predicted second color component block is equal to the intra-prediction mode selected for the intra-predicted first color component block at the same location, if the intra-prediction mode selected for the intra-predicted first color component block at the same location is one of the matrix-based intra-prediction modes. In the second option, the intra-prediction mode for each intra-predicted second color component block is selected based on the intra-mode index present / notified in the data stream for each intra-predicted second color component block. Furthermore, the block-based decoder / encoder is configured to divide the further picture into multiple color components and different color sampling formats, thereby sampling the multiple color components differently, and to decode / encode another first color component of the further picture in units of further blocks by selecting one from a first set of intra-prediction modes for each further block of the intra-predicted first color component of the further picture. The block-based decoder / encoder may be configured to select one from a first option and a second option for each further block of the intra-predicted second color component of the further picture.In the first option, the intra-prediction mode for each intra-predicted second color component further block is derived based on the intra-prediction mode selected for the same location of the first color component further block, such that it is equivalent to a planar intra-prediction mode if the intra-prediction mode selected for the same location of the first color component further block is one of the matrix-based intra-prediction modes. In the second option, the intra-prediction mode for each intra-predicted second color component further block is selected based on the intra-mode index present / notified in the data stream for each intra-predicted second color component further block. As already outlined above, the first option shows, for example, that the prediction mode for an intra-predicted second color component block is directly derived from the prediction mode for an intra-predicted first color component block at the same location. This derivation may depend on the picture's color sampling format. Therefore, the implementation of the first option may differ for a picture compared to further pictures. This is because a picture has a color sampling format in which each color component is sampled equally, while further pictures have different color sampling formats in which multiple color components are sampled differently.
[0017] According to one embodiment, a block-based decoder / encoder is configured to select a partitioning scheme from a set of partitioning schemes, the set of partitioning schemes includes further partitioning schemes in which a picture is partitioned for a first color component using first partitioning information in a data stream, and the picture is partitioned for a second color component using second partitioning information that is present / notified in a data stream separate from the first partitioning information. The block-based decoder / encoder is configured to partition further pictures for multiple color components using the partitioning scheme or further partitioning schemes.
[0018] According to one embodiment, a block-based decoder / encoder is configured to perform a selection from first and second options, in response to signal transmission, for each intra-predicted second color component block for each intra-predicted second color component block if the residual coding color conversion mode is notified to be deactivated for each intra-predicted second color component block in the data stream, and to perform a selection by inferring that the first option is selected if the residual coding color conversion mode is notified to be activated for each intra-predicted second color component block in the data stream.
[0019] One embodiment relates to a block-based decoding / coding method, comprising dividing a picture of multiple color components and color sampling formats into blocks using a division scheme in which the picture is divided equally with respect to each color component, and each color component is sampled equally. Furthermore, the method comprises decoding / coding the first color component of a block-based picture by selecting one from a first set of intra-prediction modes, including matrix-based intra-prediction modes for each mode in which the block interior is predicted, by selecting one for each intra-predicted first color component block of the picture, by deriving a sample value vector of a reference sample inside the block interior, being adjacent to the block interior, calculating the matrix-vector product between the sample value vector and the prediction matrix associated with each matrix-based intra-prediction mode to obtain a prediction vector, and predicting the sample inside the block based on the prediction vector. Furthermore, the method comprises decoding / coding the second color component of a block-based picture by intra-predicting a given second color component block of the picture using the selected matrix-based intra-prediction mode for the intra-predicted first color component block located in the same place.
[0020] The method described above is based on the same considerations as the encoder / decoder described above. The method may also be completed by all the features and functionalities described for the encoder / decoder.
[0021] One embodiment relates to a data stream having a picture or video encoded therein using the encoding method described herein.
[0022] One embodiment relates to a computer program having program code for performing the method described herein when executed on a computer.
[0023] The drawings are not necessarily to scale; instead, the focus is on illustrating the principles of the invention. Various embodiments of the invention are described in the following explanatory sections with reference to the following drawings. [Brief explanation of the drawing]
[0024] [Figure 1] An embodiment for encoding into a data stream is shown. [Figure 2] An embodiment of the encoder is shown. [Figure 3] This shows an embodiment of picture reconstruction. [Figure 4] An embodiment of the decoder is shown. [Figure 5] A schematic diagram of block prediction for encoding and / or decoding according to one embodiment is shown. [Figure 6] One embodiment demonstrates a matrix operation for predicting blocks for encoding and / or decoding. [Figure 7.1] This demonstrates block prediction using a reduced sample value vector according to one embodiment. [Figure 7.2] This demonstrates block prediction using sample interpolation according to one embodiment. [Figure 7.3] This example demonstrates block prediction using a reduced sample value vector, where only some boundary samples are averaged, according to one embodiment. [Figure 7.4] One embodiment shows block prediction using a reduced sample value vector in which a group of four boundary samples is averaged. [Figure 8] This shows a matrix operation performed by a device according to one embodiment. [Figure 9-1] This shows detailed matrix operations performed by an apparatus according to one embodiment. [Figure 9-2] This shows detailed matrix operations performed by an apparatus according to one embodiment. [Figure 10] This describes a detailed matrix operation performed by the device using offset and scaling parameters according to one embodiment. [Figure 11] This describes a detailed matrix operation performed by the device using offset and scaling parameters according to one embodiment. [Figure 12] This shows one embodiment of predicting a predetermined second color component block of a picture. [Figure 13] This illustrates one embodiment of a different prediction of a given second color component block of a picture. [Modes for carrying out the invention]
[0025] Equal or equivalent elements, or elements with equal or equivalent functionality, are represented in the following description by equal or equivalent reference numerals, even if they occur in different figures.
[0026] The following description provides several details to give a more complete description of embodiments of the present invention. However, it will be apparent to those skilled in the art that embodiments of the present invention can be carried out without those specific details. In other instances, known structures and devices are shown in the form of undetailed block diagrams to avoid obscuring embodiments of the present invention. In addition, features of different embodiments described later herein may be combined with each other unless otherwise specifically mentioned.
[0027] 1. Introduction The following describes examples, embodiments, and aspects of different inventions. At least some of these examples, embodiments, and aspects refer to methods and / or apparatus for video coding and / or performing block-based prediction, among other things, such as for video applications and / or virtual reality applications, using linear or affine transforms for, for example, adjacent sample reduction and / or optimization of video distribution (broadcast, streaming, file playback, etc.).
[0028] Furthermore, examples, embodiments, and aspects may refer to High Efficiency Video Coding (HEVC) or subsequent methods. Further embodiments, examples, and aspects are defined by the appended claims.
[0029] It should be noted that any embodiments, examples, and aspects defined by the claims may be supplemented by any of the details (features and functions) described in the following chapters.
[0030] Furthermore, the embodiments, examples, and aspects described in the following chapters may be used individually and may be supplemented by any of the features in other chapters or any of the features included in the claims.
[0031] Furthermore, it should be noted that the individual examples, embodiments, and aspects described herein can be used individually or in combination. Therefore, details can be added to each of the individual aspects described above without adding details to another one of the examples, embodiments, and aspects.
[0032] It should also be noted that this disclosure explicitly or implicitly describes the features of the decoding and / or encoding system and / or method.
[0033] Furthermore, the features and functions disclosed herein in relation to the methods can also be used in the apparatus. Moreover, any features and functions disclosed herein in relation to the apparatus can also be used in the corresponding methods. In other words, the methods disclosed herein can be supplemented by any of the features and functions described in relation to the apparatus.
[0034] Furthermore, any of the features and functions described herein can be implemented in hardware or software, or using both hardware and software, as described in the "Alternative Implementation Methods" section.
[0035] Furthermore, some of the features highlighted in parentheses ("(...)" or "[...]") may be considered optional in some examples, embodiments, or aspects.
[0036] 2 Encoders, Decoders The following describes various examples that can help achieve more efficient compression when using block-based prediction. In some examples, high compression efficiency is achieved by using a set of intra-prediction modes. Intra-prediction modes may be provided in addition to, for example, other heuristically designed intra-prediction modes, or exclusively. Other examples even utilize both of the specialties described here. However, as a variation of these embodiments, intra-prediction may be replaced with inter-prediction by using a reference sample in a different picture instead.
[0037] To facilitate understanding of the following examples of this application, the description begins with a presentation of possible encoders and decoders adapted thereto, which can then build upon the outlined examples of this application. Figure 1 shows an apparatus for blockwise encoding a picture 10 into a data stream 12. The apparatus is indicated by reference numeral 14 and may be a static picture encoder or a video encoder. In other words, the picture 10 may be a current picture from a video 16 when the encoder 14 is configured to encode the video 16 containing the picture 10 into a data stream 12, or the encoder 14 may exclusively encode the picture 10 into a data stream 12.
[0038] As described above, the encoder 14 performs encoding in a blockwise or block-based manner. For this purpose, the encoder 14 subdivides the picture 10 into blocks, which are the units into which the encoder 14 encodes the picture 10 into the data stream 12. Examples of possible subdivisions of the picture 10 into blocks 18 are shown in more detail below. In general, the subdivision may end up in blocks 18 of a fixed size or blocks of different block sizes, such as an array of blocks arranged in rows and columns, by starting with the entire picture area of the picture 10 or by using hierarchical multitree subdivision by starting with pre-partitioning of the picture 10 into an array of tree blocks or multitree subdivision, and these examples should not be considered as excluding other possible forms of subdivision of the picture 10 into blocks 18.
[0039] Furthermore, the encoder 14 is a predictive encoder configured to predictively encode the picture 10 into the data stream 12. For a particular block 18, this means that the encoder 14 determines the predictive signal for block 18 and encodes the predictive residual, i.e., the prediction error that causes the predictive signal to deviate from the actual picture content within block 18, into the data stream 12.
[0040] The encoder 14 can support different prediction modes to derive a prediction signal for a particular block 18. In the following example, the important prediction mode is the intra-prediction mode, which depends on which parts of the block 18 are spatially predicted from adjacent already encoded samples of picture 10. The encoding of picture 10 into the data stream 12, and the corresponding decoding procedure thereafter, may be based on a specific encoding order 20 defined among the blocks 18. For example, the encoding order 20 may traverse the blocks 18 in a raster scan order, such as top to bottom row by row, traversing each row from left to right. In the case of hierarchical multitree-based subdivision, raster scan ordering may be applied within each hierarchical level, or depth-first traversal ordering may be applied, i.e., leaf nodes in a block at a particular hierarchical level may precede blocks at the same hierarchical level that have the same parent block, according to the encoding order 20. Depending on the encoding order 20, adjacent already encoded samples of block 18 may typically be located on one or more edges of block 18. In the examples presented herein, for example, adjacent already encoded samples in block 18 are located above and to the left of block 18.
[0041] The intra-prediction mode is not the only mode supported by the encoder 14. If the encoder 14 is, for example, a video encoder, it may also support an inter-prediction mode in which block 18 is temporarily predicted from a previously encoded picture of video 16. Such an inter-prediction mode may be a motion-compensated prediction mode in which a motion vector is notified to such block 18, indicating the relative spatial offset of the portion from which the predicted signal of block 18 will be derived as a duplicate. Additionally or alternatively, other non-intra-prediction modes may also be available, such as an inter-prediction mode when the encoder 14 is a multi-view encoder, or a non-predictive mode in which the interior of block 18 is encoded as is, i.e., without any prediction.
[0042] Before beginning the description of this application with a focus on the intra-predictive mode, we will describe a more specific example of a possible block-based encoder, namely, a possible implementation of encoder 14 such that, as described with respect to Figure 2, we will then present two corresponding examples of decoders that fit Figures 1 and 2, respectively.
[0043] Figure 2 shows a possible implementation of the encoder 14 of Figure 1, i.e., an implementation in which the encoder is configured to use transform coding to encode the predicted residuals, but this is merely an example and the application is not limited to predicted residual coding of that classification. According to Figure 2, in order to obtain the predicted residual signal 26 to be encoded into the data stream 12 by the predicted residual encoder 28, the encoder 14 includes a subtractor 22 configured to subtract the corresponding predicted signal 24 from the inbound signal, i.e., picture 10, or block-based, from the current block 18. The predicted residual encoder 28 consists of an irreversible coding stage 28a and a reversible coding stage 28b. The irreversible stage 28a includes a quantizer 30 that receives the predicted residual signal 26 and quantizes samples of the predicted residual signal 26. As already stated above, this embodiment uses transform coding of the predicted residual signal 26, and accordingly, the irreversible coding stage 28a includes a transform stage 32 connected between the subtractor 22 and the quantizer 30 to transform such spectrally decomposed predicted residuals 26 by quantization of the quantizer 30 performed on the transformed coefficients that present the residual signal 26. The transform may be DCT, DST, FFT, Hadamard transform, etc. The transformed and quantized predicted residual signals 34 are then subject to reversible coding by a reversible coding stage 28b, which is an entropicorder that entropi codes the quantized predicted residual signals 34 into a data stream 12. The encoder 14 further includes a predicted residual signal reconstruction stage 36 connected to the output of the quantizer 30 to reconstruct the predicted residual signals from the transformed and quantized predicted residual signals 34 in a manner also available to the decoder, i.e., taking into account the coding loss in the quantizer 30. For this purpose, the predictive residual reconstruction stage 36 includes a dequantizer 38 that performs the inverse of the quantization of the quantizer 30, and a subsequent inverse converter 40 that performs the inverse transform of the transform performed by the converter 32, such as the inverse of spectral decomposition, such as the inverse of any of the examples of the particular transforms described above. The encoder 14 includes an adder 42 that adds the reconstructed predictive residual signal as the output of the inverse converter 40 and the predictive signal 24 to output the reconstructed signal, i.e., the reconstructed sample.This output is fed to the predictor 44 of the encoder 14, which then determines the predicted signal 24 based on it. It is the predictor 44 that supports all the prediction modes already discussed above with respect to Figure 1. Figure 2 also illustrates that if the encoder 14 is a video encoder, the encoder 14 may also include an in-loop filter 46 that filters out a fully reconstructed picture that, after filtering, forms a reference picture about the predictor 44 with respect to the interpredicted block.
[0044] As already mentioned above, the encoder 14 operates in a block-based manner. For the purposes of the subsequent explanation, the block-based operation in question is the operation of subdividing the picture 10 into blocks, in which an intra-prediction mode is selected from a set or multiple intra-prediction modes supported by the predictor 44 or the encoder 14, and the selected intra-prediction mode is executed individually. However, there may be other classifications of blocks into which the picture 10 is subdivided. For example, the above-mentioned decision on whether the picture 10 is inter-coded or intra-coded may be made at a block granularity or unit that deviates from block 18. For example, the inter / intra mode decision may be made at the level of the coded block into which the picture 10 is subdivided, and each coded block is subdivided into prediction blocks. Each prediction block, by coding a block into which it has been decided that intra-prediction is used, is subdivided into an intra-prediction mode decision. For this purpose, for each of those prediction blocks, it is determined whether the supported intra-prediction modes should be used for that prediction block. These prediction blocks form block 18, which is the block of interest here. Prediction blocks within a coded block associated with interpretation are handled differently by the predictor 44. They are interpreted from a reference picture by determining motion vectors and replicating the prediction signal for this block from the position in the reference picture indicated by the motion vectors. Further subdivision of another block involves subdividing into transformed blocks in the unit in which the transformation by the converter 32 and inverse converter 40 is performed. The transformed block may be, for example, the result of further subdivision of a coded block. Essentially, the examples shown herein should not be treated as limiting, and other examples may exist. For completeness only, it should be noted that the subdivision into coded blocks may also be performed using, for example, multi-tree subdivision, and prediction blocks and / or transformed blocks can also be obtained by further subdividing a coded block using multi-tree subdivision.
[0045] A decoder 54 or apparatus for blockwise decoding that conforms to the encoder 14 of Figure 1 is described in Figure 3. This decoder 54 performs the opposite of the encoder 14, namely, it decodes the picture 10 from the data stream 12 in a blockwise manner and supports multiple intra-prediction modes for this purpose. Decoder 54 may include, for example, a residual provider 156. All other possibilities discussed above with respect to Figure 1 are also valid for decoder 54. For this purpose, decoder 54 may be a still picture decoder or a video decoder, and all prediction modes and predictability are also supported by decoder 54. The difference between encoder 14 and decoder 54 lies primarily in the fact that encoder 14 selects or chooses coding decisions according to certain optimizations, such as minimizing certain cost functions which may depend on the coding rate and / or coding distortion. One of those coding options or coding parameters may involve the selection of an intra-prediction mode to be used for the current block 18 among the available or supported intra-prediction modes. The selected intra-predictive mode may then be communicated by the encoder 14 of the current block 18 in the data stream 12, by the decoder 54 making a selection again using this signaling in the data stream 12 of block 18. Similarly, the subdivision of picture 10 into block 18 may be subject to optimization in encoder 14, and the corresponding subdivision information may be carried in the data stream 12 by decoder 54, which recovers the subdivision of picture 10 into block 18 based on the subdivision information. In summary, decoder 54 may be a block-based and predictive decoder that operates in addition to the intra-predictive mode, and decoder 54 may support other predictive modes, such as inter-predictive mode, if decoder 54 is a video decoder. When decoding, decoder 54 may also use the coding order 20 discussed with respect to Figure 1, and since this coding order 20 is adhered to in both encoder 14 and decoder 54, the same adjacent samples are available for the current block 18 in both encoder 14 and decoder 54.Therefore, in order to avoid unnecessary repetition, the description of the mode of operation of the encoder 14 may also apply to the decoder 54, insofar as it relates to the subdivision of picture 10 into blocks, for example, insofar as it relates to prediction and insofar as it relates to coding the prediction residual. The difference lies in the fact that the encoder 14, by optimization, selects some coding options or coding parameters and then performs prediction and subdivision again, so that the coding parameters derived from the data stream 12 by the decoder 54 are notified within the data stream 12 or inserted into the data stream 12.
[0046] Figure 4 shows a possible implementation of the decoder 54 in Figure 3, i.e., an adaptation to the implementation of the encoder 14 in Figure 1 as shown in Figure 2. Since many elements of the encoder 54 in Figure 4 are identical to those in the corresponding encoder in Figure 2, the same reference code with an apostrophe is used in Figure 4 to indicate those elements. In particular, the adder 42', the optional in-loop filter 46', and the predictor 44' are connected to the predictor loop in the same way as they are in the encoder in Figure 2. The reconstructed, i.e., quantized and retransformed predictor residual signal applied to the adder 42' is derived by a sequence of residual signal reconstruction stages 36' consisting of an entropy decoder 56 that inverts the entropy coding of the entropy encoder 28b, followed by a dequantizer 38' and an inverse converter 40', just as is the case on the encoding side. The output of the decoder is the reconstruction of picture 10. Reconstruction of picture 10 may be available at the output of adder 42', or alternatively, directly at the output of in-loop filter 46'. To improve picture quality, some post-filters may be arranged at the output of the decoder to subject picture 10 reconstruction to some post-filtering, but this option is not described in Figure 4.
[0047] Again, regarding Figure 4, the explanations presented above for Figure 2 should also be valid for Figure 4, except that the encoder simply performs an optimization task and related decisions regarding the encoding options. However, all explanations regarding block subdivision, prediction, dequantization, and retransformation are also valid for the decoder 54 in Figure 4.
[0048] 3. ALWIP (Affine Linear Weighted Intra Predictor) While ALWIP is not always necessary to implement the technologies discussed here, several non-restrictive examples of ALWIP are discussed here.
[0049] This application relates, in particular, to the concept of an improved block-based prediction mode for block-wise picture coding usable in video codecs such as HEVC or HEVC successors. The prediction mode may be an intra-prediction mode, but theoretically, the concept described herein can also be applied to an inter-prediction mode where the reference sample is part of another picture.
[0050] A block-based predictive concept is needed to enable efficient implementations, such as hardware-friendly implementations.
[0051] This objective is achieved by the subject matter of the independent claim of the present application.
[0052] Intra-prediction mode is widely used in picture and video coding. In video coding, intra-prediction mode competes with other prediction modes such as inter-prediction modes, including motion-compensated prediction mode. In intra-prediction mode, the current block is predicted based on neighboring samples, i.e., samples that have already been coded as far as the encoder side and decoded as far as the decoder side. Prediction residuals are transmitted in the data stream of the current block, and neighboring sample values are extrapolated into the current block to form the prediction signal of the current block. The better the prediction signal, the lower the prediction residuals, and therefore the fewer bits are required to encode the prediction residuals.
[0053] To be effective, several aspects must be considered in order to form an effective framework for intra-prediction in a block-wise picture coding environment. For example, the more intra-prediction modes supported by the codec, the greater the consumption of side information rate to notify the decoder of the selection. On the other hand, the set of supported intra-prediction modes must be able to provide good prediction signals, i.e., prediction signals that yield low prediction residuals.
[0054] Disclosed below, as a comparative embodiment or basic example, is a device (encoder or decoder) for blockwise decoding of pictures from a data stream. This device supports at least one intra-prediction mode in which the prediction signal for a given size block of the picture is determined by applying a first template of adjacent samples to the current block to an affine linear predictor, called an affine linear weighted intra-predictor (ALWIP).
[0055] The device may have at least one of the following characteristics (the same may apply to a method or other technique, such as being implemented in a non-temporary memory unit that stores instructions, which, when performed by a processor, causes the processor to carry out the method and / or acts as a device).
[0056] 3.1 Predictors can complement other predictors. The intra-prediction modes, which may form the subject of implementation improvements described below, may complement other intra-prediction modes of the codec. Therefore, they can complement the DC prediction mode, plane prediction mode, or angle prediction mode defined in each HEVC codec and the JEM reference software. Hereafter, the latter three types of intra-prediction modes will be referred to as conventional intra-prediction modes. Therefore, for a given block of intra-mode, the decoder needs to analyze a flag indicating whether one of the intra-prediction modes supported by the device is used.
[0057] 3.2 Multiple Proposed Prediction Modes A device can have multiple ALWIP modes. Therefore, if the decoder knows that one of the ALWIP modes supported by the device will be used, the decoder needs to parse additional information indicating which of the ALWIP modes supported by the device will be used.
[0058] The signal transmission of supported modes may have the characteristic that some ALWIP modes require fewer bins to encode than others. Which of these modes requires fewer bins and which require more may depend on information that can be extracted from the already decoded bitstream or on information that can be modified beforehand.
[0059] 4. Several aspects Figure 2 shows a decoder 54 for decoding a picture from a data stream 12. The decoder 54 may be configured to decode a given block 18 of the picture. In particular, the predictor 44 may be configured to map a set of P adjacent samples adjacent to a given block 18 to a set of Q predicted values for the samples in the given block, using a linear or affine linear transformation (e.g., ALWIP).
[0060] As shown in Figure 5, a given block 18 contains Q predicted values (which become the "predicted values" at the end of the operation). If block 18 has M rows and N columns, then Q = M·N. The Q values of block 18 may be in a spatial domain (e.g., pixels) or a transformation domain (e.g., DCT, discrete wavelet transform, etc.). The Q values of block 18 can generally be predicted based on P values obtained from adjacent blocks 17a-17c that are adjacent to block 18. The P values of adjacent blocks 17a-17c may be in the closest position to block 18 (e.g., adjacent). The P values of adjacent blocks 17a-17c have already been processed and predicted. The P values are shown as the values of parts 17'a-17'c, distinguishing them from the blocks they are part of (in some examples, 17'b is not used).
[0061] As shown in Figure 6, to perform predictions, it is possible to operate with a first vector 17P having P entries (each entry associated with a specific position within the adjacent portion 17'a~17'c), a second vector 18Q having Q entries (each entry associated with a specific position within block 18), and a mapping matrix 17M (each row associated with a specific position within block 18, and each column associated with a specific position within the adjacent portion 17'a~17'c). Thus, the mapping matrix 17M performs predictions of the P values in the adjacent portion 17'a~17'c to the values in block 18 according to a predetermined mode. Therefore, the entries in the mapping matrix 17M can be understood as weight coefficients. In the following text, the sign 17a~17c is used instead of 17'a~17'c to refer to the adjacent portion of the boundary.
[0062] In this technical field, several conventional modes are known, including DC modes, planar modes, and 65 directional prediction modes. For example, 67 modes may be known.
[0063] However, it has been noted that it is also possible to use various modes called linear or affine linear transformations. A linear or affine linear transformation includes P·Q weight coefficients, of which at least 1 / 4P·Q are non-zero weight values, and for each of the Q predicted values, it includes a set of P weight coefficients associated with each predicted value. When the series are arranged vertically according to the raster scan order between samples of a given block, they form an envelope that is nonlinear in all directions.
[0064] It is possible to map the P positions of adjacent values 17'a~17'c (template), the Q positions of adjacent samples 17'a~17'c, and the P*Q weight coefficients of matrix 17M. The plane is an example of the envelope of a sequence for DC transformation (it is a plane for DC transformation). Since the envelope is obviously a plane, it is excluded by the definition of linear or affine linear transformation (ALWIP). Another example is a matrix that results in the emulation of angular modes, where the envelope is excluded from the definition of ALWIP and, simply put, looks like a hill extending diagonally from top to bottom along the direction in the P / Q plane. Planar modes and directional predictive modes of 65 have different envelopes, but the envelope is linear in at least one direction, i.e., all directions for the exemplified DC, and in the hill direction for the angular modes, for example.
[0065] Conversely, the envelopes of linear or affine transformations are not linear in all directions. It is understood that such types of transformations may be optimal for making predictions for block 18 in some situations. Note that it is preferable that at least 1 / 4 of the weight coefficients are non-zero (i.e., at least 25% of the P*Q weight coefficients are non-zero).
[0066] The weight coefficients may be independent of each other, following the usual mapping rules. Therefore, matrix 17M may be such that the values of its entries do not have any obvious, recognizable relationships. For example, the weight coefficients cannot be described by any analytical or differential function.
[0067] In the example, the ALWIP transformation is the average of the maximum cross-correlations between a first set of weight coefficients associated with each predicted value and a second set of weight coefficients associated with the other predicted values, or an inverted version of the latter set, which may be lower than a given threshold (e.g., 0.2 or 0.3 or 0.35 or 0.1, e.g., a threshold in the range of 0.05 to 0.035), even if it leads to a higher maximum. For example, for each row join (i1,i2) of the ALWIP matrix 17M, the cross-correlations can be calculated by multiplying the P-value of row i1 by the P-value of row i2. For each cross-correlation obtained, the maximum value can be obtained. Thus, the mean (average) can be obtained over the entire matrix 17M (i.e., the maximum cross-correlations in all combinations are averaged). The threshold can then be, for example, 0.2 or 0.3 or 0.35 or 0.1, e.g., a threshold in the range of 0.05 to 0.035.
[0068] The P adjacent samples in blocks 17a to 17c may be located along a one-dimensional path extending along the boundary of a given block 18 (e.g., 18c, 18a). For each of the Q predicted values in a given block 18, a set of P weight coefficients associated with each predicted value may be ordered to traverse the one-dimensional path in a given direction (e.g., left to right, top to bottom, etc.).
[0069] In the example, the ALWIP matrix 17M may be off-diagonal or non-block diagonal.
[0070] An example of an ALWIP matrix 17M for predicting a 4x4 block 18 from four already predicted adjacent samples may be as follows: { {37,59,77,28}, {32,92,85,25}, {31,69,100,24}, {33,36,106,29}, {24,49,104,48}, {24,21,94,59}, {29,0,80,72}, {35,2,66,84}, {32,13,35,99}, {39,11,34,103}, {45,21,34,106}, {51,24,40,105}, {50,28,43,101}, {56,32,49,101}, {61,31,53,102}, {61,32,54,100} }.
[0071] (Here, {37, 59, 77, 28} are the first row. {32, 92, 85, 25} are the second row. {61, 32, 54, 100} are the 16th row of matrix 17M). Matrix 17M has dimensions 16x4 and contains 64 weight coefficients (as a result of 16*4=64). This is because matrix 17M has dimensions QxP, where Q=M*N is the number of samples in the predicted block 18 (block 18 is a 4x4 block) and P is the number of samples in the already predicted samples. Here, M=4, N=4, Q=16 (as a result of M*N=4*4=16), and P=4. The matrix is off-diagonal and non-block diagonal and is not described by any particular rule.
[0072] As can be seen, less than 1 / 4 of the weight coefficients are 0 (in the case of the matrix above, one of the 64 weight coefficients is zero). The envelope formed by these values, when arranged vertically according to the raster scan order, forms a nonlinear envelope in all directions.
[0073] Even if the above explanation is discussed primarily with reference to a decoder (e.g., decoder 54), the same can be done in an encoder (e.g., encoder 14).
[0074] In some examples, for each block size (within a set of block sizes), the ALWIP transformations of the intra-prediction modes in a second set of intra-prediction modes for each block size are different from one another. Additionally or alternatively, the cardinality of the second set of intra-prediction modes for a block size within a set of block sizes can match, but the associated linear or affine linear transformations of the intra-prediction modes in the second set of intra-prediction modes for different block sizes can be made inexchangeable by scaling.
[0075] In some examples, an ALWIP translation may be defined as having "nothing in common" with a conventional translation (for example, an ALWIP translation may be considered to have "nothing in common" with a corresponding conventional translation even if it is mapped via one of the above mappings).
[0076] In one example, ALWIP mode is used for both the luminous and chroma components, while in another example, ALWIP mode is used for the luminous component but not for the chroma component.
[0077] 5. Affine linear weighted intra-prediction mode with faster encoder speed (e.g., Test CE3-1.2.1)
[0078] 5.1 Description of method or apparatus The affine linear weighted intra-prediction (ALWIP) mode tested in CE3-1.2.1 can be the same as the one proposed in JVET-L0199 under test CE3-2.2.2, except for the following changes.
[0079] • Multiple Reference Lines (MRL) intra-prediction, particularly the harmonization of encoder estimation and signaling, i.e., MRL is not combined with ALWIP, and transmission of MRL indices is restricted to non-ALWIP blocks.
[0080] Subsampling is mandatory for all blocks W×H ≥ 32×32 (previously it was optional for 32×32). Therefore, additional testing in the encoder and transmission of the subsampling flag have been eliminated.
[0081] ALWIP for 64×N and N×64 blocks (N≦32) is added by downsampling to 32×N and N×32 respectively and applying the corresponding ALWIP mode.
[0082] Furthermore, test CE3-1.2.1 includes the following encoder optimizations for ALWIP.
[0083] Combinatorial Mode Estimation: Conventional and ALWIP modes use a shared Hadamard candidate list for complete RD estimation. That is, ALWIP mode candidates are added to the same list as conventional (and MRL) mode candidates based on their Hadamard costs.
[0084] EMT Intrafast and PB Intrafast are supported in the combined mode list and have additional optimizations to reduce the number of full RD checks.
[0085] Only the MPMs in the available left and top blocks are added to the list for the complete RD estimation of ALWIP, following the same methodology as in the conventional mode.
[0086] 5.2 Complexity Assessment In test CE3-1.2.1, generating the predicted signal required up to 12 multiplications per sample, excluding calculations that invoke the discrete cosine transform. Furthermore, a total of 136,492 parameters were needed, each of 16 bits. This corresponds to 0.273 megabytes of memory.
[0087] 5.3 Experimental Results The test evaluation was performed using VTM software version 3.0.1 for intranet-only (AI) and random access (RA) configurations, according to common test conditions JVET-J1010[2]. The corresponding simulations were run on an Intel Xeon cluster (E5-2697Av4, AVX2 on, Turbo Boost off) with Linux® OS and GCC 7.2.1 compiler.
[0088] [Table 1] [Table 2]
[0089] 5.4 Affine linear weighted intra-prediction with complexity reduction (e.g., Test CE3-1.2.2) The techniques tested in CE2 are related to the “affine linear intraprediction” described in JVET-L0199[1], but are simplified in terms of memory requirements and computational complexity, as follows:
[0090] Only three different sets of prediction matrices (e.g., S0, S1, S2, see below) and bias vectors (e.g., to provide offset values) can exist to cover all block shapes. As a result, the number of parameters is reduced to 14,400 10-bit values, which requires less memory than storing them in 128 × 128 CTUs.
[0091] The input and output sizes of the predictor can be further reduced. Furthermore, instead of transforming the boundary via DCT, averaging or downsampling can be performed on the boundary samples, and linear interpolation can be used instead of inverse DCT for generating the predicted signal. Therefore, generating the predicted signal may require up to four multiplications per sample.
[0092] 6. Example This section describes how to perform several predictions (for example, as shown in Figure 6) using ALWIP prediction.
[0093] As a general rule, referring to Figure 6, in order to obtain the Q=M*N values of the predicted M×N block 18, the multiplication of Q*P samples from the QxP ALWIP prediction matrix 17M and P samples from the Px1 neighbor vector 17P should be performed. Therefore, in general, to obtain each of the Q=M*N values of the predicted MxN block 18, at least the multiplication of P=M+N values is required.
[0094] These multiplications have a very undesirable effect. The dimension P of the boundary vector 17P generally depends on the number M+N of boundary samples (bins or pixels) 17a, 17c adjacent to (e.g., very close to) the predicted M×N block 18. This means that if the size of the predicted block 18 is large, the number of boundary pixels (17a, 17c) M+N will be large accordingly, and therefore the dimension Px1 boundary vector P=M+N, Q×P ALWIP prediction matrix 17M, the length of each column, and thus the required number of multiplications (generally Q=M*N=W*H, where W (width) is another symbol for N, H (height) is another symbol for M, and P is P=M+N=H+W if the boundary vector is formed by only one row and / or one column of samples).
[0095] This problem is generally exacerbated in microprocessor-based systems (or other digital processing systems) by the fact that multiplication is generally a power-consuming operation. A large number of multiplications carried over a very large number of samples across a large number of blocks results in a waste of computational power, which is generally undesirable.
[0096] Therefore, it is desirable to reduce the number of multiplications Q*P required to predict MxN block 18.
[0097] It is understood that by intelligently selecting an operation that is easier to process instead of multiplication, the computational power required for each intra prediction of each predicted block 18 can be reduced in some form.
[0098] In particular, referring to FIGS. 7.1 to 7.4, an encoder or decoder uses a plurality of adjacent samples (e.g., 17a, 17c) to (e.g., in step 811), (e.g., by averaging or downsampling), reduce the plurality of adjacent samples (e.g., 17a, 17c) to obtain a set of reduced sample values with a lower number of samples compared to the plurality of adjacent samples, and then (e.g., in step 812) subject the reduced set of sample values to a linear transformation or an affine linear transformation to obtain a predicted value of a predetermined sample of a predetermined block, thereby predicting a predetermined block (e.g., 18) of a picture.
[0099] In some cases, the decoder or encoder can also derive a predicted value of a further sample of a predetermined block based on a predetermined sample and predicted values of a plurality of adjacent samples, for example, by interpolation. Thus, an upsampling strategy can be obtained.
[0100] In the example, it is possible to perform some averaging on the samples at the boundary 17 to reach the reduced set 102 of samples with a reduced number of samples (FIGS. 7.1 to 7.4) (e.g., in step 811) (at least one of the samples of the reduced number of samples 102 may be the average of two samples of the original boundary samples or a selection of the original boundary samples). For example, if the original boundary has samples of P = M + N, the reduced set of samples can have P red <P such that M red <M and N red <N, at least one of which is P red = M red + N red Thus, the boundary vector 17P actually used for prediction (e.g., in step 812b) does not have a Px1 entry, and Pred <P is, P red has x1 entries. Similarly, the ALWIP prediction matrix 17M selected for prediction does not have QxP dimensions and has at least P red <P(M red <M and N red <for at least one of N) the number of elements of the matrix is reduced to QxP red (or Q red xP red , see below).
[0101] In some examples (e.g., Figures 7.2, 7.3), the block obtained by ALWIP (at step 812) is
Number
Number
Number
Number
Number
Number
[0102] These techniques reduce the number of multiplications (Q) in matrix multiplication. red *P red or Q*P red While including ), both the initial reduction (e.g., averaging or downsampling) and the final transformation (e.g., interpolation) can be performed by reducing (or avoiding) multiplication, which can be advantageous. For example, downsampling, averaging and / or interpolation may be performed by employing binary operations that do not require computational power, such as addition and shifting (e.g., in steps 811 and / or 813).
[0103] Furthermore, addition, which can be easily performed without much computational effort, is a very simple operation.
[0104] This shift operation can be used, for example, to average two boundary samples and / or to interpolate two samples (support values) of the reduced prediction block (or obtained from the boundary) in order to obtain the final prediction block. (Two sample values are needed for interpolation. While there are always two given values within a block, as shown in Figure 7.2, to interpolate samples along the left and top boundaries of the block, there is only one given value, so the boundary samples are used as support values for interpolation.)
[0105] The following two-step procedure may be used. First, sum the values of the two samples. Next, we halve the total value (for example, by right-shifting).
[0106] Alternatively, the following is possible: First, halve each sample (for example, by left-shifting). Next, sum the values of the two half samples.
[0107] Since it is only necessary to select one sample amount for a group of samples (e.g., adjacent samples to each other), simpler operations can be performed when downsampling (e.g., at step 811).
[0108] Therefore, it is possible to define a technique (or techniques) for reducing the number of multiplications performed here. Some of these techniques may be based, among other things, on at least one of the following principles. Even if the actually predicted size of block 18 is MxN, the block is reduced (in at least one of the two dimensions) to have a reduced size of Q red xP red for the ALWIP matrix (
Number
Number
Number
number
number
number
[0109] As shown in the example in Figure 7.1, a 4x4 block 18 (M=4, N=4, Q=M*N=16) is predicted, and the neighbors 17 of sample 17a (a vertical matrix containing 4 predicted samples) and 17c (a horizontal row containing 4 predicted samples) have already been predicted in previous iterations (the neighbors 17a and 17c may be collectively represented by 17). Priorily, by using the formula shown in Figure 5, the prediction matrix 17M should be a QxP=16x8 matrix (Q=M*N=4*4 and P=M+N=4+4=8), and the boundary vector 17P should have dimensions of 8x1 (by P=8). However, this would necessitate performing 8 multiplications for each of the 16 samples in the predicted 4x4 block of 18, resulting in a total of 16*8=128 multiplications (it should be noted that the average number of multiplications per sample is a good indicator of computational complexity. Conventional intra-prediction requires 4 multiplications per sample, which increases the associated computational effort. Therefore, it is possible to use this as an upper limit for ALWIP to ensure that the complexity is reasonable and does not exceed the complexity of conventional intra-prediction).
[0110] Nevertheless, by using this technology, in step 811, the number of adjacent samples 17a and 17c to the predicted block 18 is reduced from P to P redIt is understood that it is possible to reduce to . In particular, it is possible to average adjacent boundary samples (17a, 17c) (e.g., at 100 in FIG. 7.1) to obtain a reduced boundary 102 having two horizontal rows and two vertical columns. Thus, it is understood that the operation as block 18 is a 2x2 block (the reduced boundary is formed by the average value). Alternatively, it is possible to perform downsampling. Thus, two samples are selected for row 17c and two samples are selected for column 17a. Thus, the horizontal row 17c is processed as having two samples (e.g., averaged samples) instead of having four original samples, and the vertical column 17a, which originally has four samples, is processed as having two samples (e.g., averaged samples). After subdividing row 17c and column 17a into groups 110 of two samples each, it can also be understood that one single sample is maintained (e.g., the average of the samples in group 110 or a simple selection between the samples of group 110). Thus, a so-called reduced set 102 of sample values is obtained by the set 102 having only four samples (M red =2, N red =2, P red =M red +N red =4, P red <P).
[0111] It is understood that operations (such as averaging or downsampling 100) can be performed without performing too many multiplications at the processor level. The averaging or downsampling 100 performed in step 811 can be easily obtained by simple and non-power-consuming operations such as addition and shift.
[0112] At this point, it is understood that the reduced set of sample values 102 can be subjected to a linear or affine linear (ALWIP) transformation 19 (for example, using a prediction matrix such as matrix 17M in Figure 5). In this case, the ALWIP transformation 19 directly maps the four samples 102 to the sample values 104 of block 18. In this case, interpolation is not necessary.
[0113] In this case, the dimension of the ALWIP matrix 17M is QxP red = 16x4, which follows the fact that all Q=16 samples of the predicted block 18 are obtained directly by ALWIP multiplication (no interpolation is needed).
[0114] Therefore, in step 812a, dimension Q × P red A suitable ALWIP matrix 17M having A is selected. The selection can be based, for example, at least in part on signaling from the data stream 12. The selected ALWIP matrix 17M is A k This can be shown as follows, where k can be understood as an index that can be notified in data stream 12 (in some cases the matrix is
number
[0115] In step 812b, the selected QxP red ALWIP matrix 17M(A k (Also shown as) and P red Multiplication is performed between x1 and the boundary vector 17P.
[0116] In step 812c, the offset value (for example, b k ) can be added to all acquired values 104 of vector 18Q obtained by ALWIP, for example. The offset value (b k Or in some cases
number
[0117] Therefore, we will now resume the comparison between using and not using this technology. Without this technology, Block 18 to be predicted, having dimensions M=4 and N=4. The Q to be predicted is M*N=4*4=16 values. P = M + N = 4 + 4 = 8 boundary samples. For each of the 16 values of the target Q to be predicted, P is multiplied 8 times. P*Q = 8*16 = 128 total number of multiplications. Having this technology, Block 18 to be predicted, having dimensions M=4 and N=4. Finally, the values to be predicted are Q = M * N = 4 * 4 = 16 values. Dimension of the boundary vector reduction: P red =M red +N red =2+2=4; For each of the 16 Q values predicted by ALWIP, P red = 4 multiplications, P red *Q = 4 * 16 = 64 total multiplications (half of 128!) The ratio of the number of multiplications to the number of final values obtained is P red *Q / Q=4, which is half the multiplication of P=8 for each predicted sample!
[0118] To make it understandable, it is possible to obtain the appropriate value in step 812 by relying on simple and computationally inefficient operations such as averaging (and in some cases, addition and / or shifting and / or downsampling).
[0119] Referring to Figure 7.2, the predicted block 18 is an 8x8 block (M=8, N=8) with 64 samples. Here, a priori, the prediction matrix 17M must have a size of QxP=64x16 (Q=M*N=8*8=64, given by M=8 and N=8, and Q=64 given by P=M+N=8+8=16). Thus, a priori, for each of the Q=64 samples in the 8x8 block 18 to be predicted, P=16 multiplications are required, resulting in 64*16=1024k multiplications for the entire 8x8 block 18!
[0120] However, as can be seen in Figure 7.2, instead of using all 16 samples of the boundary, we can provide a method 820 in which only 8 values are used (e.g., 4 in the horizontal boundary row 17c and 4 in the vertical boundary column 17a between the original samples of the boundary). From the boundary row 17c, 4 samples can be used instead of 8 (e.g., they may be a 2x2 mean and / or a selection of one of the two samples). Thus the boundary vector is not a Px1 = 16x1 vector, but P red x1 = 8x1 is the only vector (P red =M red +N red =4+4). Instead of the original P=16 samples, P red It is understood that it is possible to select or average (e.g., two of each) samples from horizontal row 17c and vertical column 17a so that there are only 8 boundary values, thereby forming a reduced set 102 of sample values. This reduced set 102 makes it possible to obtain a reduced version of block 18, and the reduced version is Q red =M red *N red =4*4=16 samples (instead of Q=M*N=8*8=64). Size Mred xN red It is possible to apply the ALWIP matrix to predict blocks having a 4x4 grid. The reduced version of block 18 includes the sample shown in gray in scheme 106 in Figure 7.2. The samples shown in gray squares (including samples 118' and 118") are the Q obtained in step 812. red = Forms a 4x4 reduced block with 16 values. The 4x4 reduced block is obtained by applying the linear transformation 19 in the target step 812. After obtaining the values of the 4x4 reduced block, it is possible to obtain the values of the remaining samples (samples shown as white samples in scheme 106) for example by interpolation.
[0121] Regarding method 810 in Figure 7.1, method 820 predicts the remaining QQ of the MxN=8x8 block 18. red Step 813 may further include deriving the predicted values for 64-16=48 samples (white squares) by, for example, interpolation. The remaining QQ red =64-16=48 samples were obtained directly by interpolation (for example, interpolation can also use the values of boundary samples). red =This can be obtained from 16 samples. As seen in Figure 7.2, samples 118' and 118'' were obtained in step 812 (indicated by the gray squares), while sample 108' (which is midway between samples 118' and 118'' and is indicated by the white square) is obtained in step 813 by interpolation between samples 118' and 118''. It is understood that interpolation can also be obtained by operations similar to those for averaging, such as shifting and adding. Therefore, in Figure 7.2, the value 108' can generally be determined as an intermediate value between the value of sample 118' and the value of sample 118'' (it could be the average).
[0122] By performing interpolation, in step 813, we can also arrive at the final version of block 18, MxN=8x8, based on the multiple sample values shown in 104.
[0123] Therefore, the comparison between using and not using this technology is Without this technology, The prediction target block 18 has dimensions M = 8, N = 8, and Q = M * N = 8 * 8 = 64 samples within the prediction target block 18, P = M + N = 8 + 8 = 16 samples within the boundary 17, Each of the Q = 64 values to be predicted is multiplied P = 16 times, The total number of multiplications is P * Q = 16 * 64 = 1028 multiplications The ratio of the number of multiplications to the number of final values obtained is P * Q / Q = 16. With this technology, The prediction target block 18 has dimensions M = 8, N = 8 The last Q = M * N = 8 * 8 = 64 values to be predicted. However, the Q used red xP red The ALWIP matrix has P red = M red + N red Q red = M red * N red M red = 4, N red = 4, and P red < P within the boundary, P red = M red + N red = 4 + 4 = 8 samples, Q for the 4x4 reduced block to be predicted red For each of the Q = 16 values, P red = 8 multiplications (formed by the gray squares in scheme 106), P red * Q red = 8 * 16 = 128 multiplications in total (much less than 1024!) The ratio of the number of multiplications to the number of final values obtained is P red * Q red / Q = 128 / 64 = 2 (much less than 16 obtained without this technology!)
[0124] Therefore, the technology presented herein requires eight times less power than the previous one.
[0125] FIG. 7.3 shows another example where the block 18 to be predicted is a rectangular 4×8 block (M = 8, N = 4) having Q = 4 * 8 = 32 samples to be predicted (which can be based on method 820). The boundary 17 is formed by a horizontal row 17c having N = 8 samples and a vertical column 17a having M = 4 samples. Therefore, a priori, the boundary vector 17P has dimension Px1 = 12x1, but the prediction ALWIP matrix needs to be a QxP = 32x12 matrix, and thus Q * P = 32 * 12 = 384 multiplications are required.
[0126] However, for example, it is possible to average or downsample at least 8 samples of the horizontal row 17c to obtain a reduced horizontal row having only 4 samples (e.g., averaged samples). In some examples, the vertical column 17a remains as it is (e.g., without averaging). In total, the dimension of the reduced boundary is P red = 8, and P red < P. Therefore, the boundary vector 17P has dimension P red x1 = 8x1. The ALWIP prediction matrix 17M becomes a matrix having dimension M * N red * P red = 4 * 4 * 8 = 64. The 4×4 reduced block (formed by the gray columns of schema 107) directly obtained in the target step 812 has size Q red = M * N red = 4 * 4 = 16 samples (instead of Q = 4 * 8 = 32 of the original 4x8 block 18 to be predicted). When the reduced 4×4 block is obtained by ALWIP, an offset value b k can be added (step 812c), and interpolation can be performed in step 813. As can be seen in step 813 of FIG. 7.3, the reduced 4×4 block is expanded to the 4×8 block 18, and the value 108' not obtained in step 812 is obtained in step 813 by interpolating the values 118' and 118'' (gray squares).
[0127] Therefore, the comparison between using and not using this technology is Without this technology, Block 18 to be predicted with dimensions M = 4, N = 8, Q = M * N = 4 * 8 = 32 values to be predicted, P = M + N = 4 + 8 = 12 samples within the boundary, 12 multiplications for each of the Q = 32 values to be predicted, Total number of multiplications of P * Q = 12 * 32 = 384 The ratio of the number of multiplications to the number of final values obtained is P * Q / Q = 12. With this technology, Block 18 to be predicted with dimensions M = 4, N = 8, Finally, Q = M * N = 4 * 8 = 32 values to be predicted. However, Q red × P red = 16x8 ALWIP matrix can be used, where M = 4, N red = 4, Q red = M * N red = 16, P red = M + N red = = 4 + 4 = 8, and Pred < P, P samples within the boundary red = M red + N red = 4 + 4 = 8 samples, Q of the reduced block to be predicted red = 16 multiplications for each of the P red = 8 values, Q red * P red = 16 * 8 = 128 multiplications in total (less than 384!) The ratio of the number of multiplications to the number of れる final values obtained is Pred * Qred / Q = 128 / 32 = 4 (much less than れる 12 obtained without this technology!).
[0128] Therefore, using this technology reduces the computational effort to one-third.
[0129] Figure 7.4 shows an example of block 18, which is predicted with dimension MxN=16x16, has 256 predicted values Q=M*N=16*16=256, and has 32 boundary samples P=M+N=16+16=32. This results in a prediction matrix with dimension QxP=256x32, meaning 256*32=8192 multiplications!
[0130] However, by applying method 820, in step 811, it is possible to reduce the number of boundary samples from, for example, 32 to 8 (e.g., by averaging or downsampling), so that for every group of four consecutive samples 120 in row 17a, one single sample remains (e.g., one selected from the four samples, or the mean of the samples). Similarly, for every group of four consecutive samples in column 17c, one single sample remains (e.g., one selected from the four samples, or the mean of the samples).
[0131] Here, the ALWIP matrix 17M is Q red XP red = This is a 64x8 matrix. This is P red This is due to the fact that =8 was selected (by using 8 averaged or selected samples from 32 samples of the boundary) and the fact that the reduced block to be predicted in step 812 is an 8x8 block (in scheme 109, the gray squares are 64).
[0132] Therefore, once 64 samples of the reduced 8x8 block are obtained in step 812, in step 813 the remaining QQ of block 18 to be predicted. red =256-64=192 values can be derived from 104.
[0133] In this case, it is chosen to use all samples in boundary column 17a and only the alternative samples in boundary row 17c to perform interpolation. Other choices may be made.
[0134] In this method, the ratio of the number of multiplications to the number of finally obtained values is Q red *P red / Q = 8 * 64 / 256 = 2, which is much less than the 32 multiplications for each value without using this technology!
[0135] Therefore, the comparison between using and not using this technology is Without this technology, a block 18 to be predicted having dimensions M = 16, N = 16, Q = M * N = 16 * 16 = 256 values to be predicted, P = M + N = 16 + 16 = 32 samples within the boundary, 32 multiplications for each of the Q = 256 values to be predicted, a total of P * Q = 32 * 256 = 8192 multiplications, The ratio of the number of multiplications to the number of finally obtained values is P * Q / Q = 32. With this technology, a block 18 to be predicted having dimensions M = 16, N = 16, the last Q = M * N = 16 * 16 = 256 values to be predicted, However, the ALWIP matrix Q used red ×P red = 64x8 is the sample M predicted by ALWIP red = 4, N red = 4, Q red = 8 * 8 = 64, and P red = M red +N red = 4 + 4 = 8, P red <P, the P within the boundary red = M red +N red = 4 + 4 = 8 samples, the reduced block Q to be predicted red = 8 multiplications for each of the 64 values, red P Q red *P red = 64 * 4 = 256 total multiplications (less than 8192!) The ratio of the number of multiplication operations to the number of final values obtained is P red *Q red / Q = 8 * 64 / 256 = 2 (much less than 32 obtained without this technology!)
[0136] Therefore, the computing power required for this technology is 16 times smaller than that of the conventional technology!
[0137] Therefore, by reducing a plurality of adjacent samples (100, 813) to obtain a reduced set of sample values (102) with fewer sample numbers compared to the plurality of adjacent samples (17), and subjecting the reduced set of sample values (102) to linear or affine linear transformation (19, 17M), it is possible to obtain predicted values of predetermined samples (104, 118’, 188”) of a predetermined block (18) (812), and use a plurality of adjacent samples (17) to predict a predetermined block (18) of a picture
[0138] In particular, it is possible to perform reduction (100, 813) by downsampling a plurality of adjacent samples to obtain a reduced set of sample values (102) with fewer sample numbers compared to the plurality of adjacent samples (17).
[0139] Alternatively, it is possible to perform reduction (100, 813) by averaging a plurality of adjacent samples to obtain a reduced set of sample values (102) with fewer sample numbers compared to the plurality of adjacent samples (17).
[0140] Furthermore, it is possible to derive predicted values of further samples (108, 108’) of a predetermined block (18) (813) based on interpolation on the basis of predicted values of predetermined samples (104, 118’, 118’’) and a plurality of adjacent samples (17).
[0141] A plurality of adjacent samples (17a, 17c) can extend one-dimensionally along two sides of a predetermined block (18) (e.g., in the rightward and downward directions in FIGS. 7.1 to 7.4). A predetermined sample (e.g., the one obtained by ALWIP in step 812) can also be arranged in rows and columns, and can be positioned at every nth position from the samples (112) of the predetermined samples 112 adjacent to two sides of the predetermined block 18 along at least one of the rows and columns.
[0142] Based on the plurality of adjacent samples (17), it is possible to determine a support value (118) for one (118) of the plurality of adjacent positions aligned with each of at least one of the rows and columns. Based on the predicted values of the predetermined samples (104, 118', 118'') and the support values of the adjacent samples (118) aligned with at least one of the rows and columns, it is also possible to derive the predicted values 118 of further samples (108, 108') of the predetermined block (18) by interpolation.
[0143] The predetermined sample (104) may be arranged at every nth position from the adjacent samples (112) along two sides of the predetermined block 18 along a row, and the predetermined sample may be arranged at every mth position from the samples (112) of the predetermined samples adjacent to two sides of the predetermined block (18) along a column, where n, m > 1. In some cases, n = m (e.g., in FIGS. 7.2 and 7.3, the samples 104, 118', 118'' directly obtained by ALWIP in 812 and shown as gray squares are alternately displayed with the samples 108, 108' subsequently obtained in step 813 along the rows and columns).
[0144] Along at least one of the rows (17c) and columns (17a), for example, for each support value, it may be possible to perform the determination of the support value by downsampling or averaging (122) a group (120) of adjacent samples within a plurality of adjacent samples, including the adjacent sample (118) for which each support value is determined. Thus, in Figure 7.4, in step 813, the value of sample 119 can be obtained by using the values of a given sample 118'' (previously obtained in step 812) and adjacent sample 118 as support values.
[0145] Multiple adjacent samples may extend one-dimensionally along two edges of a given block (18). It may be possible to perform reduction (811) by grouping multiple adjacent samples (17) into one or more consecutive groups (110) of adjacent samples and performing downsampling or averaging on each of the one or more groups (110) of adjacent samples having two or three or more adjacent samples.
[0146] In the example, a linear or affine linear transformation is P red *Q red or P red *Q can include weighting coefficients, P red Q is the number of sample values (102) in the reduced set of sample values. red Alternatively, Q is the number of samples in a given block (18). At least 1 / 4P red *Q red or 1 / 4P red *Q weight coefficients are non-zero weight values. P red *Q red or P red *The Q weighting coefficient is Q or Q red For each of the specified samples, a series of P related to each specified sample red Weighting coefficients may be included, and a series of weighting coefficients, when arranged vertically according to the raster scan order between predetermined samples of a given block (18), form an envelope that is nonlinear in all directions.red *Q or P red *Q red The weight coefficients may be independent of each other through a regular mapping rule. The average of the maximum values of the cross-correlation between a first series of weight coefficients associated with each given sample and a second series of weight coefficients associated with all other given samples, or the inverted version of the latter series, is lower than a predetermined threshold, even though the maximum values are higher. The predetermined threshold may be 0.3 [or, in some cases, 0.2 or 0.1]. red The adjacent samples (17) may be located along a one-dimensional path extending along two sides of a given block (18), and there may be Q or Q red For each of the specified samples, a series of P related to each specified sample red The weight coefficients are ordered so as to traverse the one-dimensional path in a predetermined direction.
[0147] 6.1 Description of Method and Apparatus
number
number
[0148] The generation of a predictive signal (e.g., the complete block 18 values) can be based on at least some of the following three steps: 1. Of the boundary samples 17, 102 samples (e.g., 4 samples if W=H=4 and / or 8 samples in other cases) can be extracted by averaging or downsampling (e.g., step 811). 2. Matrix-vector multiplication followed by offset addition can be performed with the averaged samples (or samples remaining from downsampling) as input. The result may be a reduced prediction signal relating to the subsampled set of samples in the original block (e.g., step 812). 3. Predicted signals at the remaining positions can be generated from the predicted signals on the subsampled set, for example by upsampling or by linear interpolation (e.g., step 813).
[0149] Thanks to step 1 (811) and / or step 3 (813), the total number of multiplications required to calculate the matrix-vector product is always
number
[0150] In some examples, the matrix (e.g., 17M) and offset vector (e.g., b) required to generate the prediction signal are needed. k ) is a set of matrices (for example, three sets), for example,
number
[0151] In some examples, set
number
number
[0152] In some examples, the set [Number] each having 16 rows and 8 columns and 18 offset vectors each of size 16, for performing the technique according to Figure 7.2 or 7.3 [Number] may have an n1 matrix (e.g., n1 = 8 or n1 = 18 or another numerical value), [Number] and can include (e.g., consist of). This set of matrices and offset vectors of S1 may be used for blocks of sizes 4x8, 4x16, 4x32, 4x64, 16x4, 32x4, 64x4, 8x4, and 8x8. Further, [Number] having a size of [Number] Blocks, i.e., blocks of size 4x16 or 16x4, 4x32 or 32x4, and 4x64 or 64x4, may also be used. The 16×8 matrix refers to a reduced version of block 18, which is a 4×4 block, as obtained in FIGS. 7.2 and 7.3.
[0153] Additionally, or alternatively, the set
Number
Number
Number
[0154] This set of these matrices and offset vectors or a part of the matrices and offset vectors may be used for all other block shapes.
[0155] 6.2 Averaging or Downsampling of Boundaries Here, features regarding step 811 are provided.
[0156] As described above, boundary samples (17a, 17c) can be averaged and / or downsampled (e.g., from P samples to P red <P samples).
[0157] In the first step, the input boundary
number
number
number
number
number
number
[0158] In the case of a 4x4 block,
number
number
number
[0159] In all other cases (for example, for blocks with a width or height different from 4), the block width W is
number
number
number
[0160] In yet another case, the boundary can be downsampled (for example, by selecting one specific boundary sample from a group of boundary samples) to reach a reduction in the number of samples. For example,
number
number
number
[0161] Two reduced boundaries
number
number
number
[0162] Here, if it is a mode (or the number of matrices in a set of matrices),
number
[0163] It is possible to define the case of mode -17, which corresponds to the transposed mode of mode ≥ 18.
number
[0164] Therefore, according to a specific state (one state: mode < 18; another state: mode ≥ 18), the predicted values of the output vector are obtained in different scan orders (e.g., one scan order (e.g., one scan order)).
number
number
[0165] Other strategies can also be employed. In other examples, the mode index "mode" is not necessarily limited to the range of 0 to 35 (other ranges may be defined). Furthermore, each of the three sets S0, S1, and S2 does not necessarily have 18 matrices (therefore, instead of an expression like mode ≥ 18, it is the number of matrices in each set of matrices S0, S1, and S2).
number
[0166] Mode and transpose information is not necessarily stored and / or transmitted as a single combination mode index "mode". In some examples, it may be explicitly signaled as a transpose flag and matrix index (0-15 for S0, 0-7 for S1, 0-5 for S2).
[0167] In some cases, a combination of the transpose flag and the matrix index may be interpreted as a set index. For example, there may be one bit that acts as the transpose flag and several bits that indicate the matrix index, which is collectively referred to as the "set index".
[0168] 6.3 Generation of Reduced Prediction Signals by Matrix-Vector Multiplication Here, features relating to step 812 are provided.
[0169] Shrink input vector
number
number
number
number
number
[0170] Shrinkage prediction signal
number
number
number
number
[0171]
number
number
number
number
[0172] Matrix A and
number
number
number
number
number
number
number
number
[0173] Other strategies can also be employed. In other examples, the mode index "mode" is not necessarily limited to the range of 0 to 35 (other ranges may be defined). Furthermore, each of the three sets S0, S1, and S2 does not necessarily have 18 matrices (therefore, instead of an expression like mode < 18, it is the number of matrices in each set of matrices S0, S1, and S2).
number
[0174] 6.4 Linear interpolation for generating the final predicted signal Here, features relating to step 812 are provided.
[0175] Interpolating subsampled prediction signals in large blocks may require a second version of the averaged boundary. That is,
number
number
number
number
number
number
number
[0176]
number
number
number
[0177] Linear interpolation can be given as follows (other examples are also possible):
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
[0178] This is an example of interpolation that uses a reduced boundary sample for the first interpolation (horizontal or vertical) and the original boundary sample for the second interpolation (vertical or horizontal). Depending on the block size, only the second interpolation may be needed, or no interpolation may be needed at all. If both horizontal and vertical interpolation are needed, the order depends on the width and height of the block.
[0179] However, different techniques can also be implemented; for example, the original boundary samples can be used for both the first and second interpolations, and the order can be fixed, for example, horizontal first, then vertical (otherwise, vertical first, then horizontal).
[0180] Therefore, the interpolation order (horizontal / vertical) and the use of reduced / original boundary samples may be changed.
[0181] 6.5 Explanation of an example of the entire ALWIP process The entire process of averaging, matrix-vector multiplication, and linear interpolation is illustrated for different shapes in Figures 7.1–7.4. Note that the remaining shapes are treated as just one of the illustrated examples. Given a 1.4 × 4 block, ALWIP can take two averages along each axis of the boundary using the technique shown in Figure 7.1, and the resulting four input samples are input to a matrix-vector multiplication. The matrix is:
number
number
number
number
number
[0182] 6.6 Evaluation of the number and complexity of required parameters The parameters required for all possible proposed intra-predictive modes are:
number
[0183] 6.7 Proposed Intra Predictive Mode Signaling For the Luma Block, for example, 35 ALWIP modes have been proposed (other numbers of modes may also be used). For each coding unit (CU) in the intra-mode, a flag is sent in the bitstream indicating whether the ALWIP mode should be applied to the corresponding prediction unit (PU). The signaling of the latter index can be harmonized with the MRL in the same way as in the initial CE test. If the ALWIP mode is applied, the ALWIP mode
number
[0184] Here, the derivation of the MPM may be performed using the intra-modes of the above and left PUs, as follows: Each conventional intra-prediction mode
number
number
number
number
number
number
number
number
[0185] The above PU is available, belongs to the same CTU as the current PU, is in intra mode, and is in conventional intra predictive mode.
number
number
[0186] In all other cases,
number
number
[0187] Finally, three fixed default lists
number
number
[0188] 6.8 Derivation of Adaptive MPM Lists for Conventional Luma and Chroma Intra Prediction Modes The proposed ALWIP mode can be harmonized with the conventional intra-predictive mode's MPM-based coding as follows: The conventional intra-predictive mode's luma and chroma MPM list derivation process is fixed table
number
number
number
[0189] In the case of derivation of the Luma MPM list, ALWIP mode
number
number
[0190] 7. Efficient Implementation Embodiments Let's briefly summarize the above example, as it can form the basis for further extending the embodiments described below.
[0191] Multiple adjacent samples 17a, c are used to predict a given block 18 of picture 10.
[0192] The reduction 100 by averaging multiple neighboring samples is performed to obtain a reduced set 102 of sample values with fewer samples compared to multiple neighboring samples. This reduction is optional in the embodiments herein and generates the so-called sample value vector described below. The reduced set of sample values is subjected to a linear transformation or affine linear transformation 19 to obtain the predicted values of a given sample 104 in a given block. This transformation is obtained by machine learning (ML) and is shown later using matrix A and offset vector b, which should be implemented efficiently.
[0193] Through interpolation, the predicted values of 108 additional samples of a given block are derived based on the predicted values of the given sample and multiple adjacent samples. Theoretically, the results of the affine / linear transformation can be associated with non-fullpel sample locations in block 18, such that all samples in block 18 can be obtained by interpolation according to an alternative embodiment. Interpolation may not even be necessary.
[0194] Multiple adjacent samples may extend one-dimensionally along two sides of a given block, and the given samples may be arranged in rows and columns, and along at least one of the rows and columns, and the given samples may be positioned at every nth position from the sample (112) of the given samples adjacent to two sides of the given block. Based on the multiple adjacent samples, a support value for one of the multiple adjacent positions (118) may be determined for each of at least one of the rows and columns. This is aligned to each of at least one of the rows and columns, and by interpolation, the predicted value for a further sample 108 of the given block may be derived based on the predicted value for the given sample and the support value of the adjacent sample aligned to at least one of the rows and columns. The given sample may be positioned along the rows at every nth position from the sample 112 of the given samples adjacent to two sides of the given block, and the given sample may be positioned along the columns at every mth position from the sample 112 of the given samples adjacent to two sides of the given block, where n, m>1. n=m may also be the case. Along at least one of the rows and columns, the determination of support values may be performed for each support value by averaging (122) a group of neighboring samples 120 within a group of neighboring samples that include the neighboring sample 118 for which each support value is determined. The group of neighboring samples may extend one-dimensionally along two sides of a given block, and the reduction may be performed by grouping the group of neighboring samples into one or more consecutive groups 110 of neighboring samples and performing averaging for each of the one or more groups of neighboring samples having three or more neighboring samples.
[0195] For a given block, the predicted residuals may be transmitted in the data stream. These are derived in the decoder, and the given block is reconstructed using the predicted residuals and predicted values for a given sample. In the encoder, the predicted residuals are encoded into the data stream.
[0196] A picture may be subdivided into multiple blocks of different block sizes, each containing a given block. Next, a linear or affine linear transformation of block 18 is selected based on the width W and height H of the given block. As a result, the linear or affine linear transformation selected for the given block is selected from the first set of linear or affine linear transformations, as long as the width W and height H of the given block are within a first set of width / height pairs, and is selected from the second set of linear or affine linear transformations, as long as the width W and height H of the given block are within a second set of width / height pairs that are different from the first set of width / height pairs. Similarly, it will later become clear that affine / linear transformations can be represented by other parameters, namely the weights of C, and optionally offset and scale parameters.
[0197] The decoder and encoder can be configured to subdivide a picture into multiple blocks of different block sizes, including a predetermined block, and to select a linear or affine linear transformation depending on the width W and height H of the predetermined block, and as a result, the linear or affine linear transformation selected for the predetermined block is As long as the width W and height H of a given block are within the first set of width / height pairs, As long as the width W and height H of a given block are within a second set of width / height pairs that is different from the first set of width / height pairs, the second set of linear or affine linear transformations, and A third set of linear or affine linear transformations is selected, provided that the width W and height H of a given block are within one or more third sets of width / height pairs that are different from the first and second sets of width / height pairs.
[0198] A third set of one or more width / height pairs simply contains one width / height pair W', H', and each linear or affine linear transformation in the first set of linear or affine linear transformations is for converting N' sample values into W'*H' predicted values of the W'xH' array of sample positions.
[0199] Each of the width / height pairs in the first and second sets is W p H p The first width / height is not equal to W. p H p And, H q =W p and W q =H p The second width / height is W q H q It can include this.
[0200] Each of the first and second sets of width / height pairs corresponds to the third width / height pair W p is H p It can further include W p is H p Equal to H p >H q That is the case.
[0201] For a given block, a set index indicating which linear or affine linear transformation from a given set of linear or affine linear transformations should be selected for block 18 may be transmitted in the data stream.
[0202] Multiple adjacent samples can extend one-dimensionally along two sides of a given block, and the reduction can be performed by grouping a first subset of multiple adjacent samples adjacent to a first side of a given block into a first group 110 of one or more consecutive adjacent samples, and a second subset of multiple adjacent samples adjacent to a second side of a given block into a second group 110 of one or more consecutive adjacent samples, and then performing averaging on each of the first and second groups of one or more adjacent samples having three or more adjacent samples in order to obtain a first sample value from the first group and a second sample value for the second group. Next, a linear or affine linear transformation can be selected from a predetermined set of linear or affine linear transformations according to the set index, so that two different states of the set index result in a selection of one of the linear or affine linear transformations from a predetermined set of linear or affine linear transformations, and in order to generate an output vector of predicted values, for a set index assuming a first of two different states in the form of a first vector, the reduced set of sample values can be subjected to a predetermined linear or affine linear transformation, and the predicted values of the output vector can be distributed to predetermined samples of predetermined blocks along a first scan order, and for a set index assuming a second of two different states in the form of a second vector, the first and second vectors are different such that, in order to generate an output vector of predicted values, the component input by one of the first sample values of the first vector is input by one of the second sample values of the second vector, and the component input by one of the second sample values of the first vector is input by one of the first sample values of the second vector, and the predicted values of the output vector can be distributed to predetermined samples of predetermined blocks transposed with respect to the first scan order, along a second scan order.
[0203] Each linear or affine linear transformation in a first set of linear or affine linear transformations may be for converting N1 sample values into w1*h1 predicted values for a w1xh1 array of sample positions, each linear or affine linear transformation in a first set of linear or affine linear transformations may be for converting N2 sample values into w2*h2 predicted values for a w2xh2 array of sample positions, for a first predetermined one of the first set of width / height pairs, w1 may exceed the width of the first predetermined width / height pair, or h1 may exceed the height of the first predetermined width / height pair, and for a second predetermined one of the first set of width / height pairs, w1 may not exceed the width of the second predetermined width / height pair, and h1 may not exceed the height of the second predetermined width / height pair. Shrinking multiple adjacent samples (100) to obtain a shrunken set (102) of sample values by averaging may be done such that the shrunken set of sample values 102 has N1 sample values, where a given block is a first predetermined width / height pair and a given block is a second predetermined pair, and applying the shrunken set of sample values to a selected linear or affine linear transformation may be done entirely along the width dimension if w1 exceeds the width of one width / height pair, or along the height dimension if h1 exceeds the height of one width / height pair, where a given block is a first predetermined width / height pair, or where a given block is a second predetermined width / height pair, using only the first subpart of the selected linear or affine linear transformation relating to the subsampling of the w1xh1 array of sample positions.
[0204] Each linear or affine linear transformation in the first set of linear or affine linear transformations may be for transforming N1 sample values into w1*h1 predicted values for a w1xh1 array of sample positions where w1=h1, and each linear or affine linear transformation in the first set of linear or affine linear transformations may be for transforming N2 sample values into w2*h2 predicted values for a w2xh2 array of sample positions where w2=h2.
[0205] All embodiments described above are merely illustrative in that they can form the basis for the embodiments described below herein. That is, the above concepts and details are helpful in understanding the embodiments described below and serve as a repository for possible extensions and modifications of the embodiments described below herein. Many of the details described above, in particular, such as the averaging of adjacent samples and the fact that adjacent samples are used as reference samples, are optional.
[0206] More generally, the embodiments described herein assume that the prediction signal on a rectangular block is generated from already reconstructed samples, such that the intra-prediction signal on the rectangular block is generated from adjacent already reconstructed samples to the left and above the block. The generation of the prediction signal is based on the following steps: 1. However, among the reference samples called boundary samples, samples can be extracted by averaging, except for the possibility of transferring explanations to reference samples located elsewhere. Here, averaging is performed on both the left and top boundary samples of the block, or on only one of the boundary samples on either side. If averaging is not performed on one side, the sample on that side remains unchanged. 2. Optionally, an offset is added, and matrix-vector multiplication is performed. The input vector for the matrix-vector multiplication is either the concatenation of the averaged boundary sample on the left side of the block if averaging is applied only to the left, or the concatenation of the original boundary sample on the left of the block and the averaged boundary sample on top of the block if averaging is applied only to the top of the block, or the concatenation of the averaged boundary sample on the left of the top of the block and the averaged boundary sample on top of the block if averaging is applied to both sides of the block. Again, alternatives exist, such as not using averaging at all. 3. The result of matrix-vector multiplication and optional offset addition may optionally be a reduced prediction signal of a subsampled set of samples within the original block. The prediction signal at the remaining positions may be generated from the prediction signal on the subsampled set by linear interpolation.
[0207] The matrix-vector product calculation in step 2 should preferably be performed using integer arithmetic. Therefore,
number
number
number
number
number
number
number
number
number
number
number
number
number
[0208] Figure 8 shows the improved ALWIP prediction. Samples for a given block can be predicted based on a first matrix-vector product between a matrix A1100 derived by some machine learning-based training algorithm and a sample value vector 400. Optionally, an offset b1110 can be added. To achieve an integer or fixed-point approximation of this first matrix-vector product, the sample value vector can undergo a reversible linear transformation 403 to determine a further vector 402. A second matrix-vector product between the further matrix B1200 and the further vector 402 may be equal to the result of the first matrix-vector product.
[0209] Due to the characteristics of the further vector 402, the second matrix-vector product may be an integer approximated by the matrix-vector product 404 between a given prediction matrix C405, the further vector 402, and the further offset 408. The further vector 402 and the further offset 408 can consist of integer values or fixed-point values. For example, all components of the further offset are the same. The given prediction matrix 405 may be a quantized matrix or a matrix to be quantized. The result of the matrix-vector product 404 between the given prediction matrix 405 and the further vector 402 can be understood as the prediction vector 406.
[0210] The following provides further details regarding this integer approximation.
[0211] Possible solutions according to Example I: Subtraction and addition of average values Expressions usable in the above scenario
number
number
number
number
[0212]
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
[0213]
number
number
number
number
[0214] Alternatively, the predetermined value 1400 is the default value, or the value transmitted in the data stream in which the picture is encoded.
[0215] For example, a predetermined value of 1400 is advantageous when the deviation from the predicted value of a given block of samples is small.
[0216] According to one embodiment, the device is configured to include a plurality of reversible linear transformations 403, each associated with one component of a further vector 402. Furthermore, the device is configured to, for example, select a predetermined component 1500 from the components of a sample value vector 400 and use a reversible linear transformation 403 from among a plurality of reversible linear transformations associated with the predetermined component 1500 as the predetermined reversible linear transformation. This depends, for example, on the i0th row, i.e., the position of the predetermined component in the further vector, which is the position of the row of the reversible linear transformation 403 corresponding to the predetermined component. For example, if the first component of the further vector 402, i.e., y1, is the predetermined component, then i o The nth row replaces the first row of the invertible linear transformation.
[0217] As shown in Figure 9b, the column 412 of a predetermined prediction matrix 405 corresponding to a predetermined component 1500 of the further vector 402, i.e., the matrix component 414 of the predetermined prediction matrix C405 in the i0th column, is, for example, all zero. In this case, the device is configured to calculate the matrix-vector product 404 by performing a multiplication by calculating the matrix-vector product 407 between the reduced prediction matrix C'405 resulting from the predetermined prediction matrix C405 by leaving column 412, and the further vector 410 resulting from the further vector 402 by leaving a predetermined component 1500, as shown in Figure 9c. Thus, the prediction vector 406 can be calculated with fewer multiplications.
[0218] As shown in Figures 8, 9b, and 9c, the device may be configured to calculate the sum of each component of the prediction vector 406 with a, i.e., a predetermined value of 1400, when predicting a sample of a given block based on the prediction vector 406. This sum can be represented by the sum of the prediction vector 406 and vector 409, where all components of vector 409 are equal to the predetermined value of 1400, as shown in Figures 8 and 9c. Alternatively, the sum can be represented by the sum of the prediction vector 406 and the matrix-vector product 1310 between the integer matrix M1300 and a further vector 402, as shown in Figure 9b, where the matrix components of the integer matrix 1300 are 1 in the column of the integer matrix 1300 corresponding to a predetermined component 1500 of the further vector 402, i.e., the i0th column, and all other components are, for example, zero.
[0219] The sum of the given prediction matrix 405 and the integer matrix 1300 is equal to or approximates a further matrix 1200, for example, shown in Figure 8.
[0220] In other words, the matrix obtained by summing the matrix obtained by multiplying each matrix component of the predetermined prediction matrix C405 in column 412, i.e., the i0th column, which corresponds to a predetermined component 1500 of the further vector 402, by the inverse linear transformation 403 (i.e., matrix B), i.e., the further matrix B1200, corresponds, for example, to a quantized version of the machine learning prediction matrix A1100, as shown in Figures 8, 9a, and 9b. As shown in Figure 9b, summing each matrix component of the predetermined prediction matrix C405 in column 412 with 1 can correspond to the sum of the predetermined prediction matrix 405 and the integer matrix 1300. As shown in Figure 8, the machine learning prediction matrix A1100 can be equal to the result of multiplying the further matrix 1200 by the inverse linear transformation 403.
number
[0221] Matrix multiplication using only integer arithmetic. For low-complexity implementations (in terms of the complexity of scalar addition and multiplication, as well as the storage required for entries in the participating matrix), it is preferable to perform matrix multiplication 404 using only integer arithmetic.
[0222]
number
number
number
number
[0223] A matrix-vector product 404 with a size m × n matrix, i.e., a given prediction matrix 405, can be performed as shown in this pseudocode, where << and >> are left and right shift operations of arithmetic binary, and +, - and * operate only on integer values. (1) final_offset=1<<(right_shift_result-1); for i in 0…m-1 { accumulator=0 for j in 0…n-1 { accumulator:=accumulator + y[j]*C[i,j] } z[i]=(accumulator+final_offset)>>right_shift_result; }
[0224] Here, array C, i.e., the predetermined prediction matrix 405, stores fixed-point numbers, for example, as integers. The final addition of final_offset and the right-shift operation output of right_shift_result reduce precision by rounding in order to obtain the fixed-point format required for the output.
[0225] To expand the range of real numbers that can be represented by integers in C, two additional features are provided, as shown in the embodiments of Figures 10 and 11.
number
number
number
number
number
number
[0226] In other words, the device predicts parameters, for example
number
[0227] According to one embodiment, the prediction parameters include weights associated with the corresponding matrix components of the prediction matrix. In other words, a given prediction matrix is replaced or represented, for example, by the prediction parameters. The weights are, for example, integers and / or fixed-point values.
[0228] According to one embodiment, the prediction parameter is one or more scaling factors, for example, the value scale. i,j The further includes a weight associated with one or more corresponding matrix components of a given prediction matrix 405, for example
number
number
[0229]
number
number
number
[0230]
number
[0231] These solutions and their broader embodiments The above solution refers to the following embodiments. 1. A prediction method as described in Section I, in which step 2 of Section I performs the following for integer approximations of the matrix-vector product involved:
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
[0232] In other words, according to embodiments of the present application, the encoder and decoder operate as follows to predict a given block 18 of picture 10, with reference to Figure 8 in combination with one of Figures 7.1 to 7.4: Multiple reference samples are used for prediction. As outlined above, embodiments of the present application are not limited to intra-coding, and therefore the reference samples are not limited to adjacent samples, i.e., samples of picture 10 adjacent to block 18. In particular, the reference samples are not limited to those arranged along the outer edge of block 18, such as samples abutting the outer edge of the block. However, this situation is certainly one embodiment of the present application.
[0233] To perform the prediction, a sample value vector 400 is formed from reference samples such as reference samples 17a and 17c. Possible formations are described above. The formation may involve averaging, thereby reducing the number of samples 102 or the number of components in vector 400 compared to the reference sample 17 contributing to the formation. The formation may also depend in some way on dimensions or sizes such as the width and height of block 18, as described above.
[0234] To obtain the prediction for block 18, the vector 400 should be subjected to an affine or linear transformation. Different nomenclature is used above. The goal is to perform the prediction by applying vector 400 to matrix A by matrix-vector product within the range of performing a sum with offset vector b, using the most recent method. The offset vector b is of arbitrary choice. The affine or linear transformation determined by A or A and b can be determined for the prediction by the encoder and decoder, or more precisely, based on the size and dimensions of block 18 as already described above.
[0235] However, in order to achieve the computational efficiency improvements outlined above, or to make predictions more effective from an implementation standpoint, the affine or linear transformations are quantized, and the encoder and decoder, or their predictors, use C and T as described above to represent and perform the linear or affine transformations using C and T applied in the manner described above, which represent the quantized version of the affine transformation. In particular, instead of directly applying vector 400 to matrix A, the predictors of the encoder and decoder apply it by subjecting vector 402, obtained from sample value vector 400, to a mapping via a given reversible linear transformation T. The transformation T used here is the same as long as vector 400 is the same, i.e., does not depend on the dimensions of the block, i.e., width and height, or at least the same for different affine / linear transformations. In the above, vector 402 is denoted by y. The exact matrix for performing the affine / linear transformation determined by machine learning was B. However, instead of performing B exactly, predictions in the encoder and decoder are made by approximating it or by a quantized version thereof. In particular, the representation is done by appropriately representing C in the manner outlined above, where C+M represents a quantized version of B.
[0236] Therefore, prediction in the encoder and decoder is further performed by computing a matrix-vector product 404 between vector 402 and a predetermined prediction matrix C appropriately represented and stored in the encoder and decoder in the manner described above. Next, the vector 406 obtained from this matrix-vector product is used to predict sample 104 of block 18. As described above, for prediction, each component of vector 406 may be subject to summing with parameter a, as shown in 408, to compensate for the corresponding definition of C. An optional sum of vector 406 with offset vector b may also be involved in deriving the prediction of block 18 based on vector 406. As described above, each component of vector 406, and therefore each component of the sum of vector 406, all vectors a shown in 408, and an optional vector b, directly correspond to sample 104 of block 18 and can therefore represent the predicted value of the sample. It is also possible that only a subset of sample 104 of the block is predicted in this way, and the remaining samples of block 18, such as 108, are derived by interpolation.
[0237] As described above, there are different embodiments for setting α. For example, it may be the arithmetic mean of the components of vector 400. For that case, see, for example, Figure 9. The inverse linear transformation T may be as shown in Figure 9a. i0 represents the given components of the sample value vector and vector 402, respectively, and is replaced by a. However, as described above, there are other possibilities. However, as far as the representation of C is concerned, as shown above, it can be embodied in a different way. For example, the matrix-vector product 404 may, in its actual calculation, end up being the actual calculation of a smaller matrix-vector product with lower dimensions, for example, see Figure 9c. In particular, as described above, by the definition of C, the entire i0th column 412 may be 0, so the actual calculation of product 404 is
number
[0238] The weights of C or C', i.e., the components of this matrix, can be represented and stored in fixed-point representation. However, these weights 414 can also be stored in a manner related to different scales and / or offsets, as described above. The scales and offsets may be defined for the entire matrix C, i.e., equal for all weights 414 of matrix C or matrix C', or they may be defined to be constant or equal for all weights 414 of the same row or column of matrices C and C', respectively. In this regard, Figure 10 shows that the calculation of the matrix-vector product, i.e., the result of the product, can actually be performed in a slightly different way, i.e., by shifting the multiplication with the scale(s) towards the vector 402 or 404, for example, thereby reducing the number of further multiplications that must be performed. Figure 11 shows the case where one scale and one offset are used for all weights 414 of C or C', as done in calculation (2) above.
[0239] 8. Embodiments using block-based intra-predictive mode for all channels in special cases All of the above descriptions are to be considered details of any optional implementation of the embodiments described herein. Note that, hereafter, the term matrix-based intra-prediction (MIP) is used to refer to intra-prediction modes such as block-based intra-prediction modes, which may be embodied by or equivalent to those shown by the ALWIP described above with respect to Figures 5 to 11.
[0240] This paper proposes that, in the case of a 4:4:4 chroma format and a single tree, on a chroma intra block where the chroma intra mode is direct mode (DM) and the chroma intra mode is MIP mode, the chroma intra prediction signal is generated using this MIP mode.
[0241] In direct mode, the chromine intra-prediction mode is derived from, for example, the lumine intra-prediction mode. For example, the chromine intra-prediction mode is equal to the lumine intra-prediction mode. There is an exception when the lumine intra-prediction mode is the MIP mode, and the chromine intra-prediction mode is the MIP mode only when it is a 4:4:4 color sampling format; in the case of a single tree, and in all other cases, the chromine intra-prediction mode is the planar mode. The 4:4:4 color sampling format represents a color sampling format in which each color component is sampled equally.
[0242] 8.1 Description of the proposed method In current VTMs, MIP is used only for the luma component [3]. If the intra-mode of a chroma intra-block is direct mode (DM) and the intra-mode of a luma block in the same location is MIP mode, the chroma block must use planar mode to generate the intra-prediction signal. The main reason for treating the DM mode in this way in the case of MIP is that, in the case of 4:2:0 or dual trees, luma blocks in the same location may have a different shape than the color blocks. Therefore, since the MIP mode is not applicable to all block shapes, the MIP mode of a luma block in the same location may not be applicable to the chroma block in that case.
[0243] On the other hand, if the chroma format is 4:4:4 and a single tree is used, it is not true that the above luma MIP mode cannot be applied to the chroma component. Therefore, in this case, if the chroma intra mode on the intra block is DM mode and the luma intra mode is MIP mode, it is proposed that the chroma intra prediction signal is generated in this MIP mode.
[0244] This proposed change is particularly relevant when ACT (Adaptive Color Conversion) is enabled, as the chroma mode is presumed to be DM mode when ACT is used in a particular block. Specifically for ACT, a strong correlation of predictive signals across all channels is beneficial, and such correlation can be enhanced when the same intra-predictive mode is used for all three components of a block. The experimental results reported below support this view, as the proposed change significantly impacts camera-captured content in RGB format, which is claimed to be the intersection of MIP and ACT when they are primarily designed.
[0245] 8.2 Embodiments of the Proposed Method Figure 12 shows one embodiment of the prediction of a predetermined second color component block 182 of picture 10. A block-based decoder, block-based encoder, and / or apparatus for predicting picture 10 can be configured to perform this prediction.
[0246] According to one embodiment, the picture 10 is divided equally with respect to each color component 101, 102 using a division scheme 11', for example, a first color component block 18' 11 ~18' 1n and the second color component block 18' 21 ~18' 2nA picture 10 of multiple color components 101, 102 and a color sampling format 11, in which each color component 101, 102 is sampled equally into blocks. According to the color sampling format 11, each color component 101, 102 is sampled equally, for example, and according to the division scheme 11', each color component 101, 102 is divided similarly, for example. Thus, according to the division scheme 11', for example, the first color component block 18' 11 ~18' 1n This is the second color component block 18' 21 ~18' 2n It has the same geometric characteristics, and according to the color sampling format 11, the first color component block 18' 11 ~18' 1n This is the second color component block 18' 21 ~18' 2n It has the same sampling. The sampling is indicated by sample points in a predetermined second color component block 182 and an intra-predicted first color component block 181 located at the same location.
[0247] The first color component 101 of picture 10 is the intra-predicted first color component block 18' of picture 10. 11 ~18' 1n For each of these, the block interior 18 derives the sample value vector 514 of the reference sample 17, is adjacent to the block interior 18, and in order to obtain the prediction vector 514, the sample value vector 514 and the respective matrix-based intra prediction modes 5101~510 m By calculating the matrix-vector product 512 between the associated prediction matrix 516 and the prediction vector 514, and by predicting the samples of the block interior 18 based on the prediction vector 514, matrix-based intra-prediction modes (MIP modes) 5101~510 are determined for each mode in which the block interior 18 is predicted. mOne is selected from a first set of 508 intra-prediction modes, including the above, and decoded / encoded in blocks. The sample value vector 514 has the features and / or functions described with respect to the sample value vector 400, for example, with respect to the further vector 402 in Figures 6 or 8-11. In other words, matrix-based intra-prediction 5101-510 m This is performed, for example, as explained with respect to one of Figures 5 to 11.
[0248] Optionally, the number of components in the prediction vector 518 is less than the number of samples in the block interior 18. In this case, the samples in the block interior 18 can be predicted based on the prediction vector 518 by interpolating the samples based on the components of the prediction vector 518 assigned to the support sample positions in the block interior 18.
[0249] Optionally, a matrix-based intra-prediction mode 5101-510 consisting of a first set of 508 intra-prediction modes. m However, a set of coprime subsets of matrix-based intra-prediction modes is selected according to the block dimension of each intra-predicted first color component block 181, such that one subset of matrix-based intra-prediction modes is obtained from a set of coprime subsets of matrix-based intra-prediction modes. The prediction matrices 516 associated with the set of coprime subsets of matrix-based intra-prediction modes are, for example, machine-trained, and prediction matrices 516 consisting of one subset of matrix-based intra-prediction modes are equal in size, while prediction matrices 516 consisting of two subsets of matrix-based intra-prediction modes selected for different block sizes are different in size. In other words, prediction matrices consisting of one subset are equal in size, but prediction matrices of different subsets have different sizes.
[0250] Optionally, a matrix-based intra-prediction mode 5101-510 consisting of a first set of 508 intra-prediction modes. m The prediction matrices 516 associated with this are of equal size and are machine-learned.
[0251] Optionally, matrix-based intra-prediction modes 5101-510 consisting of a first set of intra-prediction modes 508 m The intra-prediction modes, consisting of a first set 508 of intra-prediction modes other than the one mentioned above, include DC mode 506, planar mode 504, and directional modes 5001-5001.
[0252] The second color component 102 of picture 10 is decoded and / or encoded in block units by intra-predicting a predetermined second color component block 182 of picture 10 using a matrix-based intra-prediction mode selected for the intra-predicted first color component block 181 located at the same location. Thus, the decoder and / or encoder uses MIP modes 5101-510 for at least one intra-encoded second / third color component block. m It is necessary to be able to adopt this. This is done, for example, when the initial prediction mode of a given second color component block 182 is direct mode.
[0253] In direct mode, the intra-prediction mode of a predetermined second color component block 182 is derived, for example, from the intra-prediction mode of an intra-predicted first color component block 181 located at the same location. For example, the intra-prediction mode of a predetermined second color component block 182 is equal to the intra-prediction mode of an intra-predicted first color component block 181 located at the same location. In other words, the intra-prediction mode of an intra-predicted first color component block 181 located at the same location is equal to the intra-prediction mode of MIP modes 5101~510 m If so, the predetermined second color component block 182 is also the same MIP mode 5101~510 m This is predicted by [the program / platform].
[0254] According to one embodiment, if an index indicating the prediction mode is not notified in a predetermined second color component block 182 in the data stream, the direct mode can be inferred. Otherwise, the mode notified in the data stream is used to predict the predetermined second color component block 182.
[0255] In other words, the second color component block 18' of picture 10. 21 ~18' 2n For each of these, one can be selected from the first option and the second option. According to the first optional selection, the intra-predicted first color component block 18' located at the same location 11 -18' 1n The selected intra prediction mode is matrix-based intra prediction mode 5101 to 510. m In that case, each second color component block 18' 21 -18' 2n The intra-prediction mode for the intra-predicted first color component block 18' is located at the same location. 11 ~18' 1n Each second color component block 18' is equal to 21 ~18' 2n The intra-prediction mode for the intra-predicted first color component block 18' is located at the same location. 11 ~18' 1n Derived based on the selected intra-prediction mode for the first color component block 18' located at the same location. 11 ~18' 1n The selected intra-prediction mode is planar intra-prediction mode 504, DC intra-prediction mode 506 and / or directional intra-prediction mode 5001~500. l As shown, matrix-based intra prediction modes 5101~510 m If it is one of the other intra prediction modes, then each second color component block 18' 21 ~18' 2n The intra-prediction mode also uses the intra-predicted first color component block 18' located in the same place.11 ~18' 1n The selected intra prediction mode may be equal to the first option. The first option can represent the direct mode. According to the second option, each second color component block 18' 21 ~18' 2n The intra-prediction mode for each second color component block 18' 21 ~18' 2n The intra-mode index is selected based on the intra-mode index present in the data stream. Optionally, the intra-mode index is selected from a first set of 508 intra-predicted first color component blocks 18' located in the same place, for the intra-predicted first color component block 18'. 11 ~18' 1n In addition to further intra-mode indices present in the data stream for each second color component block 18' present in the data stream. 21 ~18' 2n The intra-mode indices present in the data stream for each second color component block 18' 21 ~18' 2n For example, the prediction modes from the first set of intra-prediction modes (508) are shown to be selected for the prediction.
[0256] According to one embodiment, when the picture 10 has a 4:4:4 color sampling format 11 and the prediction mode of a predetermined second color component block 182 is direct mode, the device for predicting a predetermined block of the picture, a block-based decoder and / or block-based encoder, predicts the same MIP mode 5101-510 as the intra-predicted first color component block 181 located at the same location for the predetermined second color component block 182. mConfigured to use, the intra-predicted first color component block 181 and the predetermined second color component block 182 located in the same place represent predetermined blocks associated with different color components, namely the first and second color components. The intra-predicted first color component block 181 and the predetermined second color component block 182 located in the same place include, for example, the same geometric features and the same spatial positioning in picture 10. In other words, the first color component coding tree, e.g., the luma coding tree, is equivalent to the second color component coding tree, e.g., the chroma coding tree. For example, a single tree is used. Picture 10 is divided equally with respect to each color component 101, 102. The division scheme 11' defines, for example, single-tree processing of picture 10.
[0257] According to one embodiment, dual-tree processing of picture 10 is also possible. Dual-tree processing can be associated with a further division scheme 11' in which picture 10 is divided with respect to a first color component 101 using first division information in a data stream, and picture 10 is divided with respect to a second color component 2 using second division information that exists in a data stream separate from the first division information.
[0258] If it is not a dual tree or single tree, the device predicting a given block of the picture, a block-based decoder and / or block-based encoder, for example, the prediction mode of the intra-predicted first color component block 181 located at the same location is MIP mode 5101~510 mTherefore, if the prediction mode of a predetermined second color component block 182 is direct mode, it is configured to use planar intra-prediction mode 504. However, if the prediction mode of the intra-predicted first color component block 181 at the same location is not MIP mode, for example, if the prediction mode of the intra-predicted first color component block 181 at the same location is DC intra-prediction mode 506, planar intra-prediction mode 504, or directional intra-prediction mode 5001-5001 (e.g., angular intra-prediction mode), and the prediction mode of a predetermined second color component block 182 is direct mode, then the prediction mode of the predetermined second color component block 182 is equal to the prediction mode of the intra-predicted first color component block 181 at the same location. This can also be applied when using a single tree but not a 4:4:4 color sampling format 11, for example, when using a 4:1:1 color sampling format 11, a 4:2:2 color sampling format 11, or a 4:2:0 color sampling format 11.
[0259] The color sampling format 11 can be represented as x:y:z, where the first number x refers to, for example, the size of the color component block 18', and the next two numbers y and z both refer to the second and / or third color component samples. These, i.e., y and z, are both relative to the first number and define horizontal and vertical sampling, respectively. A signal with 4:4:4 is not compressed (and therefore not subsampled) and carries the first color component and further color component data, e.g., the second and / or third color component samples, i.e., the entire chroma sample. In a 4x2 array of pixels, 4:2:2 has half the chroma sample of 4:4:4, and 4:2:0 and 4:1:1 have one-quarter of the available saturation information. A 4:2:2 signal has half the horizontal sampling rate but maintains full sampling in the vertical direction. On the other hand, 4:2:0 samples only the color from half of the pixels in the first row, completely ignoring the second row of the sample, while 4:1:1 samples the color of one pixel in the first row and the color of one pixel in the second row of the sample. In other words, a 4:1:1 signal has a 1 / 4 horizontal sampling rate, but maintains full sampling in the vertical direction.
[0260] According to one embodiment, if the prediction mode of an intra-predicted first color component block 181 located in the same place is not one of the MIP modes 5101-5101, such as DC intra-prediction mode 506, planar intra-prediction mode 504, or directional intra-prediction modes 5001-5001, and the prediction mode of a predetermined second color component block 182 is a direct mode, then the device for predicting a predetermined block of the picture, a block-based decoder and / or block-based encoder, is configured to use the prediction mode of an intra-predicted first color component block 181 located in the same place as the prediction mode of a predetermined second color component block, independently of the color sampling format 11 and / or the division scheme of the picture 10 with respect to each color component, i.e., using a single tree or a dual tree.
[0261] In DC intra-prediction mode 506, for example, one value, which is a quasi-DC value, is derived based on adjacent samples 17 that are spatially adjacent to a predetermined block 18, for example, an intra-predicted first color component block 181 and / or a predetermined second color component block 181 located in the same place, and this one DC value is attributed to all samples in the predetermined block 18 in order to obtain an intra-prediction signal.
[0262] In the planar intra-prediction mode 504, for example, a two-dimensional linear function defined by horizontal gradient, vertical gradient, and offset is derived based on adjacent samples 17 that are spatially adjacent to a given block 18, for example, an intra-predicted first color component block 181 and / or a given second color component block 182 located in the same place, and this linear function defines the predicted sample values of the given block 18.
[0263] For example, in directional intra-prediction modes 5001-5001, such as angular intra-prediction mode, a reference sample 17 adjacent to a predetermined block 18, for example, an intra-predicted first color component block 181 and / or a predetermined second color component block 182 located in the same place, is used to fill the predetermined block 18 in order to obtain the intra-prediction signal of the predetermined block 18. In particular, a reference sample 17 positioned along the boundary of the predetermined block 18, such as along the top and left edge of the predetermined block 18, represents picture content extrapolated or copied into the predetermined block 18 along a predetermined direction 502. Prior to extrapolation or copying, the picture content represented by the adjacent sample 17 may be subject to interpolation filtering, i.e., it can be derived from the adjacent sample 17 by interpolation filtering. Angular intra-prediction modes 5001-500 l The intra-prediction directions 502 are different from each other. Each angle intra-prediction mode 5001~500 l It can have an associated index, and the angle intra prediction mode 5001~500 lThe association of the index to the angle intra prediction mode 5001-500 is according to the associated mode index. l When ordering them, direction 502 may rotate monotonically clockwise or counterclockwise.
[0264] Figure 13 shows in more detail the prediction of a given second color component block 182 of picture 10, depending on different division schemes and color sampling formats.
[0265] The predictions are shown for picture 10a, another picture 10b, and yet another picture 10c, with different conditions set for different pictures.
[0266] According to one embodiment, for example, a device such as a block-based decoder, a block-based encoder, and / or a device for predicting a picture includes or is accessible a set of partitioning schemes 11' to select and / or obtain a partitioning scheme for picture 10. The set of partitioning schemes 11' includes partitioning scheme 11'1, i.e., a first partitioning scheme in which picture 10 is divided equally with respect to each color component 101, 102, and picture 10 is divided with respect to a first color component 101 using first partitioning information 11'2 in the data stream 12, and picture 10 is divided with respect to the first partitioning information 11' 2a A second segmentation information 11' exists separately within the data stream 12. 2b The first division scheme 11'1 can represent a single-tree processing of the picture 10, and the second division scheme 11'2 can represent a dual-tree processing of the picture.
[0267] In a further subdivision scheme 11'2, the first color component 101 is subdivided in a different way than the second color component 102. One of the color components 101 and 102 can be subdivided finely, while the other color components 101 and 102 can be subdivided coarsely. Second color component block 18' 21 ~18'24 The boundary is the first color component block 18' 11 ~18' 116 It is also possible that it is not located in the same place as the boundary. Second color component block 18' 21 ~18' 24 The boundary is 10 2alternative Or, as shown in the further division of picture 10b, the first color component block 18' 11 ~18' 116 It is possible to pass through the inside of the block.
[0268] For picture 10a, the same prediction as described with respect to Figure 12 can be applied, and a color sampling format is used in which the first division scheme 11'1 is selected from the set of division schemes 11', and each color component 10a1, 10a2 is sampled equally.
[0269] The apparatus in Figure 13 is configured to select one of the first and second options for each of the intra-predicted second color component blocks of picture 10a, for example. In the first option, as shown in Figure 12, the selected intra-prediction mode for the intra-predicted first color component block 18a1 at the same location is the matrix-based intra-prediction mode 5101-510 mIf one of these is the case, the intra-prediction mode for each intra-predicted second color component block 18a2 is derived based on the intra-prediction mode selected for the intra-predicted first color component block 18a1 located in the same place, such that the intra-prediction mode for each intra-predicted second color component block 18a2 is equal to the intra-prediction mode selected for the intra-predicted first color component block 18a1 located in the same place. In the second option, the intra-prediction mode for each intra-predicted second color component block 18a2 is selected based on an intra-mode index 509 present in the data stream for each intra-predicted second color component block 18a2. The intra-mode index 509 is present in the data stream in addition to a further intra-mode index 507 present in the data stream for the intra-predicted first color component block 18a1 located in the same place, for selection from a first set of intra-prediction modes.
[0270] Optionally, the device is configured to divide a further picture 10b into further blocks using a further division scheme 11'2, with each color component sampled equally, and with a color sampling format. Thus, the first color component 10b1 has the same color sampling as the second color component 10b2, but with a different block size. As shown in Figure 13, a further block 18b2 of a given second color component in the further picture 10b has, for example, one-quarter the size of a further block 18b1 of the first color component located in the same place in the further picture 10b. The further picture 10b is sampled, for example, in a 4:4:4 color sampling format.
[0271] When using a further division scheme 11'2, a further block 18b1 of the first color component located in the same place in a further block 18b2 of a given second color component is determined, for example, based on a pixel located in both blocks 18b1 and 18b2. According to one embodiment, this pixel is located in the upper left corner, upper right corner, lower left corner, lower right corner and / or in the middle of the further block 18b2 of the given second color component.
[0272] The first color component 10b1 of the further picture 10 is selected from the first set of intra-prediction modes for each of the further blocks of the intra-predicted first color component of the further picture 10b, and decoded in units of further blocks.
[0273] For each of the further blocks of the further picture 10b of the intra-predicted second color component, one of the first option and the second option can be selected. According to the first option, the intra-prediction mode for each further block of the intra-predicted second color component, for example, a given further block 18b2 of the second color component, is the intra-prediction mode for each further block 18b2 of the intra-predicted second color component, and the intra-prediction mode selected for the further block 18b1 of the first color component located in the same place is the matrix-based intra-prediction mode 5101~510. m If it is one of the above, it is derived based on the intra-prediction mode selected for a further block 18b1 of the first color component in the same location, so as to be equal to the planar intra-prediction mode 504. The intra-prediction mode selected for a further block 18b1 of the first color component in the same location is, for example, the intra-prediction mode is planar intra-prediction mode 504, DC intra-prediction mode 506, or directional intra-prediction modes 5001-5001, and matrix-based intra-prediction modes 5101-510 mIf not one of the above, the intra-prediction mode for each intra-predicted second color component further block 18b2 is equal to the intra-prediction mode selected for the first color component further block 18b1 located in the same place. According to the second option, the intra-prediction mode for each intra-predicted second color component block 18b2 is selected based on the intra-mode index 5092 present in the data stream 12 for each intra-predicted second color component block 18b2. The intra-mode index 5092 present in the data stream 12 is, for example, a further intra-mode index 5072 present in the data stream for the intra-predicted first color component block 18b1 located in the same place, for selection from a first set of intra-prediction modes.
[0274] Optionally, the device is configured to split further pictures 10c into multiple color components and different color sampling formats, such that multiple color components are sampled differently. As shown in Figure 13, the first color component 10c1 is sampled differently from the second color component 10c2. Sampling is indicated by sample points. According to the embodiment shown in Figure 13, the further pictures are sampled according to a 4:2:1 color sampling format. However, it is clear that other color sampling formats other than the 4:4:4 color sampling format can also be used. This is shown once for the use of the first splitting scheme 11'1 and once for the use of a further splitting scheme 11'2. Regardless of the splitting scheme 11', further predictions of the further pictures 10c can be made.
[0275] Furthermore, the first color component 10c1 of the further picture 10c is selected from the first set of intra-prediction modes for each of the further blocks of the intra-predicted first color component of the further picture 10c, and decoded in units of the further blocks.
[0276] Furthermore, the device is configured to select, for example, one of the first and second options for an additional block of picture 10c for each of the intra-predicted second color components. According to the first option, the intra-prediction mode for each additional block 18c2 of the intra-predicted second color component is the intra-prediction mode for each additional block 18c2 of the intra-predicted second color component, and the intra-prediction mode selected for the additional block 18c1 of the first color component in the same location is the matrix-based intra-prediction mode 5101~510. m If it is one of the above, it is derived based on the intra-prediction mode selected for any further block 18c1 of the first color component in the same location, so as to be equal to the planar intra-prediction mode 504. The intra-prediction mode selected for any further block 18c1 of the first color component in the same location is, for example, the intra-prediction mode is the planar intra-prediction mode 504, the DC intra-prediction mode 506, or the directional intra-prediction modes 5001-500. l This is the matrix-based intra prediction mode 5101~510 m If not one of the above, the intra-prediction mode for each intra-predicted second color component for any further block 18c2 is equal to the intra-prediction mode selected for any further block 18c1 of the first color component located in the same place. According to the second option, the intra-prediction mode for each intra-predicted second color component for any further block 18c2 is selected based on the intra-mode index 5093 present in the data stream for each intra-predicted second color component for any further block 18c2. Optionally, the intra-mode index 5093 is present in the data stream 12 in addition to the further intra-mode index 5073 present in the data stream for any further block 18c1 of the intra-predicted first color component located in the same place for selection from the first set of intra-prediction modes.
[0277] According to one embodiment, if blocks 18a2, 18b2 and / or 18c2 in the data stream 12 are notified to deactivate a residual coding color conversion mode, such as by deactivating ACT (Adaptive Color Conversion), the selection from the first and second options depends on signal transmission present in the data stream 12 to each intra-predicted second color component block 18a2, each intra-predicted second color component further block 18b2, and / or each intra-predicted second color component yet further block 18c2. If blocks 18a2, 18b2, and / or 18c2 in the data stream 12 are notified to activate a residual coding color conversion mode, such as by activating ACT (Adaptive Color Conversion), no signal transmission is necessary. In this case, i.e., the residual coding color conversion mode is notified to be activated for blocks 18a2, 18b2 and / or 18c2, and the device may be configured to infer, for example, that when ACT (adaptive color conversion) is activated, the first option should be selected.
[0278] When ACT is used on a given block, for example, the chroma mode is presumed to be DM mode, i.e., the first option. In particular for ACT, a strong correlation of the predicted signals across all channels is beneficial, and such correlation can be enhanced if the same intra-prediction mode is used for all three components of the block. The experimental results reported below support this view in that the proposed changes have a significant impact on camera-captured content in RGB format, which is claimed to be the intersection when MIP and ACT are primarily designed.
[0279] 8.3 Experimental Results This section reports experimental results for 4:4:4 according to general test conditions. Tables 1 and 2 report the results of the proposed modifications compared to VTM-7.0 anchors in AI and RA configurations, respectively. The corresponding simulations were run on an Intel Xeon cluster (E5-2697A v4, AVX2 on, Turbo Boost off) with Linux® OS and GCC 7.2.1 compiler. Results are reported for RGB and YUV according to CE-8 conditions. Regarding single-tree off, the proposed modifications do not alter the bitstream, so single-tree is always enabled.
[0280] [Table 3] [Table 4] [Table 5] [Table 6]
[0281] 8.4 Conclusion This paper proposes enabling MIP on all three channels in the case of 4:4:4 content and single trees. In this case, it is proposed that the chroma component on the intrablock use MIP mode, and the chroma intramode is DM mode, and that the chroma intraprediction signal be generated in this MIP mode. It is proposed that the techniques described in this paper be adopted in the next working draft of VVC.
[0282] 9 References [1] P. Helle et al., “Non-linear weighted intra prediction”, JVET-L0199, Macao, China, October 2018. [2] F.Bossen, J.Boyce, K.Suehring, X.Li, V.Seregin, “JVET common test conditions and software reference configurations for SDR video”, JVET-K1010, Ljubljana, SI, July 2018. [3] B.Bross, J.Chen, S.Liu, Y.-K.Wang,Verstatile Video Coding (Draft 7),Document JVET-P2001,Version 14 ,Geneva,Switzerland,October 2019 […]Common test conditions for 4:4:4
[0283] Further embodiments and examples Generally, an example can be implemented as a computer program product having program instructions, which are operable to perform one of the methods when the computer program product is executed on a computer. The program instructions can be stored, for example, on a machine-readable medium.
[0284] Other examples include computer programs stored on machine-readable media for performing one of the methods described herein.
[0285] In other words, one example of the method is a computer program having program instructions for performing one of the methods described herein when the computer program is executed on a computer.
[0286] Further examples of the method include a data carrier medium (or digital storage medium, or computer-readable medium) and a computer program for performing one of the methods described herein, which is recorded thereon. The data carrier medium, digital storage medium, or recorded medium is tangible and / or non-temporary, rather than intangible and temporary signals.
[0287] A further example of the method is a data stream or signal sequence, which thus represents a computer program for performing one of the methods described herein. The data stream or signal sequence may be transmitted over a data communication connection, such as over the Internet.
[0288] Further examples include, for example, processing means such as a computer or programmable logic device that performs one of the methods described herein.
[0289] Further examples include computers on which computer programs for performing one of the methods described herein are installed.
[0290] Further examples include devices or systems that transfer (e.g., electronically or optically) a computer program for performing one of the methods described herein to a receiver. The receiver may be, for example, a computer, a mobile device, a memory device, etc. The device or system may include, for example, a file server for transferring the computer program to the receiver.
[0291] In some examples, a programmable logic device (e.g., a field-programmable gate array) may be used to perform some or all of the functions of the method herein. In some examples, a field-programmable gate array may work with a microprocessor to perform one of the methods herein. In general, the method may be performed by any suitable hardware device.
[0292] The examples described above are merely illustrative to illustrate the principles described above. It will be understood that modifications and changes to the arrangements and details described herein will become apparent. Therefore, it is intended that this specification be limited only by the imminent claims and not by the specific details presented as part of the description and explanation of the examples herein.
[0293] Equal or equivalent elements, or elements with equal or equivalent functionality, are represented in the following description by equal or equivalent reference numerals, even if they occur in different figures.
Claims
[Claim 1] A video decoder comprising one or more processors, wherein the one or more processors are To determine that the picture is in a 4:4:4 color sampling format, Determining the matrix-based intra-prediction (MIP) mode based at least partially on the index signaled to the data stream, Using the aforementioned MIP mode, decode the Rumablock of the picture, Regarding the chroma block of the aforementioned picture, it is determined whether the encoding tree of the chroma block is a single tree, This involves selecting an intra-prediction mode for decoding the chroma block, In response to the determination that the coding tree of the chroma block is a single tree, the selected intra-prediction mode is MIP mode, and, In response to the determination that the coding tree of the chroma block is not a single tree, the selected intra-prediction mode is a planar intra-prediction mode. Selecting the aforementioned intra-prediction mode, The chroma block is decoded using the selected intra prediction mode, A video decoder configured to perform the following actions.