Intra-prediction using linear or affine transforms with adjacent sample reduction

This method reduces the computational complexity and efficiency of video coding by employing a reduced set of sample values to improve the efficiency of video coding.

JP2026062940APending Publication Date: 2026-04-10FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
Filing Date
2026-01-06
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing video coding technologies face challenges in achieving efficient compression of video data, particularly in intra-prediction modes, due to high computational complexity and resource consumption in performing affine linear transformations.

Method used

The use of a reduced set of sample values obtained by downsampling or averaging neighboring samples, followed by linear or affine linear transformations, to predict block values in video coding, reducing the number of multiplications required for prediction, and, thereby optimizing the prediction process.

Benefits of technology

This approach reduces the number of multiplications required for prediction, enhancing the efficiency of video coding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026062940000001_ABST
    Figure 2026062940000001_ABST
Patent Text Reader

Abstract

The present invention provides a method for encoding / decoding a video signal, which is implemented in a decoder, encoder, method, and non-temporary memory unit that stores instructions for performing the method. [Solution] The method predicts a given block of a picture by using multiple adjacent samples, reducing the number of adjacent samples 811 to obtain sample values ​​102 of a reduced set with fewer samples compared to the multiple adjacent samples, and subjecting the sample values ​​of the reduced set to a linear or affine linear transformation 812 to obtain predicted values ​​of a given block 18.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] 1. Introduction Different embodiments, examples, and aspects of the present invention are described below. At least some of these examples, embodiments, and aspects relate, among other things, to methods and / or apparatus for video coding, and / or for performing intra-prediction using linear or affine transforms with, for example, adjacent sample reduction, and / or for optimizing video delivery (e.g., broadcast, streaming, file playback, etc.) for, for example, video applications and / or virtual reality applications. Furthermore, the examples, embodiments, and aspects relate to High Efficiency Video Coding (HEVC) or its successors. Further embodiments, examples, and aspects are also defined by the appended claims. [Background technology]

[0002] Please note that any embodiments, examples, and aspects defined by the claims may be supplemented by any of the details (features and functions) described in the following chapters.

[0003] Furthermore, the embodiments, examples, and aspects described in the following chapters may be used individually and may be supplemented by any of the features in other chapters or any features included in the claims.

[0004] Furthermore, it should be noted that the individual examples, embodiments, and aspects described herein can be used individually or in combination. Therefore, details can be added to each of the examples, embodiments, and individual aspects without adding further detail to another aspect. Note that this disclosure also explicitly or implicitly describes the features of the decoding and / or encoding systems and / or methods.

[0005] Furthermore, the features and functions disclosed herein in relation to the method may also be used in the apparatus. Additionally, any features and functions disclosed herein in relation to the apparatus may also be used in a corresponding manner. In other words, the methods disclosed herein can be complemented by any of the features and functions described in relation to the apparatus. Moreover, any of the features and functions described herein may be implemented in hardware, software, or a combination of hardware and software, as described in other sections such as "Further Embodiments and Examples."

[0006] Furthermore, any of the features indicated in parentheses ("(...)" or "[...]") may be considered optional in some examples, embodiments, or aspects. [Overview of the Initiative] [Problems that the invention aims to solve]

[0007] [Means for solving the problem]

[0008] 1.1 Overview According to one embodiment, a decoder for decoding a picture from a data stream, using multiple adjacent samples,

[0009] Compared to multiple neighboring samples, reducing the number of samples to obtain a reduced set of sample values ​​by reducing the number of neighboring samples (e.g., by averaging or downsampling), In order to obtain predicted values ​​for a given sample in a given block, a reduced set of sample values ​​is subjected to a linear or affine linear transformation. A decoder is provided, configured to predict a given block of a picture.

[0010] In the example, the decoder may be further configured to perform, for example, averaging downsampling of multiple neighboring samples to obtain a reduced set of sample values ​​with fewer samples compared to multiple neighboring samples.

[0011] In some cases, the decoder may also derive predicted values ​​for further samples in a given block based on the predicted values ​​of a given sample and several adjacent samples, for example, by interpolation. Therefore, upsampling operations can be applied. According to one embodiment, an encoder for encoding a picture from a data stream, using a plurality of adjacent samples,

[0012] Compared to multiple neighboring samples, reducing the number of samples to obtain a reduced set of sample values ​​by reducing the number of neighboring samples (e.g., by averaging or downsampling), In order to obtain predicted values ​​for a given sample in a given block, a reduced set of sample values ​​is subjected to a linear or affine linear transformation. An encoder is provided, configured to predict a given block of a picture.

[0013] In the example, the encoder may be further configured to perform a reduction by downsampling multiple neighboring samples to obtain a reduced set of sample values ​​with fewer samples compared to multiple neighboring samples.

[0014] In some cases, the encoder may also derive predicted values ​​for further samples in a given block based on the predicted values ​​of a given sample and several adjacent samples, for example, by interpolation. Therefore, upsampling operations can be applied.

[0015] In some examples, a system comprising an encoder and / or a decoder as described above may be provided. In some examples, the encoder hardware and / or at least some procedural routines may be the same as those of the decoder.

[0016] In the embodiment, multiple neighboring samples are used to predict a given block of a picture by reducing the number of neighboring samples, for example by downsampling or averaging, in order to obtain a sample value for a reduced set of fewer samples compared to multiple neighboring samples. In order to obtain predicted values ​​for a given sample of a given block, the sample values ​​of the reduced set are subjected to a linear or affine linear transformation. A decryption method including the above may be provided.

[0017] In the embodiment, multiple neighboring samples are used to predict a given block of a picture by reducing the number of neighboring samples, for example by downsampling or averaging, in order to obtain a sample value for a reduced set of fewer samples compared to multiple neighboring samples. In order to obtain predicted values ​​for a given sample of a given block, the sample values ​​of the reduced set are subjected to a linear or affine linear transformation. An encoding method including this may be provided. In the embodiment, a non-temporary memory unit may be provided that, when executed by the processor, stores instructions that cause the processor to perform the above method. 1.3 Drawings [Brief explanation of the drawing]

[0018] [Figure 1] An example of an encoder is shown. [Figure 2] An example of an encoder is shown. [Figure 3] An example of a decoder is shown. [Figure 4] An example of a decoder is shown. [Figure 5] This is a diagram showing the predicted blocks. [Figure 6] This demonstrates matrix operations. [Figure 7.1] An example of operation according to the embodiment is shown. [Figure 7.2] An example of operation according to the embodiment is shown. [Figure 7.3] An example of operation according to the embodiment is shown. [Figure 7.4] An example of operation according to the embodiment is shown. [Figure 8.1] An example of the method described in the examples is shown. [Figure 8.2] An example of the method described in the examples is shown. [Figure 9] An example of a component is shown below. [Figure 10] An example of an encoder is shown. [Figure 11] An example of a decoder is shown. [Figure 12] This shows a scheme for associating the dimension of the block to be predicted with the prediction mode. [Figure 13] A useful scheme for understanding the present invention is shown. [Modes for carrying out the invention]

[0019] 2 encoders, decoders The following are various examples that may help achieve more effective compression when using intra-prediction. Some examples achieve improved compression efficiency by consuming a set of intra-prediction modes. The latter may be added to, for example, other heuristically designed intra-prediction modes, or provided exclusively. Still other examples utilize both of the aforementioned specialties.

[0020] To facilitate understanding of the following examples of this application, the description begins with a presentation of possible encoders and corresponding decoders that can construct the examples of this application outlined thereafter. Figure 1 shows a device for encoding a picture 10 into a data stream 12 in blocks. The device is indicated by reference numeral 14 and may be a still image encoder or a video encoder. In other words, if the encoder 14 is configured to encode a video 16 containing the picture 10 into a data stream 12, the picture 10 may be the current picture in the video 16, or the encoder 14 may exclusively encode the picture 10 into the data stream 12.

[0021] As described above, the encoder 14 performs encoding on a block-by-block or block-based basis. For this purpose, the encoder 14 subdivides the picture 10 into blocks, and in units of blocks, the encoder 14 encodes the picture 10 into the data stream 12. Examples of possible subdivisions of the picture 10 into blocks 18 are described in more detail below. In general, the subdivision can end up as blocks 18 of a fixed size, such as an array of blocks arranged in rows and columns, or blocks 18 of different block sizes, starting from the entire picture area of ​​the picture 10, or from a pre-subdivision of the picture 10 to an array of tree blocks, such as using a hierarchical multitree subdivision to start a multitree subdivision into an array of tree blocks, and these examples are not treated as excluding other possible ways of subdividing the picture 10 into blocks 18.

[0022] Furthermore, the encoder 14 is a predictive encoder configured to predictively encode the picture 10 into the data stream 12. For a particular block 18, this means that the encoder 14 determines the predictive signal for block 18 and encodes the predictive residual, i.e., the prediction error by which the predictive signal deviates from the actual picture content within block 18, into the data stream 12.

[0023] The encoder 14 may support different prediction modes to derive the prediction signal for a given block 18. The prediction mode important in the following example is the intra-prediction mode, in which the interior of block 18 is spatially predicted from adjacent already encoded samples of picture 10. The encoding of picture 10 into the data stream 12, and therefore the corresponding decoding procedure, may be based on a specific encoding order 20 defined between the blocks 18. For example, the encoding order 20 may traverse the blocks 18 in a raster scan order such as top to bottom row by row, traversing each row from left to right. In the case of a hierarchical multitree-based subdivision, the raster scan ordering may be applied within each hierarchical level, and a depth-first traverse order may be applied, i.e., leaf nodes in a block at a particular hierarchical level may precede blocks at the same hierarchical level that have the same parent block according to the encoding order 20. Depending on the encoding order 20, adjacent already encoded samples of block 18 may typically be located on one or more sides of block 18. In the examples presented herein, for example, adjacent already encoded samples in block 18 are located at the top and left of block 18.

[0024] The intra-prediction mode may not be the only mode supported by encoder 14. For example, if encoder 14 is a video encoder, encoder 14 may also support an inter-prediction mode in which block 18 is temporarily predicted from a previously encoded picture of video 16. Such an inter-prediction mode may also be a motion-compensated prediction mode in which motion vectors are signaled to block 18, such as indicating the relative spatial offset of the portion of block 18 from which the predicted signal of block 18 is derived as a duplicate. Additionally or alternatively, other non-intra-prediction modes may also be available, along with an inter-prediction mode when encoder 14 is a multi-view encoder, or a non-prediction mode in which the inside of block 18 is encoded as is, i.e., without prediction.

[0025] Before beginning to focus the description of this application on the intra-predictive mode, we will show a more specific example of a possible block-based encoder, i.e., a possible implementation of encoder 14, as described with respect to Figure 2, and then two corresponding examples of decoders that fit Figures 1 and 2, respectively.

[0026] Figure 2 shows a possible implementation of the encoder 14 of Figure 1, i.e., the encoder is configured to use transform coding to encode the predicted residuals, but this is merely an example and the application is not limited to this type of predicted residual coding. According to Figure 2, the encoder 14 includes a subtractor 22 configured to obtain a predicted residual signal 26 by subtracting the corresponding predicted signal 24 from the inbound signal, i.e., picture 10, or current block 18 on a block basis, which is then encoded into a data stream 12 by the predicted residual encoder 28. The predicted residual encoder 28 consists of an irreversible coding stage 28a and a reversible coding stage 28b. The irreversible stage 28a includes a quantizer 30 that receives the predicted residual signal 26 and quantizes samples of the predicted residual signal 26. As already mentioned above, this example uses transform coding of the predicted residual signal 26. Therefore, the irreversible coding stage 28a comprises a transform stage 32 connected between the subtractor 22 and the quantizer 30, which transforms the spectrally decomposed predicted residual 26 by quantization of the quantizer 30 performed on the transformed coefficients that present the residual signal 26. The transform may be DCT, DST, FFT, Hadamard transform, etc. The transformed and quantized predicted residual signal 34 is then subjected to reversible coding by the reversible coding stage 28b, which is the entropi-coded quantized predicted residual signal 34 of the entropicorder, to become the data stream 12. The encoder 14 further comprises a predicted residual signal reconstruction stage 36 connected to the output of the quantizer 30, thereby reconstructing the predicted residual signal from the transformed and quantized predicted residual signal 34 in a manner also usable by the decoder, i.e., taking coding loss into account in the quantizer 30. For this purpose, the prediction residual reconstruction stage 36 includes an inverse quantizer 38, which performs the inverse of the quantization of the quantizer 30, followed by an inverse converter 40, which performs the inverse transform for the transform performed by the converter 32, such as the inverse of spectral decomposition, such as the inverse of any of the specific transform examples described above. The encoder 14 includes an adder 42 that adds the reconstructed prediction residual signal output by the inverse converter 40 and the prediction signal 24 to output a reconstructed signal, i.e., a reconstructed sample. This output is fed to a predictor 44 of the encoder 14, which determines the prediction signal 24 based on it.This is a predictor 44 that supports all the prediction modes already described above with respect to Figure 1. Figure 2 also shows the case where the encoder 14 is a video encoder, and the encoder 14 may also include an in-loop filter 46 that filters a fully reconstructed picture, which, after being filtered, forms a reference picture for the predictor 44 with respect to the interprediction block.

[0027] As already mentioned above, the encoder 14 operates on a block basis. For the purposes of the following description, the block basis of interest involves subdividing the picture 10 into blocks, selecting an intra-prediction mode from a set or multiple intra-prediction modes supported by the predictor 44 or encoder 14 for each block, and executing the selected intra-prediction mode individually. However, there may be other types of blocks into which the picture 10 is subdivided. For example, the above-mentioned decision of whether the picture 10 is inter-encoded or intra-encoded may be made at a granular level, or in units of blocks that deviate from block 18. For example, the inter / intra mode decision may be made at the level of encoded blocks, where the picture 10 is subdivided and each encoded block is subdivided into prediction blocks. Prediction blocks having encoded blocks for which intra-prediction has been determined to be used are each subdivided into intra-prediction mode decisions. For each of these prediction blocks, it is determined which supported intra-prediction mode should be used for each prediction block. These prediction blocks form block 18 of interest here. Prediction blocks within a coded block associated with interprediction are processed differently by the predictor 44. They are interpredicted from a reference picture by determining a motion vector and replicating the prediction signal for this block from the position in the reference picture indicated by the motion vector. Another block subdivision relates to subdivision into transformed blocks in units in which transformations are performed by the converter 32 and the inverse converter 40. The transformed blocks may, for example, be the result of further subdivision of a coded block. Naturally, the embodiments described herein should not be treated as limiting, and other embodiments exist. For completeness only, it should be noted that subdivision into coded blocks may use, for example, multitree subdivision, and prediction blocks and / or transformed blocks may also be obtained by further subdivision of a coded block using multitree subdivision.

[0028] Figure 3 shows a decoder 54 or device for block-level decoding that is compatible with the encoder 14 in Figure 1. This decoder 54 does the opposite of the encoder 14, namely decoding from the data stream 12 picture 10 block by block, and for this purpose supports multiple intra-prediction modes. The decoder 54 may include, for example, a residual provider 156. All other possibilities described above with respect to Figure 1 are also valid for the decoder 54. In contrast, the decoder 54 may be a still image decoder or a video decoder, and all prediction modes and predictability are also supported by the decoder 54. The difference between the encoder 14 and the decoder 54 is mainly the fact that the encoder 14 selects coding decisions according to some optimization to minimize some cost function which may depend, for example, on the coding rate and / or coding distortion. One of these coding options or coding parameters may include the selection of an intra-prediction mode to be used for the current block 18 from among the available or supported intra-prediction modes. Next, the selected intra-prediction mode may be signaled by the encoder 14 of the current block 18 in the data stream 12 by the decoder 54 re-executing the selection using this signaling in the data stream 12 of block 18. Similarly, the subdivision of picture 10 into block 18 may be subjected to optimization within encoder 14, and the corresponding subdivision information may be transmitted in the data stream 12 by the decoder 54 restoring the subdivision of picture 10 into block 18 based on the subdivision information. To summarize the above, decoder 54 may be a block-based predictive decoder, and in addition to the intra-prediction mode, decoder 54 may support other predictive modes, such as inter-prediction mode, if decoder 54 is a video decoder. In decoding, decoder 54 may also use the coding order 20 described with respect to Figure 1, which is observed in both encoder 14 and decoder 54 so that the same neighboring samples are available in the current block 18 in both encoder 14 and decoder 54.Therefore, to avoid unnecessary repetition, the description of the operating modes of the encoder 14 shall also apply to the decoder 54 as far as subdivision of the picture 10 into blocks, for example, as far as prediction and as far as coding of prediction residuals. The difference lies in the fact that the encoder 14, by optimization, selects several coding options or coding parameters and signals, or inserts coding parameters into the data stream 12, which are then derived from the data stream 12 for the decoder 54 to re-execute prediction, subdivision, etc.

[0029] Figure 4 shows possible implementations of the decoder 54 in Figure 3, i.e., implementations that are compatible with the implementation of the encoder 14 in Figure 1 shown in Figure 2. Since many elements of the encoder 54 in Figure 4 are the same as those that occur in the corresponding encoder in Figure 2, the same apostrophe-bearing reference codes are used in Figure 4 to indicate these elements. In particular, the adder 42', the optional in-loop filter 46', and the predictor 44' are connected to the predictive loop in the same way as the encoder in Figure 2. The reconstructed, i.e., inversely quantized and retransformed predictive residual signal applied to the adder 42' is derived by a sequence of entropy decoders 56 that reverse the entropy coding of the entropy encoder 28b, followed by a residual signal reconstruction stage 36' consisting of an inverse quantizer 38' and an inverse converter 40', as in the case of the coding side. The output of the decoder is the reconstruction of picture 10. The reconstruction of picture 10 may be available directly at the output of the adder 42' or at the output of the in-loop filter 46'. To improve picture quality, several post-filters could be placed at the decoder output to subject the reconstruction of picture 10 to some form of post-filtering, although this option is not shown in Figure 4.

[0030] Here again, with respect to Figure 4, the explanation given above for Figure 2 is equally valid for Figure 4, except that only the encoder performs the relevant decisions regarding the optimization task and coding options. However, all explanations regarding block subdivision, prediction, inverse quantization, and retransformation are also valid for the decoder 54 in Figure 4. 3 ALWIP

[0031] Some non-exclusive examples of ALWIP are described herein, even if ALWIP is not necessarily required to embody the technologies described herein.

[0032] This application relates, in particular, to an improved intra-predictive mode concept for block-based picture coding, such as that usable with video codecs such as HEVC or any successor to HEVC.

[0033] Intra-prediction mode is widely used in picture and video coding. In video coding, intra-prediction mode competes with other prediction modes such as inter-prediction modes, including motion-compensated prediction mode. In intra-prediction mode, the current block is predicted based on neighboring samples, i.e., samples that have already been coded as far as the encoder side is concerned and already decoded as far as the decoder side is concerned. The neighboring sample values ​​are extrapolated to the current block to form the predicted signal for the current block, and the predicted residual is transmitted in the data stream of the current block. The better the predicted signal, the smaller the predicted residual will be, and therefore fewer bits will be needed to encode the predicted residual.

[0034] To be effective, several aspects should be considered in order to form an effective framework for intra-prediction in a block-based picture coding environment. For example, the more intra-prediction modes supported by the codec, the greater the side information rate consumption for signaling the selection to the decoder. On the other hand, the set of supported intra-prediction modes must be able to provide good prediction signals, i.e., prediction signals that yield low prediction residuals.

[0035] When using the improved intra-predictive mode concept, there is a need for an intra-predictive mode concept that enables more efficient compression of block-based picture codecs.

[0036] This objective is achieved, in particular, by a so-called affine linear weighted intra-predictor (ALWIP) transformation. A device (encoder or decoder) for decoding pictures in blocks from a data stream is disclosed, which supports at least one intra-prediction mode in which an intra-prediction signal for a given size block of the picture is determined by applying a first template of adjacent samples to the current block to an affine linear predictor called an affine linear weighted intra-predictor (ALWIP) in the sequence.

[0037] The device may have at least one of the following characteristics (the same may apply to a method or other technique, for example, which is implemented in a non-temporary memory unit that stores instructions causing a processor to perform a method and / or to operate as a device when performed by a processor): 3.1 The proposed predictors may be complementary to other predictors.

[0038] The intra-prediction modes supported by the device are complementary to the other intra-prediction modes of the codec. Therefore, they may be complementary to the DC prediction mode, planar prediction mode, or angular prediction mode defined in the HEVC codec response. (See JEM reference software.) Hereafter, the latter three types of intra-prediction modes will be referred to as conventional intra-prediction modes. Therefore, for a given block of intra-modes, the decoder must parse a flag indicating whether one of the intra-prediction modes supported by the device should be used. 3.2 Two or more proposed prediction modes

[0039] A device may have two or more ALWIP modes. Therefore, if the decoder knows that one of the ALWIP modes supported by the device should be used, the decoder needs to parse additional information indicating which of the ALWIP modes supported by the device should be used.

[0040] The signal transmission of supported modes may have the characteristic that encoding some ALWIP modes may require fewer bins than other ALWIP modes. Which of these modes requires fewer bins and which require more may depend on information that can be extracted from the already decoded bitstream or fixed in advance. 4. Several aspects

[0041] Figure 2 shows a decoder 54 for decoding a picture from a data stream 12. The decoder 54 may be configured to decode a given block 18 of the picture. In particular, the predictor 44 may be configured to map a set of P adjacent samples adjacent to a given block 18 to a set of Q predicted values ​​of the samples in the given block using a linear or affine linear transformation [e.g., ALWIP].

[0042] As shown in Figure 5, a given block 18 contains Q predicted values ​​(which become the “predicted values” at the end of the operation). If block 18 has M rows and N columns, then Q = M * N. The Q values ​​of block 18 can be in the spatial domain (e.g., pixels) or the transformation domain (e.g., DCT, discrete wavelet transform, etc.). The Q values ​​of block 18 can generally be predicted based on the P values ​​obtained from adjacent blocks 17a to 17c adjacent to block 18. The P values ​​of adjacent blocks 17a to 17c may be the closest to block 18 (e.g., adjacent). The P values ​​of adjacent blocks 17a to 17c have already been processed and predicted. The P values ​​are shown as values ​​in parts 17'a to 17'c to distinguish the P values ​​from the blocks they are part of (in some examples, 17'b is not used).

[0043] As shown in Figure 6, to perform the prediction, it is possible to operate using a first vector 17P with P entries (each entry associated with a specific position in the adjacent portion 17'a~17'c), a second vector 18Q with Q entries (each entry associated with a specific position in block 18), and a mapping matrix 17M (each row associated with a specific position in block 18, each column associated with a specific position in the adjacent portion 17'a~17'c). Thus, the mapping matrix 17M performs the prediction of the P values ​​of the adjacent portion 17'a~17'c to the values ​​in block 18 according to a predetermined mode. Therefore, the entries in the mapping matrix 17M can be understood as weight coefficients. In the following sections, the signs 17a~17c are used instead of 17'a~17'c to refer to the adjacent portions of the boundary.

[0044] In this technical field, several conventional modes are known, including DC modes, planar modes, and 65 directional predictive modes. For example, 67 modes are known.

[0045] However, it should be noted that it is also possible to use a different mode called linear or affine linear transformation. A linear or affine linear transformation includes P*Q weight coefficients, of which at least 1 / 4 of the P*Q weight coefficients are non-zero weight values, and for each of the Q predicted values, it includes a set of P weight coefficients associated with each predicted value. When arranged vertically in the order of the raster scan between samples of a given block, the continuity forms an envelope that is nonlinear in all directions.

[0046] Figure 13 shows an example of Figure 70 mapping P positions (templates) of adjacent values ​​17'a~17'c, Q positions of adjacent samples 17'a~17'c, and P*Q values ​​of weight coefficients in matrix 17M. Plane 72 is an example of the envelope of a continuum of DC transformations (it is a plane for DC transformations). The envelope is obviously a plane and is therefore excluded by the definition of linear or affine linear transformations (ALWIP). Another example is a matrix that results in the emulation of angular modes, where the envelope is excluded from the ALWIP definition and, simply put, looks like a hill that slopes diagonally from top to bottom along the direction in the P / Q plane. Planar modes and directional predictive modes have different envelopes, but this is linear in at least one direction, i.e., all directions of the exemplified DC and, for example, the hill direction of the angular modes.

[0047] Conversely, the envelopes of linear or affine transformations are not omnidirectionally linear. It is understood that such types of transformations may, in some situations, be optimal for performing the predictions in block 18. Note that it is preferable that at least 1 / 4 of the weight coefficients are non-zero (i.e., at least 25% of the P*Q weight coefficients are non-zero).

[0048] The weight coefficients may be independent of each other according to any regular mapping rule. Therefore, matrix 17M may be such that the values ​​of its entries do not have any obvious, recognizable relationships. For example, the weight coefficients cannot be described by any analytical or differential function.

[0049] In the embodiment, the ALWIP transformation may result in the mean of the maximum cross-correlation between a first set of weight coefficients associated with each predicted value and a second set of weight coefficients associated with the other predicted values, or an inverted version of the latter, being lower than a predetermined threshold (e.g., 0.2 or 0.3 or 0.35 or 0.1, e.g., a threshold in the range of 0.05 to 0.035), regardless of whether the maximum is higher. For example, for each join (i1,i2) of the rows of the ALWIP matrix 17M, the cross-correlation may be calculated by multiplying the P-value of the i1st row by the P-value of the i2nd row. For each resulting cross-correlation, a maximum value may be obtained. Thus, the mean can be obtained for the entire matrix 17M (i.e., the maximum cross-correlation values ​​for all combinations are averaged). The threshold may then be, for example, 0.2 or 0.3 or 0.35 or 0.1, e.g., a threshold in the range of 0.05 to 0.035.

[0050] The P adjacent samples in blocks 17a to 17c may be arranged along a one-dimensional path extending along the boundary of a given block 18 (e.g., 18c, 18a). For each of the Q predicted values ​​in a given block 18, a set of P weight coefficients associated with each predicted value may be ordered to traverse the one-dimensional path in a predetermined direction (e.g., left to right, top to bottom, etc.). In this example, the ALWIP matrix 17M may be off-diagonal or non-block diagonal. An example of an ALWIP matrix 17M for predicting a 4x4 block 18 from four already predicted adjacent samples may be as follows: { {37,59,77,28}, {32,92,85,25}, {31,69,100,24}, {33,36,106,29}, {24,49,104,48}, {24,21,94,59}, {29,0,80,72}, {35,2,66,84}, {32,13,35,99}, {39,11,34,103}, {45,21,34,106}, {51,24,40,105}, {50,28,43,101}, {56,32,49,101}, {61,31,53,102}, {61,32,54,100} }.

[0051] (Here, {37,59,77,28} is the first row, {32,92,85,25} is the second row, and {61,32,54,100} is the 16th row of matrix 17M.) Matrix 17M has dimensions 16x4 and contains 64 weight coefficients (as a result of 16*4=64). This is because matrix 17M has dimensions QxP, where Q=M*N is the number of samples in the block 18 to be predicted (block 18 is a 4x4 block), and P is the number of samples of the already predicted samples. Here, M=4, N=4, Q=16 (as a result of M*N=4*4=16), and P=4. The matrix is ​​off-diagonal and non-block diagonal and is not described by any particular rule.

[0052] As can be seen, weight coefficients less than 1 / 4 are 0 (in the case of the matrix above, one of the 64 weight coefficients is 0). The envelope formed by these values, when arranged vertically according to the raster scan order, forms an envelope that is nonlinear in all directions.

[0053] Even if the above explanation primarily refers to a decoder (e.g., decoder 54), the same may be performed on an encoder (e.g., encoder 14).

[0054] In some examples, for each block size (in a set of block sizes), the ALWIP transforms of the intra-prediction modes in a second set of intra-prediction modes for each block size are different from one another. Additionally or alternatively, the concentrations of the second set of intra-prediction modes for a block size in a set of block sizes may coincide, but the associated linear or affine linear transforms of the intra-prediction modes in the second set of intra-prediction modes for different block sizes may not be transferable to each other by scaling.

[0055] In some examples, an ALWIP transformation may be defined as having "nothing in common" with a conventional transformation (for example, an ALWIP transformation has "nothing in common" with the corresponding conventional transformation even if it is mapped via one of the above mappings).

[0056] In one example, ALWIP mode is used for both the luminous and chroma components, while in another example, ALWIP mode is used for the luminous component but not for the chroma component. 5. Affine linear weighted intra-prediction mode achieved by speeding up the encoder (e.g., Test CE3-1.2.1:) 5.1 Description of method or apparatus

[0057] The affine linear weighted intra-prediction (ALWIP) mode tested in CE3-1.2.1 may be the same as that proposed in JVET-L0199 under test CE3-2.2.2, except for the following changes:

[0058] Multiple Reference Lines (MRL) intra-prediction, particularly the matching with encoder estimation and signaling, means that MRL is not combined with ALWIP, and the transmission of MRL indices is restricted to non-ALWIP blocks.

[0059] Subsampling is mandatory for all blocks, W×H ≥ 32×32 (previously optional for 32×32), and the transmission of additional tests and subsampling flags in the encoder has been removed.

[0060] ALWIP for 64×N and N×64 blocks (N≦32) is added by downsampling to 32×N and N×32, respectively, and applying the corresponding ALWIP mode. Furthermore, test CE 3-1.2.1 includes the following encoder optimizations for ALWIP:

[0061] Combined mode estimation: Conventional modes and ALWIP modes use a shared Hadamard candidate list for full RD estimation; i.e., ALWIP mode candidates are added to the same list as conventional (and MRL) mode candidates based on their Hadamard costs. EMT intrafast and PB intrafast are supported for join mode lists and have additional optimizations to reduce the number of full RD checks. Only the MPMs from the available left and top blocks are added to the list for the full RD estimation of ALWIP, following the same methodology as in the conventional mode. 5.2 Complexity Assessment

[0062] Excluding the calculations that invoke the discrete cosine transform, test CE3-1.2.1 required up to 12 multiplications per sample to generate the predicted signal. Furthermore, a total of 136,492 parameters, each 16 bits, were required. This corresponds to 0.273 megabytes of memory. 5.3 Experimental Results

[0063] The test evaluation was performed using VTM software version 3.0.1 for intra-only (AI) and random access (RA) configurations, following common test conditions JVET-J 1010[2]. The corresponding simulations were run on an Intel Xeon cluster (E5-2697A v4, AVX2 on, Turbo Boost off) with Linux® OS and GCC 7.2.1 compiler.

[0064] [Table 1] 5.4 Test CE3-1.2.2: Affine linear weighted intraprediction with complexity reduction

[0065] The technique tested in CE2 is related to the “affine linear intraprediction” described in JVET-L0199[1], but simplifies it in terms of memory requirements and computational complexity.

[0066] Only three different sets of prediction matrices (e.g., S0, S1, S2, see below) and bias vectors (e.g., to provide offset values) can exist to cover all block shapes. As a result, the number of parameters is reduced to 14,400 10-bit values, which requires less memory than 128 × 128 CTUs to store.

[0067] The input and output sizes of the predictor are further reduced. Furthermore, instead of transforming the boundary via DCT, averaging or downsampling can be performed on the boundary samples, and the generation of the predictive signal can use linear interpolation instead of inverse DCT. As a result, up to four multiplications per sample may be required to generate the predictive signal. 6. Examples This section describes how to perform several predictions (for example, as shown in Figure 6) using ALWIP prediction.

[0068] In principle, as shown in Figure 6, to obtain the Q=M*N values ​​for the predicted MxN block 18, the multiplication of Q*P samples from the QxP ALWIP prediction matrix 17M and P samples from the Px1 neighbor vector 17P should be performed. Therefore, in general, to obtain each of the Q=M*N values ​​for the MxN block 18 to be predicted, multiplication of at least P=M+N values ​​is required.

[0069] These multiplications have highly undesirable effects. The dimension P of the boundary vector 17P generally depends on the number M+N of boundary samples (bins or pixels) 17a, 17c adjacent to the MxN block 18 to be predicted (e.g., horizontally). This means that as the size of the block 18 to be predicted increases, the number of boundary pixels M+N(17a,17c) increases accordingly, which in turn increases the dimension P=M+N of the Px1 boundary vector 17P, the length of each row of the QxP ALWIP prediction matrix 17M, and consequently the number of multiplications required (generally speaking, Q=M*N=W*H, where W (width) is another symbol for N, H (height) is another symbol for M, and P=M+N=H+W if the boundary vector is formed by only one row and / or one column of samples).

[0070] This problem is generally exacerbated by the fact that multiplication is generally a power-consuming operation in microprocessor-based systems (or other digital processing systems). A large number of multiplications performed on a very large number of samples across a large number of blocks can generally be assumed to result in an undesirable waste of computational power. Therefore, it is preferable to reduce the number of multiplications Q*P required to predict the MxN block 18.

[0071] It is understood that by intelligently selecting an easier operation to perform instead of multiplication, it is possible to reduce in some way the computational power required for each intra-prediction of each predicted block 18.

[0072] In particular, referring to FIGS. 7.1 to 7.4, FIG. 8.1, and FIG. 8.2, an encoder or decoder uses a plurality of adjacent samples (e.g., 17a, 17c) to

[0073] reduce the plurality of adjacent samples (e.g., by averaging or downsampling) (e.g., step 811) to obtain a reduced set of sample values with fewer samples compared to the plurality of adjacent samples (e.g., 17a, 17c), and

[0074] subject the reduced set of sample values to a linear or affine-linear transformation (e.g., step 812) to obtain a predicted value of a predetermined sample of a predetermined block, and it is understood that a predetermined block (e.g., 18) of a picture can be predicted thereby.

[0075] In some cases, the decoder or encoder may also derive a predicted value of a further sample of a predetermined block based on a predetermined sample and a predicted value of a plurality of adjacent samples, e.g., by interpolation (e.g., step 813 in FIG. 8.1). Thus, an upsampling strategy can be obtained.

[0076] In an example, it is possible to perform some averaging on the samples of the boundary 17 (e.g., in step 811) so as to reach a reduced set 102 of samples (FIGS. 7.1 to 7.4) with fewer samples (at least one of the samples of the sample-reduced 102 may be the average of two samples of the original boundary samples or the selection of the original boundary samples). For example, if the original boundary has P = M + N samples, the reduced set of samples has P red = M red + N red where P red < M and N red < N with at least one of them, whereby P red<P can be obtained. Therefore, the boundary vector 17P actually used for prediction (e.g., in step 812b) does not have a Px1 entry, and P red <P which is P red has a Px1 entry. Similarly, the ALWIP prediction matrix 17M selected for prediction does not have a QxP dimension and has at least P red <P(M red <M and N red <N (by at least one of) reduces the number of elements of the matrix to QxP red (or Q red xP red (see below).

[0077] In some examples (e.g., FIGS. 7.2, FIGS. 7.3, and FIGS. 8.2), the block obtained by ALWIP (in step 812) has a size

Number

Number

Number

Number

[0078] These techniques reduce the number of multiplications (Q) in matrix multiplication. red *P red or Q*P red While this may include, it can be advantageous that both the initial reduction (e.g., averaging or downsampling) and the final transformation (e.g., interpolation) can be performed by reducing (or avoiding) multiplication. For example, downsampling, averaging and / or interpolation can be performed by employing binary operations that do not require computational power, such as addition and shifting (e.g., in steps 811 and / or 813).

[0079] Here, we describe an example of a shift operation at the processor level. Figure 9 shows the number 580 encoded in binary (1001000100b) in a 10-bit register 910, where the 10-bit register 910 has a 10-bit register 910a, each bit register 910a storing a 1-bit value (e.g., 1 or 0). The value 1001000100b, represented in binary encoded in the 2-byte register 910, is shown as 901 in Figure 9a as the "shifted value". The binary value shown as sign 902 in Figure 9b is the right-shifted version of the binary value shown as sign 901. As can be seen, the binary value 902 (encoded as 0100100010b) is the version of the binary value 901 after each value encoded in each bit register 910a has been simply moved to the bit register to its right, the lower bits of the binary value 901 are lost in the binary value 902, and the most significant bit of the shifted binary value 902 is added as 0. When the shifted binary value 902 is converted to a decimal, we find that we obtain the decimal value 290, which is half of 580. This could be a technique for bisecting a binary number (only quantization error exists if the value 901 is odd). This operation is extremely easy to perform and does not require high computational power. In particular, by shifting multiple times, division by a power of 2 can be obtained; for example, by shifting twice, division by 4 can be obtained, by shifting three times, division by 8 can be obtained, and so on, by shifting r times. r Division by (2^r) is obtained. This can also be expressed as f>>r, and therefore f>>1 means f / 2, f>>2 means f / 4, and so on (f is an integer). This operation is also known as right rotation. Similarly, the left rotation operation f< <rは、fに2 rThis means multiplying by (or 2^r). Shift operations are not computationally intensive for the processor, as they avoid the need to perform multiplication, and are achieved simply by moving bits between different registers. In this example, register 910 is represented as a 10-bit register. However, in this example, register 910 may have a different number of bit registers 910a, for example, an 8-bit register 910 (in which case register 910 is an 8-bit register). Furthermore, addition is a very simple operation that can be easily performed without much computational effort.

[0080] This shift operation can be used, for example, to average two boundary samples to obtain the final predicted block, and / or to interpolate two samples (support values) of the reduced predicted block (or obtained from the boundary). (Two sample values ​​are needed for interpolation. Within a block, there are always two predetermined values, but to interpolate samples along the left and top boundaries of the block, there is only one predetermined value, as shown in Figure 7.2, and therefore the boundary sample is used as the support value for interpolation.) The following two-step procedure may be used: First, sum the values ​​of the two samples, Next, the total value is halved (for example, by right-shifting). or, First, divide each sample in half (for example, by left-shifting), Next, sum the values ​​of the two halved samples. It is possible.

[0081] Since only one sample quantity and group of samples (e.g., adjacent samples) need to be selected, even easier operation can be performed during downsampling (e.g., in step 811).

[0082] Therefore, it is possible here to define a technique (or techniques) for reducing the number of multiplications to be performed. Some of these techniques may be based, among other things, on at least one of the following principles:

[0083] Even if the size of the actually predicted block 18 is MxN, the block is reduced (in at least one of the two dimensions) to a reduced size Q red xP red ( [Number] , P red = N red + M red and [Number] and / or M red < M and / or N red < N) of the ALWIP matrix may be applied. Thus, the boundary vector 17P has a size P red x1 and P red < means only P multiplications (P red = M red + N red and P = M + N). P red x 1 boundary vector 17P can be easily obtained from the original boundary 17, for example, by downsampling (for example, by selecting only some samples of the boundary) and / or by averaging multiple samples of the boundary (which can be easily obtained by addition and shift without multiplication).

[0084] Additionally or alternatively, instead of predicting all Q = M*N values of the prediction target block 18 by multiplication, a reduced block with reduced dimensions (for example, [Number] It is possible to predict only the remaining samples of block 18 to be predicted, for example, the remaining QQ to be predicted. red Q as a support value for the value red It is obtained by interpolation using a sample.

[0085] An example that can be understood as a general description of process 810 is provided by Figure 8.1, and its specific case is shown in Figure 7.1. In this case, a 4x4 block 18 (M=4, N=4, Q=M*N=16) is to be predicted, and the neighborhoods 17 of samples 17a (a vertical matrix with 4 already predicted samples) and 17c (a horizontal row with 4 already predicted samples) have already been predicted in the previous iteration (neighborhoods 17a and 17c can be collectively denoted by 17). Preferentially, by using the formula shown in Figure 5, the prediction matrix 17M should be a QxP=16x8 matrix (Q=M*N=4*4 and P=M+N=4+4=8), and the boundary vector 17P should have 8x1 dimensions (P=8). However, this leads to the need to perform 8 multiplications for each of the 16 samples of the 4x4 block 18 to be predicted, and thus a total of 16*8=128 multiplications. (Note that the average number of multiplications per sample is a good indicator of computational complexity. Conventional intraprediction requires four multiplications per sample, which increases the computational effort involved. Therefore, it is possible to use this as an upper limit for ALWIP, ensuring that the complexity is reasonable and does not exceed the complexity of conventional intraprediction.)

[0086] Nevertheless, by using this technology, in step 811, the number of adjacent samples 17a and 17c to the predicted block 18 is changed from P to P redIt is understood that it is possible to reduce to . In particular, in order to obtain a reduced boundary 102 having two horizontal rows and two vertical columns, it is possible to average adjacent boundary samples (17a, 17c) with each other (for example, at 100 in FIG. 7.1), and thus it is understood that operating as block 18 is a 2×2 block (the reduced boundary formed by the average value). Alternatively, it is possible to perform downsampling, and thus select two samples for row 17c and two samples for column 17a. Thus, the horizontal row 17c is processed (e.g., averaged samples) as having two samples instead of having four original samples, and the vertical column 17a, which originally has four samples, is processed as having two samples (e.g., averaged samples). After sub-dividing the row 17c and the column 17a into groups 110 of two samples each, it can also be understood that a single sample is maintained (e.g., the average of the samples in group 110 or a simple selection between the samples in group 110). Thus, by having only a set 102 of four samples (M red =2, N red =2, P red =M red +N red =4, P red <P), a so-called reduced set 102 of sample values is obtained.

[0087] It is understood that it is possible to perform operations (such as averaging or downsampling 100) without performing an excessive number of multiplications at the processor level, and the averaging or downsampling 100 performed in step 811 can be easily obtained by operations that do not require simple computational capabilities such as addition and shift.

[0088] At this point, it is understood that the reduced set of sample values ​​102 can be subjected to a linear or affine linear (ALWIP) transformation 19 (for example, using a prediction matrix such as matrix 17M in Figure 5). In this case, the ALWIP transformation 19 directly maps the four samples 102 to the sample values ​​104 of block 18. In this case, interpolation is not required.

[0089] In this case, the ALWIP matrix 17M has dimension QxP red It has =16x4. This follows the fact that all Q=16 samples of block 18 to be predicted can be obtained directly by ALWIP multiplication (interpolation is not required).

[0090] Therefore, in step 812a, dimension QxP red A suitable ALWIP matrix 17M having A is selected. The selection may be based, for example, at least in part on signaling from data stream 12. The selected ALWIP matrix 17M is also A k It may also be shown as follows, where k is the data stream 12 (in some cases the matrix is ​​as follows)

number

[0091] In step 812b, the selected QxP red ALWIP matrix 17M(A k (Also shown as) and P red Multiplication with the x1 boundary vector 17P is performed.

[0092] In step 812c, the offset value (for example, b k ) can be added to all acquired values ​​104 of vector 18Q obtained by, for example, ALWIP. The offset value (b k , or in some cases further

number

[0093] Therefore, the comparison between using this technology and not using it is resumed here: If this technology is not used: Prediction target block 18, dimensions M=4, N=4, The predicted Q = M * N = 4 * 4 = 16 values, P = M + N = 4 + 4 = 8 boundary samples P = predicted Q = 16 values, each multiplied by 8 times. P*Q = 8*16 = 128 total number of multiplications. When using this technology, Prediction target block 18, dimensions M=4, N=4, Finally, the predicted Q = M * N = 4 * 4 = 16 values, The reduced dimension of the boundary vector: P red =M red +N red =2+2=4, For each of the 16 values ​​of Q that should be predicted by ALWIP, P red = Multiplication by 4, Total P red *Q = 4 * 16 = 64 multiplication (half of 128!) The ratio of the number of multiplications to the number of final values ​​obtained is: P red *Q / Q = 4, which means that for each predicted sample, P = 8 multiplied by half!

[0094] As can be understood, it is possible to obtain the appropriate value in step 812 by relying on operations that do not require direct computational power, such as averaging (and in some cases, addition and / or shifting and / or downsampling).

[0095] Referring to Figures 7.2 and 8.2, the predicted block 18 is an 8x8 block (M=8, N=8) with 64 samples. Here, preferentially, the prediction matrix 17M should have size QxP=64x16 (Q=M*N=8*8=64, Q=64 due to M=8 and N=8, and P=M+N=8+8=16). Therefore, preferentially, for each of the Q=64 samples in the predicted 8x8 block 18, P=16 multiplications are required, resulting in 64*16=1024 multiplications for the entire 8x8 block 18!

[0096] However, as can be seen in Figures 7.2 and 8.2, instead of using all 16 samples of the boundary, we can provide a method 820 in which only 8 values ​​(e.g., 4 from the horizontal boundary row 17c and 4 from the vertical boundary column 17a) are used. From the boundary row 17c, 4 samples may be used instead of 8 (e.g., they may be a 2x2 mean and / or a selection of one of the two samples). Thus the boundary vector is not a Px1 = 16x1 vector, but P red x1 = 8x1 is the only vector (P red =M red +N red =4+4). Instead of the original P=16 samples, P red It is understood that it is possible to select or average (e.g., 2x2) samples from horizontal row 17c and vertical column 17a to form a reduced set of sample values ​​102, which has only 8 boundary values. This reduced set 102 makes it possible to obtain a reduced version of block 18, and the reduced version is Q (instead of Q=M*N=8*8=64). red =M red *N red=4*4=16 samples. Size M red xN red An ALWIP matrix can be applied to predict a 4x4 block. The reduced version of block 18 includes the sample shown in gray in scheme 106 in Figure 7.2, and the sample shown in gray squares (including samples 118' and 118'') is the Q obtained in step 812. red A 4x4 decreasing block with 16 values ​​is formed. The 4x4 decreasing block is obtained by applying the linear transformation 19 in the target step 812. After obtaining the values ​​of the 4x4 decreasing block, it is possible to obtain the values ​​of the remaining samples (samples shown as white samples in scheme 106) by interpolation, for example.

[0097] With respect to method 810 in Figures 7.1 and 8.1, method 820 is the remaining QQ of the predicted MxN = 8x8 block 18. red Step 813 may further include deriving the predicted values ​​for 64-16=48 samples (white squares) by interpolation, for example. The remaining QQ red =64-16=48 samples were obtained directly by interpolation (interpolation can also utilize the values ​​of boundary samples, for example). red = can be obtained from 16 samples. As seen in Figure 7.2, samples 118' and 118'' are obtained in step 812 (indicated by gray squares), while sample 108' (which is midway between samples 118' and 118'' and is indicated by a white square) is obtained in step 813 by interpolation between samples 118' and 118''. It is understood that interpolation can also be obtained by operations similar to those for averaging, such as shifting and adding. Therefore, in Figure 7.2, the value 108' can generally be determined as a value midway between the value of sample 118' and the value of sample 118'' (it could be the average).

[0098] By performing interpolation, it is also possible to reach the final version of the MxN = 8x8 block 18 based on the plurality of sample values shown in 104 in step 813. Therefore, comparing using this technique with not using it, When not using this technique: The prediction target block 18 has dimensions M = 8, N = 8 and has Q = M * N = 8 * 8 = 64 samples within the prediction target block 18. P = M + N = 8 + 8 = 16 samples within the boundary 17. For each of the Q = 64 predicted values, P = 16 multiplications. Total number of multiplications P * Q = 16 * 64 = 1028. The ratio of the number of multiplications to the number of final values obtained is P * Q / Q = 16. When using this technique, The prediction target block 18 has dimensions M = 8, N = 8. Finally, Q = M * N = 8 * 8 = 64 values are predicted.

[0099] However, Q red xP red The ALWIP matrix is used, and P red = M red + N red , Q red = M red * N red , M red = 4, N red = 4, and P within the boundary red = M red + N red = 4 + 4 = 8 samples, where P red < P. For each of the Q = 16 values of the predicted 4x4 reduced block red P red = 8 multiplications (formed by the gray squares in scheme 106), Total number of P red * Q red = 8 * 16 = 128 multiplications (substantially less than 1024!), The ratio of the number of multiplications to the number of final values obtained is P red*Q red / Q = 128 / 64 = 2 (much less than 16 obtained without using this technology!). Therefore, the technology presented in this specification requires 8 times less computing power than the previous technology.

[0100] Figure 7.3 shows another example (obtainable based on method 820), where the predicted block 18 is a rectangular 4×8 block (M = 8, N = 4) with Q = 4 * 8 = 32 predicted samples. The boundary 17 consists of a horizontal row 17c of N = 8 samples and a vertical column 17a of M = 4 samples. Therefore, preferably, the boundary vector 17P has dimension Px1 = 12x1, but the predicted ALWIP matrix should be a QxP = 32x12 matrix, and thus Q * P = 32 * 12 = 384 multiplications are required.

[0101] However, for example, it is possible to average or downsample at least 8 samples of the horizontal row 17c to obtain a reduced horizontal row of only 4 samples (e.g., averaged samples). In some examples, the vertical column 17a remains as it is (e.g., without averaging). In total, the reduced boundary has dimension P red = 8, and P red < P. Therefore, the boundary vector 17P has dimension P red x1 = 8x1. The ALWIP prediction matrix 17M becomes a matrix with dimension M * N red * P red = 4 * 4 * 8 = 64. The 4x4 reduced block (formed by the gray columns of scheme 107) directly obtained in the target step 812 has size Q red = M * N red = 4 * 4 = 16 samples (instead of Q = 4 * 8 = 32 of the original 4x8 block 18 to be predicted). When the reduced 4×4 block is obtained by ALWIP, the offset value b kIt is possible to add (step 812c) and perform interpolation in step 813. As can be seen in step 813 of Figure 7.3, the reduced 4x4 block is expanded into a 4x8 block 18, and the value 108' that was not obtained in step 812 is obtained in step 812 by interpolating the values ​​118' and 118'' (gray square) obtained in step 813.

[0102] Therefore, comparing the use of this technology with the non-use of it, If this technology is not used: The prediction target has 18 blocks, with dimensions M=4 and N=8. The predicted Q = M * N = 4 * 8 = 32 values, Within the boundary, P = M + N = 4 + 8 = 12 samples. For each of the 32 predicted values ​​of Q, multiply by P = 12. The total number P*Q = 12*32 = 384 multiplication. The ratio of the number of multiplication steps to the number of final values ​​obtained is P*Q / Q = 12. When using this technology, The block to be predicted is 18, and the block has dimensions M=4 and N=8. Finally, the predicted Q = M * N = 4 * 8 = 32 values,

[0103] However, Q red XP red =16x8 ALWIP matrix can be used, M=4,N red =4,Q red =M*N red =16,P red =M+N red =4+4=8, P within the boundary red =M+N red =4+4=8 samples, where P red <Pであり、 Predicted reduction in Q blocks red = For each of the 16 values ​​P red = Multiplication by 8, Total P red *Q red=8*16=128 multiplication (less than 384!) The ratio of the number of multiplications to the number of final values ​​obtained is P red *Q red / Q = 128 / 32 = 4 (much less than the 12 obtained without using this technique!). Therefore, this technology reduces computational effort by one-third.

[0104] Figure 7.4 shows an example of 18 blocks to be predicted with dimension MxN=16x16, where the final predicted values ​​are Q=M*N=16*16=256 and P=M+N=16+16=32 boundary samples. This results in a prediction matrix with dimension QxP=256x32, which means a multiplication of 256*32=8192!

[0105] However, by applying method 820, in step 811, it is possible to reduce the number of boundary samples from, for example, 32 to 8 (e.g., by averaging or downsampling), so that for every group 120 of four consecutive samples in row 17a, a single sample remains (e.g., selected from the four samples or the average of the samples). Also, for every group of four consecutive samples in column 17c, one sample remains (e.g., selected from the four samples or the average of the samples).

[0106] Here, the ALWIP matrix 17M is Q red XP red = 64x8 matrix. This is because it is P (by using 8 averaged samples or samples selected from 32 boundaries). red This is due to the fact that =8 was selected, and that the reduction block expected in step 812 is an 8x8 block (in scheme 109, the gray squares are 64).

[0107] Therefore, once 64 samples of the 8x8 block, reduced in step 812, are obtained, in step 813, the remaining QQ of the predicted block 18 are obtained. red=256-64=192 values ​​can be derived from 104.

[0108] In this case, it is chosen to use only all samples from boundary column 17a and alternative samples from boundary row 17c to perform interpolation. Other choices may be made.

[0109] In this method, the ratio between the number of multiplications and the number of final values ​​obtained is Q. red *P red / Q = 8 * 64 / 256 = 2, which is far less than the 32 multiplications of each value without using this technique! Comparing the use of this technology with the non-use of this technology, If this technology is not used: The prediction target has 18 blocks, with dimensions M=16 and N=16. The predicted Q = M * N = 16 * 16 = 256 values. Within the boundary, P = M + N = 16 * 16 = 32 samples. For each of the 256 predicted values ​​of Q, multiply by P = 32. The total number P*Q = 32*256 = 8192 multiplication The ratio of the number of multiplication steps to the number of final values ​​obtained is P*Q / Q = 32. When using this technology, The block to be predicted is 18, and the block has dimensions M=16 and N=16. Finally, the predicted Q = M * N = 16 * 16 = 256 values.

[0110] However, Q red XP red =64x8 ALWIP matrix can be used, M red =4,N red =4, and ALWIP predicts Q red =8*8=64 samples, P red =M red +N red =4+4=8 P within the boundary red =M red +N red=4+4=8 samples, where P red <Pであり、 Predicted reduction in Q blocks red = For each of the 64 values ​​P red = Multiplication by 8, Total Q red *P red =64*4=256 multiplication (This is less than 8192!) The ratio of the number of multiplications to the number of final values ​​obtained is P red *Q red Q = 8 * 64 / 256 = 2 (which is far less than the 32 obtained without using this technique!). Therefore, the computational power required by this technique is 16 times less than that of conventional techniques!

[0111] Therefore, by using multiple neighboring samples (17) and reducing the number of samples (100, 813) compared to multiple neighboring samples (17), we obtain a reduced set (102) with fewer sample values.

[0112] In order to obtain predicted values ​​for a given sample (104, 118', 188'') of a given block (18), a reduced set of sample values ​​(102) is subjected to a linear or affine linear transformation (19, 17M) (812) This makes it possible to predict a predetermined block (18) of the picture.

[0113] In particular, compared to multiple adjacent samples (17), it is possible to perform a reduction (100,813) by downsampling to obtain a reduced set (102) of sample values ​​with fewer samples.

[0114] Alternatively, compared to multiple adjacent samples (17), it is possible to perform a reduction (100,813) by averaging multiple adjacent samples to obtain a reduced set (102) of sample values ​​with fewer samples.

[0115] Furthermore, interpolation makes it possible to derive predicted values ​​for further samples (108, 108') of a given block (18) based on the predicted values ​​of a given sample (104, 118', 118'') and multiple adjacent samples (17) (813).

[0116] Multiple adjacent samples (17a, 17c) may extend one-dimensionally along two sides of a given block (18) (e.g., to the right and down in Figures 7.1 to 7.4). A given sample (e.g., obtained by ALWIP in step 812) may also be arranged in rows and columns, and may be positioned at every nth sample (112) of a given sample 112 adjacent to two sides of a given block 18 along at least one of the rows and columns.

[0117] Based on multiple adjacent samples (17), it is possible to determine a support value (118) for one of multiple adjacent positions (118) aligned in at least one row and column for each of the rows and columns. By interpolation, it is also possible to derive a predicted value 118 for a further sample (108, 108') in a given block (18) based on the predicted value of a given sample (104, 118', 118'') and the support value of the adjacent sample (118) aligned in at least one row and column.

[0118] A given sample (104) may be positioned every n samples (112) adjacent to two sides of a given block 18 along the row, and a given sample may be positioned every m samples (112) adjacent to two sides of a given block (18) along the column, where n,m>1. In some cases, n=m (for example, in Figures 7.2 and 7.3, samples 104, 118', 118'', which are directly acquired by ALWIP at 812 and shown as gray squares, are alternated along the row and column with samples 108, 108', which are subsequently acquired in step 813).

[0119] Along at least one of the rows (17c) and columns (17a), for each support value, it may be possible to determine the support value by, for example, downsampling or averaging (122) a group (120) of adjacent samples within a plurality of adjacent samples, including the adjacent sample (118) for which each support value is determined. Thus, in Figure 7.4, in step 813, it is possible to obtain the value of sample 119 by using the value of a given sample 118''' (previously obtained in step 812) and the value of the adjacent sample 118 as support values.

[0120] Multiple adjacent samples may extend one-dimensionally along two sides of a given block (18). It may be possible to perform reduction (811) by grouping multiple adjacent samples (17) into one or more consecutive groups (110) of adjacent samples and performing downsampling or averaging for each of the one or more groups (110) of adjacent samples having two or more adjacent samples.

[0121] In the example, a linear or affine linear transformation is P red *Q red or P red *It can include Q weighting coefficients, P red Q is the number of sample values ​​in the reduced set of sample values ​​(10²), and red Alternatively, Q is the number of samples (18) in a given block. At least 1 / 4 of P red *Q red or 1 / 4 P red *Q weight coefficients are non-zero weight values. red *Q red or P red *The Q weight coefficients are either Q or Q red For each of the specified samples, a series of P for each specified sample red It can include a number of weight coefficients, and a series of P redWhen the weight coefficients are arranged vertically according to the raster scan order between predetermined samples of a given block (18), they form an envelope that is nonlinear in all directions. red *Q or P red *Q red The weight coefficients may be independent of each other through a regular mapping rule. The mean of the maximum values ​​of the cross-correlation between a first set of weight coefficients associated with each given sample and a second set of weight coefficients associated with all other given samples, or an inverted version of the latter, is lower than a predetermined threshold, regardless of whether the maximum values ​​are higher. The predetermined threshold may be 0.3 (or possibly 0.2 or 0.1). red The adjacent samples (17) are arranged along a one-dimensional path extending along two sides of a given block (18), and Q or Q red For each of the specified samples, a series of P related to each specified sample red The weight coefficients can be ordered to traverse a one-dimensional path in a given direction. Description of Method and Apparatus

[0122]

number

number

[0123] 1. Of the boundary samples 17, 102 samples (e.g., 4 samples if W=H=4, and / or 8 samples otherwise) can be extracted by averaging or downsampling (e.g., step 811).

[0124] 2. Matrix-vector multiplication followed by offset addition can be performed with the averaged samples (or samples remaining from downsampling) as input. The result may be a reduced predicted signal on the subsampled set of samples in the original block (e.g., step 812).

[0125] 3. Predicted signals for the remaining positions can be generated, for example, by linear interpolation (e.g., step 813), from the predicted signals on the subsampled set, for example, by upsampling.

[0126] Thanks to step 1 (811) and / or 3 (813), the total number of multiplications required to calculate the matrix-vector product is always

number

[0127] In some examples, the matrix (e.g., 17M) and offset vector (e.g., b) required to generate the prediction signal are needed. k ) is, for example, a set of matrices (e.g., 3 sets) that can be stored in the memory units of the decoder and encoder, for example,

number

[0128] In some examples, set

number

Number

Number

Number

Number

[0129]

Number

Number

Number

Number

Number

Number

[0130] Additionally or alternatively, the set

Number

Number

Number

[0131] Set based on block dimensions

Number

[0132] As described above, the boundary samples (17a, 17c) can be averaged and / or downsampled (e.g., from P samples to P red <P samples).

[0133] In the first step, the input boundary

Number

Number

Number

Number

[0134] and

Number

Number

[0135] In all other cases (for example, blocks with widths or heights different from 4), the block width W is

number

number

number

number

[0136] In other cases, the number of samples can be reduced by downsampling the boundary (for example, by selecting one specific boundary sample from a group of boundary samples). For example,

number

number

number

[0137] Two reduced boundaries

number

number

number

number

number

number

[0138] Therefore, a specific state

number

number

[0139] Other strategies may be employed. In other examples, the mode index "mode" is not necessarily in the range of 0 to 35 (other ranges may be defined). Furthermore, each of the three sets S0, S1, and S2 is an 18-matrix (therefore,

number

number

[0140] Mode and transpose information are not necessarily stored and / or transmitted as a single combined mode index "mode". In some examples, they may be explicitly signaled as a transpose flag and matrix index (0-15 for S0, 0-7 for S1, 0-5 for S2).

[0141] In some cases, the combination of the transpose flag and the matrix index may be interpreted as a set index. For example, there may be one bit that acts as the transpose flag and several bits that represent the matrix index, which are collectively shown as a "set index". 5.5 Generation of Reduced Prediction Signals by Matrix-Vector Multiplication Features are provided here with respect to step 812.

[0142] Reduced input vector

number

number

number

number

[0143]

number

[0144] Reduced predictive signal

number

number

[0145] Here,

number

number

number

[0146]

number

number

number

number

[0147] Matrix A and vector

number

number

number

number

number

number

number

number

number

number

number

number

[0148] Other strategies may be employed. In other examples, the mode index "mode" is not necessarily in the range of 0 to 35 (other ranges may be defined). Furthermore, each of the three sets S0, S1, and S2 is an 18-matrix (therefore,

number

number

[0149] Interpolating subsampled prediction signals, especially in large blocks, may require a second version of the averaged boundary. That is,

number

number

number

number

number

number

number

[0150]

number

number

[0151] Linear interpolation may be given as follows (other examples are also possible):

number

number

number

number

number

number

number

number

number

number

number

[0152] Here,

number

number

[0153] This is an example of interpolation that uses reduced boundary samples for the first interpolation (horizontal or vertical) and the original boundary samples for the second interpolation (vertical or horizontal). Depending on the block size, only the second interpolation or no interpolation at all may be required. If both horizontal and vertical interpolation are needed, the order depends on the width and height of the block.

[0154] However, different techniques may be implemented, for example, the original boundary samples may be used for both the first and second interpolations, and the order may be fixed, for example, horizontal first, then vertical (otherwise, vertical first, then horizontal). Therefore, the interpolation order (horizontal / vertical) and the use of reduced / original boundary samples can be changed. 5.7 Explanation of an example of the entire ALWIP process

[0155] The entire process of averaging, matrix-vector multiplication, and linear interpolation is shown for different shapes in Figures 7.1 to 7.4. Note that the remaining shapes are treated as one of the illustrated examples.

[0156] Given a 4x4 block, ALWIP can take two averages along each axis of the boundary by using the technique shown in Figure 7.1. The resulting four input samples are then subjected to matrix-vector multiplication. The matrix is ​​set

number

[0157] Given an 8x8 block, ALWIP can take four averages along each axis of the boundary. The resulting eight input samples are then subjected to matrix-vector multiplication using the technique shown in Figure 7.2. The matrix is ​​set

number

[0158] Given an 8x4 block, ALWIP can take the four mean values ​​along the horizontal axis of the boundary and the four original boundary values ​​on the left boundary by using the technique shown in Figure 7.3. The resulting eight input samples are then subjected to matrix-vector multiplication. The matrix is ​​set

number

[0159] Given a 16x16 block, ALWIP can take four averages along each axis of the boundary. The resulting eight input samples are then subjected to matrix-vector multiplication using the technique shown in Figure 7.2. The matrix is ​​set

number

[0160] In the case of a W×8 block, since samples are given at odd horizontal and vertical positions, only horizontal interpolation is necessary. Therefore, in these cases, at most a multiplication of (8*64) / (16*8)=4 is performed per sample.

[0161] Finally, regarding W x 4 blocks where W > 8,

number

[0162] The parameters required for all possible proposed intra-prediction modes are set

number

[0163] For the Luma Block, for example, 35 ALWIP modes have been proposed (other numbers of modes may be used). For each coding unit (CU) in the intra-mode, a flag indicating whether or not the ALWIP mode should be applied to the corresponding predictive unit (PU) is transmitted in the bitstream. The signaling of the latter indicator can be harmonized with the MRL in the same way as in the initial CE trial. If an ALWIP mode is applied, the ALWIP mode index

number

[0164] Here, the derivation of the MPM may be performed using the intra-modes of the above and left PU, as follows: Each conventional intra-prediction mode

number

number

number

number

[0165] This indicates which of the three sets the ALWIP parameters should be interpreted as in Section 4 above. Prediction Units

number

number

number

number

[0166] The above PU is available, belongs to the same CTU as the current PU, is in intra mode, and is in conventional intra predictive mode.

number

[0167]

number

number

[0168] This means this mode is unavailable. Similarly, derive the modes without the restriction that the left PU must belong to the same CTU as the current PU:

number

[0169] Finally, three fixed default lists

number

number

[0170] The proposed ALWIP mode can be harmonized with conventional intra-predictive mode MPM-based coding as follows: Conventional intra-predictive mode luma and chroma MPM list derivation processes use fixed tables

number

number

[0171] In the case of Luma MPM list derivation, ALWIP mode

number

number

[0172] The test evaluation was performed using VTM software version 3.0.1 for intra-only (AI) and random access (RA) configurations, following common test conditions JVET-J 1010[2]. The corresponding simulations were run on an Intel Xeon cluster (E5-2697A v4, AVX2 on, Turbo Boost off) with Linux® OS and GCC 7.2.1 compiler.

[0173] [Table 2] 5.12 Further results from even faster encoders This provides two further results from tests that rely on the same syntax as CE3-1.2.2 but have optimized encoder lookups.

[0174] [Table 3] 6. Encoder in Figure 10

[0175] Figure 10 shows another example that can be interpreted from the examples in Figures 1, 2, and 5-9 (in particular, some features may be directly derived from Figure 2 and are therefore not repeated here).

[0176] Figure 10 shows an encoder 14, which may be a specific example of the encoder in Figure 1. Similar to Figure 2, encoder 14 may include a subtractor 22 configured to subtract a corresponding prediction signal 24 (e.g., block 18 with the reconstructed sample 104 acquired in step 812) from the inbound signal, i.e., picture 10, or block 18, by the current block 18, to obtain a prediction residual signal 26, which is then encoded into a data stream 12 by a prediction residual encoder 28. The prediction residual encoder 28 may include an irreversible coding stage 28a and a reversible coding stage (entropicorder) 28b. The irreversible coding stage 28a may include a quantizer 30 (not shown) configured to receive the prediction residual signal 26 and quantize the samples of the prediction residual signal 26. The obtained prediction residual signal 34 is then subjected to reversible coding by the reversible coding stage 28b, which is the entropi-encoded quantized prediction residual signal 34 of the entropicorder, to become a data stream 12. The encoder 14 further comprises a predictive residual signal reconstruction stage 36 connected to the output of the irreversible coding stage 28a, thereby enabling the reconstruction of the predictive residual signal from the converted and quantized predictive residual signal 34'.

[0177] The encoder 14 may include an adder 42 that adds the reconstructed predicted residual signal 34' output by the stage 36 to the predicted signal 24 (for example, including a block 18 having the reconstructed sample 104 obtained in step 813) in order to output a reconstructed signal, i.e., a reconstructed sample. This output is supplied to a predictor 44, which may determine the predicted signal 24 based on it (for example, by applying the techniques shown in Figures 8.1 to 7.4).

[0178] As seen in Figure 9, method steps 811, 812, and 813 are mapped here by stages 811', 812', and 813' within the predictor 44, respectively, and method steps 811, 812, and 813 can be implemented in hardware units and / or procedural routines within the predictor 44, collectively represented by 811', 812', and 813', or controlled by the predictor. The example shows that it is possible to skip the derivation stage 813', as in the example in Figure 7.1.

[0179] In particular, stages 811 and / or 813 may be shown as presenting registers such as register 910 for performing the shift operation described above (register 910 is not necessarily part of stage 811 or 813, and it may be a unit controlled by the stage in question). Alternatively, stage 812 may be shown as having or controlling a multiplier 1910, in which the multiplier selects adjacent samples 17 or the P of averaged samples 102. red Multiplication performed between individual elements is a matrix

number

[0180] The storage device 1044 here has an ALWIP matrix of 17M or

number

number

[0181] Even if not shown in the diagram, encoder 14 may determine the dimensions of the ALWIP matrix to be used (e.g., the set of sets S0, S1, S2) based on, for example, the dimensions of block 18. In some cases, it may not be necessary to notify the encoder of this selection as a result of the selection of the dimensions of block 18.

[0182] Therefore, the encoder 14 is further configured to insert the predicted residuals into the data stream 12 from which the predetermined block 18 can be reconstructed, using the predicted residuals 34 and predicted values ​​24(104) of the predetermined sample acquired in step 812.

[0183] Additionally or alternatively, the encoder 14 calculates the predicted residuals (26, 34) for a given block (18) by Q or Q. red For each of the given samples, the corresponding residual values ​​are inserted into a data stream (12), and as a result, the given block (18) is Q or Q red By correcting the predicted values ​​for each of the set of values, the predicted residuals (26,34) and predicted values ​​of a given sample are reconstructed, and as a result, the corresponding reconstructed values ​​are obtained in the reduced set of sample values ​​(102), except for the clipping that is optionally applied after prediction and / or correction. red It is strictly linearly dependent on the number of adjacent samples (102).

[0184] Additionally or alternatively, the encoder 14 may be configured to subdivide the picture (16) into multiple blocks of different block sizes, each comprising a predetermined block (18). The encoder 14 may select a linear or affine linear transformation (19, Ak) for a predetermined block (18) from a first set of linear or affine linear transformations, provided that the width W and height H of the predetermined block (18) are within a first set of width / height pairs (e.g., associated with S0), and a second set of linear or affine linear transformations, provided that the width W (also indicated as N) and height H (also indicated as M) of the predetermined block (18) are within a second set of width / height pairs (e.g., associated with S1) that do not belong to the first set of width / height pairs, such that the linear or affine linear transformation (19, Ak) selected for the predetermined block (18) is selected from a first set of linear or affine linear transformations, provided that the width W (also indicated as N) and height H (also indicated as M) of the predetermined block (18) are within a second set of width / height pairs (e.g., associated with S1). k ) may be configured to select.

[0185] Additionally or alternatively, the encoder may be configured such that a third set of one or more width / height pairs (e.g., S0) simply contains one width / height pair W',H', and each linear or affine linear transformation in the second set of linear or affine linear transformations is for converting N' sample values ​​to W'*H' predicted values ​​of the W'xH' array of sample positions.

[0186] Additionally or alternatively, the encoder may have a first and second set of width / height pairs, where each of the first width / height pairs is W p ,H p And, W p H p The first width / height vs. W is not equal to W. p ,H p And the second width / height ratio W q ,H q H q =W p and W q =H p The second width / height ratio is W q ,H q It can be configured to include the following.

[0187] Additionally or alternatively, the encoder may have a first and second set of width / height pairs, and a third width / height pair W p ,H p It may further include, W p is H p Equal to H p >H q It can be configured to be so.

[0188] Additionally or alternatively, the encoder may be configured to insert a set index into the data stream for a given block and select a linear or affine linear transformation from a given set of linear or affine linear transformations according to the set index.

[0189] Additionally or alternatively, the encoder may be configured such that a plurality of neighboring samples extend one-dimensionally along two sides of a given block, and the encoder groups a first subset of a plurality of neighboring samples adjacent to a first side of the given block into a first group (110) of one or more consecutive neighboring samples, and a second subset of a plurality of neighboring samples adjacent to a second side of the given block into a second group (110) of one or more consecutive neighboring samples, and performs a reduction by performing downsampling or averaging for each of the first and second groups of one or more neighboring samples having three or more neighboring samples, thereby obtaining a first sample value from the first group and a second sample value for the second group, and the encoder may be configured to select a linear or affine linear transformation from a given set of linear or affine linear transformations according to the set index, thereby obtaining two different states of the set index, which are linear or affine The system can be configured to select one of a linear or affine linear transformations from a predetermined set of linear transformations, subject the reduced set of sample values ​​to the predetermined linear or affine linear transformation, and generate an output vector of predicted values. In the case of a set index that assumes a first state of two different states in the form of a first vector, the predicted values ​​of the output vector are distributed to predetermined samples of a predetermined block along a first scan order. In the case of a set index that assumes a second state of two different states in the form of a second vector, the first and second vectors are different, and as a result, a component input by one of the first sample values ​​in the first vector is input by one of the second sample values ​​in the second vector, and a component input by one of the second sample values ​​in the first vector is input by one of the first sample values ​​in the second vector, and as a result, an output vector of predicted values ​​is generated, and the predicted values ​​of the output vector are distributed to predetermined samples of a predetermined block along a second scan order, transposed relative to the first scan order.

[0190] Additionally or alternatively, the encoder may be configured such that each linear or affine linear transformation in a first set of linear or affine linear transformations converts an N1 sample value to a w1*h1 predicted value of the w1xh1 array of sample positions, and each linear or affine linear transformation in a first set of linear or affine linear transformations converts an N2 sample value to a w2*h2 predicted value of the w2xh2 array of sample positions, and for a first predetermined set of width / height pairs, w1 exceeds the width of the first predetermined width / height pair or h1 exceeds the height of the first predetermined width / height pair, and for a second predetermined set of width / height pairs, w1 does not exceed the width of the second predetermined width / height pair or h1 exceeds the height of the second predetermined width / height pair, and the encoder may be configured to downsize multiple adjacent samples in order to obtain a reduced set (102) of sample values. The system may be configured to perform a reduction (100) by sampling or averaging, and as a result, if a given block is of a first predetermined width / height pair, and if a given block is of a second predetermined width / height pair, the reduced set of sample values ​​(102) has N1 sample values, and to perform a reduction set of sample values ​​to a selected linear or affine linear transformation by using only the first sub-part of the selected linear or affine linear transformation relating to the subsampling of the w1xh1 array of sample positions, along the width dimension if w1 exceeds the width of one width / height pair, or along the height dimension if h1 exceeds the height of one width / height pair, if a given block is of a second predetermined width / height pair.

[0191] Additionally or alternatively, the encoder may be configured such that each linear or affine linear transformation in a first set of linear or affine linear transformations is for converting N1 sample values ​​to w1*h1 predicted values ​​of a w1xh1 array of sample positions where w1=h1, and each linear or affine linear transformation in a first set of linear or affine linear transformations is for converting N2 sample values ​​to w2*h2 predicted values ​​of a w2xh2 array of sample positions where w2=h2. 7. Example in Figure 11 Figure 11 shows another example that can be interpreted from the examples in Figures 3 through 9 (in particular, some features may be directly derived from Figure 4 and are therefore not repeated here).

[0192] Figure 11 shows possible implementations of the decoder 54 in Figure 4, i.e., implementations that are compatible with the implementation of the encoder 14 in Figure 10. In particular, the adder 42' and predictor 44' may be connected to the prediction loop in the same way as the encoder 14 in Figure 10. The reconstructed, i.e., inversely quantized and retransformed prediction residual signal applied to the adder 42' may be derived by a sequence of entropy decoders that inverse the entropy coding of the entropy encoder, followed, as on the coding side, by a residual signal reconstruction stage consisting of an inverse quantizer and inverse converter 40'. The output of the decoder is the reconstruction of picture 10. The reconstruction of picture 10 may be available directly at the output of the adder 42' or at the output of the in-loop filter.

[0193] As can be seen from the diagram, the rows 813', 812', and 813' may also be encoders 14, and the storage unit 1044 may store a set of matrices similar to encoders 14. Therefore, the explanation will not be repeated here. The index 944 (e.g., one or more of the above indices such as i, k, transposed index, set index) can be obtained directly from the data stream 12. The selection between sets S0, S1, and S2 can follow size (e.g., H / K or M / N).

[0194] Additionally or alternatively, the decoder may be configured to derive a predicted residual (34'') from the data stream (12) for a given block (18), and to reconstruct the given block (18) (42') using the predicted residual (34'') and predicted value (24') for a given sample (24', 104, 108, 108').

[0195] Additionally or alternatively, the decoder may, for a given block (18), have Q or Q red To obtain the corresponding residual values ​​for each of a given set of samples, the predicted residuals (34'') are derived from the data stream (12), and the corresponding reconstructed values ​​(10) are obtained for Q samples or Q red By correcting the predicted value of each of a given set of samples by the corresponding residual value (34''), a given block (18) is reconstructed using the predicted residual (34'') and predicted value (24'', 104) of a given sample (118'', 118''), and as a result the corresponding reconstructed value (10) is obtained, with the exception of clipping applied after prediction and / or correction, which reduces the P values ​​in the set of samples. red It is strictly linearly dependent on the number of adjacent samples (102). It can be configured in this way.

[0196] Additionally or alternatively, the decoder is configured to subdivide the picture (10) into multiple blocks of different block sizes, including a predetermined block (18), and the decoder performs a linear or affine linear transformation (19, 17M, A) depending on the width W and height H of the predetermined block (18). k The linear or affine linear transformation selected for a given block (18) may be selected from a first set of linear or affine linear transformations, as long as the width W and height H of the given block (81) are within a first set of width / height pairs, and selected from a second set of linear or affine linear transformations, as long as the width W and height H of the given block are within a second set of width / height pairs that are separate from the first set of width / height pairs.

[0197] Additionally or alternatively, the decoder is configured to subdivide the picture (10) into multiple blocks of different block sizes, including a predetermined block (18), and the decoder performs a linear or affine linear transformation (19, 17M, A) depending on the width W and height H of the predetermined block (18). k The system can be configured such that, as a result of selecting a linear or affine linear transformation for a given block (18), it is selected from a first set of linear or affine linear transformations, as long as the width W and height H of the given block (18) are in a first set of width / height pairs; it is selected from a second set of linear or affine linear transformations, as long as the width W and height H of the given block (18) are in a second set of width / height pairs separated from the first set of width / height pairs; and it is selected from a third set of linear or affine linear transformations, as long as the width W and height H of the given block (18) are in a third set of width / height pairs separated from the first and second sets of width / height pairs.

[0198] Additionally or alternatively, the decoder may be configured such that a third set of one or more width / height pairs simply contains one width / height pair W',H', and each linear or affine linear transformation in the first set of linear or affine linear transformations is for converting N' sample values ​​to W'*H' predicted values ​​of the W'xH' array of sample positions.

[0199] Additionally or alternatively, the decoder may have a first and second set of width / height pairs, where the first width / height pair is W p ,H p And, W p H p The first width / height vs. W is not equal to W. p ,H p And the second width / height ratio W q ,H q H q =W p and W q =H p The second width / height ratio is W q ,H qIt can be configured to include the following.

[0200] Additionally or alternatively, the decoder may have a third width / height pair W for each of the first and second sets of width / height pairs. p ,H p It may further include, W p is H p Equal to H p >H q That is the case. Additionally or alternatively, the decoder reads a set index (k) from the data stream (12) for a given block (18). It can be configured to select a linear or affine linear transformation from a predetermined set of linear or affine linear transformations according to a set index (k).

[0201] Additionally or alternatively, the decoder may be configured such that a plurality of neighboring samples (17) extend one-dimensionally along two sides of a given block (18), and the decoder groups a first subset of a plurality of neighboring samples adjacent to a first side of the given block into a first group (110) of one or more consecutive neighboring samples, and a second subset of a plurality of neighboring samples adjacent to a second side of the given block into a second group (110) of one or more consecutive neighboring samples, and performs a reduction (811) by performing downsampling or averaging for each of the first and second groups of one or more neighboring samples having three or more neighboring samples, thereby obtaining a first sample value from the first group and a second sample value for the second group, and the decoder may be configured to select a linear or affine linear transformation from a given set of linear or affine linear transformations according to the set index, thereby obtaining two different states of the set index, which are linear or affine linear. Alternatively, the system may be configured to select one of a predetermined set of linear or affine linear transformations, subject the reduced set of sample values ​​to the predetermined linear or affine linear transformation, and to generate an output vector of predicted values, in the case of a set index that assumes a first state of two different states in the form of a first vector, the predicted values ​​of the output vector are distributed to predetermined samples of a predetermined block along a first scan order, and in the case of a set index that assumes a second state of two different states in the form of a second vector, the first vector and the second vector are different, and as a result a component input by one of the first sample values ​​in the first vector is input by one of the second sample values ​​in the second vector, and a component input by one of the second sample values ​​in the first vector is input by one of the first sample values ​​in the second vector, and as a result an output vector of predicted values ​​is generated, and the predicted values ​​of the output vector are distributed to predetermined samples of a predetermined block along a second scan order, transposed relative to the first scan order.

[0202] Additionally or alternatively, the decoder may be configured such that each linear or affine linear transformation in a first set of linear or affine linear transformations is for converting N1 sample values ​​to w1*h1 predicted values ​​of the w1xh1 array of sample positions, and each linear or affine linear transformation in a first set of linear or affine linear transformations is for converting N2 sample values ​​to w2*h2 predicted values ​​of the w2xh2 array of sample positions, and for a first predetermined set of width / height pairs, w1 exceeds the width of the first predetermined width / height pair or h1 exceeds the height of the first predetermined width / height pair, and for a second predetermined set of width / height pairs, w1 does not exceed the width of the second predetermined width / height pair or h1 exceeds the height of the second predetermined width / height pair, and the decoder may be configured to downgrade multiple adjacent samples to obtain a reduced set (102) of sample values. The system is configured to perform a reduction (100) by sampling or averaging, and as a result, if a given block is of a first predetermined width / height pair, and if a given block is of a second predetermined width / height pair, the reduced set of sample values ​​(102) has N1 sample values, and to perform a reduction set of sample values ​​to a selected linear or affine linear transformation by using only the first sub-part of the selected linear or affine linear transformation related to subsampling the w1xh1 array of sample positions, along the width dimension if w1 exceeds the width of one width / height pair, or along the height dimension if h1 exceeds the height of one width / height pair, if a given block is of a second predetermined width / height pair.

[0203] Additionally or alternatively, the decoder may be configured such that each linear or affine linear transformation in a first set of linear or affine linear transformations is for converting N1 sample values ​​to w1*h1 predicted values ​​of a w1xh1 array of sample positions where w1=h1, and each linear or affine linear transformation in a first set of linear or affine linear transformations is for converting N2 sample values ​​to w2*h2 predicted values ​​of a w2xh2 array of sample positions where w2=h2. 8. Consideration of the effects of this technology

[0204] It should also be noted that, regardless of the use of operations such as bit shifts for averaging and / or interpolation (which, among other things, results in the effect of reducing computational effort), other effects that may outweigh the effective use of bit shifts may be obtained in some cases.

[0205] In particular, in this embodiment, the prediction mode can be shared across different block shapes so that the selection of the ALWIP matrix 17M (e.g., in step 812a) is performed for a limited number of sets. For example, there may be a set of ALWIP matrices that is smaller than the possible dimensions (e.g., height / width pairs) of the block 18 to be predicted. As described above, one can refer to Figure 12, which maps different width / height pairs of the predicted block 18 to one of the sets S0 (e.g., n0 matrices, e.g., n0=16), S1 (e.g., n1 matrices, e.g., n1=8), and S2 (e.g., n2 matrices, e.g., n2=6) (different subdivisions may be possible).

[0206] For example, the 16x8 matrix of set S1 has the dimension The prediction modes of blocks having any of the following dimensions may be shared: 4×8, 4×16, 4×32, 4×64, 8×4, 8×8, 16×4, 32×4, and 64×4. The 64×8 matrix of set S2 may be shared by the prediction modes of blocks having any of the following dimensions: 8×16, 8×32, 8×64, 16×8, 16×16, 16×32, 16×64, 32×8, 32×16, 32×32, 32×64, 64×8, 64×16, 64×32, and 64×6. P required to form set 102 red All that is required is to perform the technique described in the reduction step 811 (see above) to reduce the dimension of the boundary 17 to a sample of numbers, but in step 812, the original dimension of the predicted block 18 is irrelevant. In step 813 (if implemented), it is possible to obtain a complete prediction of the block by simply performing interpolation.

[0207] It should be noted that this method makes it possible to reduce the memory space required in memory space 1044 with an unexpected dimension of 16*16*4 + 8*16*8 + 6*64*8 = 5120 values ​​(for example, each value being, for example, an 8-bit value).

[0208] In comparison, conventional techniques require a set of matrices for each width / height pair. As can be easily seen from Figure 12, 25 sets are needed! It is easy to see that 25 sets of matrices require far more memory space than 5120 values. Therefore, to reduce the required memory space, the number of matrices in each set needs to be reduced, but if there are fewer matrices available for prediction, the quality will suffer!

[0209] The reduction in memory space due to shared technology is further amplified by the reduction in the size of the stored matrices themselves. For example, a prediction of MxN = 64x64 blocks would require a matrix of size QxP = (M*N)x(M+N), i.e., (64*64)*(64+64) = 524,288 values ​​to be stored in memory space! Therefore, this technology can save even more memory space than expected. Therefore, this technology makes it possible to reduce the number of parameters that need to be stored in unit 1044.

[0210] Regardless of whether bit shifts are actually used, it can reduce the amount of memory resources available to the encoder or decoder, or conversely, allow more prediction modes to be used for parity in the memory space.

[0211] Nevertheless, the optimal effect is achieved by combining the bit shift technique (in steps 811 and / or 813) with one that shares the same prediction mode for multiple modes (in step 812).

[0212] Compared to the conventional method which uses 25 different sets for 25 different height / width pairs, this technique could be interpreted as obviously increasing complexity (since steps 811 and / or 813 are not conceivable in the conventional technique). However, the introduction of steps 811 and / or 813 can be better compensated for by the reduction in multiplication.

[0213] Furthermore, with respect to the conventional method using 25 different sets for 25 different height / width pairs, the instructions required to control this process require more memory space (because additional instructions for steps 811 and / or 813 are stored). However, the need to store instructions for steps 811 and / or 813 can be further compensated by the reduction in space implied by the reduction in the number of matrices to be stored. 9. Further Embodiments and Examples

[0214] Generally, the embodiments may be implemented as a computer program product having program instructions that operate to execute one of the methods when the computer program product runs on a computer. The program instructions can be stored, for example, on a machine-readable medium. Other embodiments include a computer program for performing one of the methods described herein, which is stored on a machine-readable carrier.

[0215] In other words, an embodiment of the method of the present invention is a computer program having program instructions for performing one of the methods of the present invention when the computer program is executed on a computer.

[0216] Accordingly, further embodiments of the methods of the present invention include a computer program for performing one of the methods described herein, and a data carrier medium (or digital storage medium or computer-readable medium) recorded thereon. The data carrier medium, digital storage medium, or recording medium is tangible and / or non-temporary, rather than intangible and transient signals.

[0217] Therefore, a further embodiment of the method of the present invention is a data stream or sequence of signals representing a computer program for performing one of the methods described herein. The data stream or sequence of signals may be transmitted, for example, over a data communication connection, such as the Internet. Further embodiments include processing means, such as a computer or programmable logic device, that perform one of the methods described herein. Further embodiments include a computer on which a computer program for performing one of the methods described herein is installed.

[0218] Further embodiments of the present invention include an apparatus or system for transferring (e.g., electronically or optically) a computer program for performing one of the methods described herein to a receiver. The receiver may be, for example, a computer, a mobile device, a memory device, etc. The apparatus or system may include, for example, a file server for transferring the computer program to the receiver.

[0219] In some examples, a programmable logic device (e.g., a field-programmable gate array) can be used to perform some or all of the functions of the methods described herein. In some embodiments, a field-programmable gate array may work with a microprocessing unit to perform one of the methods described herein. In general, these methods can be performed by any suitable hardware device.

[0220] The examples described above are merely illustrative of the principles described above. Modifications and variations of the configurations and details described herein will be obvious. Therefore, it is intended that these are not limited by the imminent claims, but not by the specific details shown in the descriptions and explanations of the embodiments herein.

[0221] Elements that are equal or equivalent, or elements that have equal or equivalent functions, are indicated in the following description by equal or equivalent reference numbers, even if they occur in different diagrams. References

[0222] [1] P. Helle et al., “Non-linear weighted intra prediction”, JVET-L0199, Macao, China, October 2018.

[0223] [2] F. Bossen, J. Boyce, K. Suehring, X. Li, V. Seregin, “JVET common test conditions and software reference configurations for SDR video”, JVET-K1010, Ljubljana, SI, July 2018.

Claims

1. A device for predicting at least a portion of a picture, The aforementioned device is Downsampling is a method of downsampling a set of sample values ​​adjacent to a block of a picture, wherein the picture includes a plurality of blocks. Determining a matrix of weight coefficients based at least partially on the height and width of the block in the picture, The process involves generating multiple predicted values ​​based at least partially on the set of downsampled sample values, It is configured to perform actions including, Determining the weight coefficient matrix involves selecting the weight coefficient matrix based in part on whether the height and width of the block of the picture are in a first set of height and width combinations or in a second set of height and width combinations separate from the first set of height and width combinations. The above generation is characterized by including the application of the weight coefficient matrix, Device.

2. The aforementioned operation is, Further predicted sample values ​​of the block are derived by upsampling the aforementioned multiple predicted values. Further including, The apparatus according to claim 1.

3. The apparatus according to claim 1, wherein the set of sample values ​​adjacent to the block of the picture extends one-dimensionally along the top of the block and extends one-dimensionally along the left side of the block.

4. The device includes a decoder configured to decode data corresponding to the picture from a data stream, or The device includes an encoder configured to encode data corresponding to the picture into a data stream. The apparatus according to claim 1.

5. A method for predicting at least a portion of a picture, The aforementioned method, Downsampling is a method of downsampling a set of sample values ​​adjacent to a block of a picture, wherein the picture includes a plurality of blocks. Determining a matrix of weight coefficients based at least partially on the height and width of the block in the picture, This includes generating a set of predicted values ​​based at least partially on the downsampled set of sample values, Determining the weight coefficient matrix involves selecting the weight coefficient matrix based in part on whether the height and width of the block of the picture are in a first set of height and width combinations or in a second set of height and width combinations separate from the first set of height and width combinations. The above generation is characterized by including the application of the weight coefficient matrix, method.

6. Further predicted sample values ​​of the block are derived by upsampling the aforementioned multiple predicted values. The method according to claim 5, further comprising:

7. The method according to claim 5, wherein the set of sample values ​​adjacent to the block of the picture extends one-dimensionally along the top of the block and extends one-dimensionally along the left side of the block.

8. Decoding data corresponding to the picture from the data stream, or encoding data corresponding to the picture into the data stream. The method according to claim 5, further comprising:

9. When executed by at least one processor, The aforementioned at least one processor, Downsampling is a method of downsampling a set of sample values ​​adjacent to a block of a picture, wherein the picture includes a plurality of blocks. Determining a matrix of weight coefficients based at least partially on the height and width of the block in the picture, The process involves generating multiple predicted values ​​based at least partially on the set of downsampled sample values, Includes instructions that cause an action to be performed, Determining the weight coefficient matrix involves selecting the weight coefficient matrix based in part on whether the height and width of the block of the picture are in a first set of height and width combinations or in a second set of height and width combinations separate from the first set of height and width combinations. The above generation is characterized by including the application of the weight coefficient matrix, Non-temporary computer-readable media.

10. The aforementioned operation is, Further predicted sample values ​​of the block are derived by upsampling the aforementioned multiple predicted values. The non-temporary computer-readable medium according to claim 9, further comprising:

11. The non-temporary computer-readable medium according to claim 9, wherein the set of sample values ​​adjacent to the block of the picture extends one-dimensionally along the top of the block and extends one-dimensionally along the left side of the block.

12. The aforementioned operation is, Decoding data corresponding to the picture from the data stream, or encoding data corresponding to the picture into the data stream. The non-temporary computer-readable medium according to claim 9, further comprising: