Concept of intra-prediction mode for image encoding in block unit
A neural network-based method for resampling intra prediction templates addresses inefficiencies in video codecs by optimizing intra prediction modes, enhancing compression efficiency and reducing side information in block-based image coding.
Patent Information
- Application Number
- JP2025068135
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2018-03-29
- Filing Date
- 2025-04-17
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2039-03-28
AI Technical Summary
Existing video codecs face challenges in efficiently determining intra prediction modes for block-based image coding, leading to increased side information rates and suboptimal compression efficiency due to the trade-off between the number of supported modes and prediction accuracy.
Implementing a neural network-based approach to determine intra prediction signals by resampling templates of adjacent samples to match the current block size, utilizing a neural network to enhance prediction accuracy and reduce side information requirements.
Improves compression efficiency by optimizing intra prediction modes, reducing the need for side information and enhancing prediction accuracy, thereby minimizing bit usage in data streams.
Smart Images

Figure 2025106572000001_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the concept of an improved intra prediction mode for block-based image coding that can be used in video codecs such as HEVC or successors to HEVC.
Background Art
[0002] Intra prediction modes are widely used in the coding of images and videos. In video coding, intra prediction modes compete with other prediction modes such as inter prediction modes like motion compensation prediction mode. In the intra prediction mode, the current block is predicted based on adjacent samples, i.e., samples that have already been coded on the encoder side and decoded on the decoder side as far as the decoder side is concerned. The adjacent sample values are extrapolated to the current block to form a prediction signal for the current block, and the prediction residual is transmitted in the data stream of the current block. The better the prediction signal, the smaller the prediction residual, and thus the fewer the number of bits required to code the prediction residual.
[0003] In order to be effective, several aspects need to be considered to form an effective framework for intra prediction in a block-based image coding environment. For example, the larger the number of intra prediction modes supported by the codec, the greater the consumption of the side information rate for notifying the decoder of the selection. On the other hand, the set of supported intra prediction modes needs to be able to provide a good prediction signal, i.e., a prediction signal with a low prediction residual.
Summary of the Invention
Problems to be Solved by the Invention
[0004] The present application aims to provide a concept of an intra prediction mode that enables more efficient compression of a block-based image codec when using the concept of an improved intra prediction mode.
Means for Solving the Problems
[0005] This object is achieved by the subject matter of the independent claims of this application.
[0006] An apparatus (e.g., a decoder) for decoding an image from a data stream in block units, the apparatus supporting at least one intra prediction mode, wherein an intra prediction signal of a block of a predetermined size of the image is determined by applying a first template of samples adjacent to the current block to a neural network, and for a current block of a size different from the predetermined size, resampling a second template of samples adjacent to the current block to match the first template to obtain a resampled template; applying the resampled template of the samples to a neural network to obtain a preliminary intra prediction; An apparatus is disclosed that is configured to resample the preliminary intra prediction signal to match the current block to obtain the intra prediction signal of the current block.
[0007] An apparatus (e.g., an encoder) for encoding an image in block units into a data stream, the apparatus supporting at least one intra prediction mode, wherein an intra prediction signal of a block of a predetermined size of the image is determined by applying a first template of samples adjacent to the current block to a neural network, and for a current block of a size different from the predetermined size, resampling a second template of samples adjacent to the current block to match the first template to obtain a resampled template; applying the resampled template of the samples to a neural network to obtain a preliminary intra prediction; An apparatus is also disclosed that is configured to resample a preliminary intra prediction signal to match a current block in order to obtain an intra prediction signal for the current block.
[0008] The apparatus can be configured to resample by downsampling a second template to obtain a first template.
[0009] The apparatus can be configured to resample the preliminary intra prediction signal by upsampling the preliminary intra prediction signal.
[0010] The apparatus can be configured to convert the preliminary intra prediction signal from the spatial domain to a transform domain and resample the preliminary intra prediction signal in the transform domain.
[0011] The apparatus can be configured to resample the transform domain preliminary intra prediction signal by scaling the coefficients of the preliminary intra prediction signal.
[0012] The apparatus increases the dimension of the intra prediction signal to match the dimension of the current block, and zero-pads the coefficients of the additional coefficients of the preliminary intra prediction signal that are related to higher frequency bins to resample the transform domain preliminary intra prediction signal.
[0013] The apparatus can be configured to construct the transform domain preliminary intra prediction signal with an inverse quantized version of the prediction residual signal.
[0014] The apparatus can be configured to resample the preliminary intra prediction signal in the spatial domain.
[0015] The apparatus can be configured to resample a preliminary intra prediction signal by performing bilinear interpolation.
[0016] The apparatus can be configured to encode information regarding resampling and / or the use of neural networks of different dimensions in a data field.
[0017] An apparatus (e.g., a decoder) for decoding an image from a data stream in block units, supporting at least one intra prediction mode in which an intra prediction signal of a current block of an image is determined by applying a first set of neighboring samples of the current block to a neural network to obtain a prediction of a set of transform coefficients of the transform of the current block.
[0018] An apparatus (e.g., an encoder) for encoding an image in block units in a data stream, supporting at least one intra prediction mode in which an intra prediction signal of a current block of an image is determined by applying a first set of neighboring samples of the current block to a neural network to obtain a prediction of a set of transform coefficients of the transform of the current block.
[0019] One of the apparatuses can be configured to inverse transform a prediction to obtain a reconstructed signal.
[0020] One of the apparatuses can be configured to decode an index from a data stream using a variable length code and perform a selection using the index.
[0021] One of the apparatuses can be configured to determine a ranking of a set of intra prediction modes and then resample a second template.
[0022] Resample a second template of samples adjacent to the current block to conform to the first template and obtain a resampled template. Apply the resampled template of the samples to a neural network to obtain a preliminary intra prediction signal. Resample the preliminary intra prediction signal to match the current block and obtain the intra prediction signal of the current block. A method comprising the above is disclosed.
[0023] A method for decoding an image from a data stream in blocks, A method is disclosed that includes applying a first set of samples adjacent to the current block to a neural network to obtain a prediction of a set of transform coefficients for the transform of the current block.
[0024] A method for encoding an image in a data stream in blocks, A method is disclosed that includes applying a first set of samples adjacent to the current block to a neural network to obtain a prediction of a set of transform coefficients for the transform of the current block.
[0025] The above and / or following methods can use a device comprising at least one of the above and / or following devices.
[0026] A computer-readable storage medium is also disclosed that, when executed by a computer, includes instructions to cause the computer to execute the above and / or following methods and / or implement the above and / or following in at least one component of a device.
[0027] A data stream obtained by the above and / or following methods and / or by the above and / or following devices is also disclosed.
[0028] As far as the design of the neural network described above is concerned, this application provides many examples for appropriately determining its parameters.
[0029] Advantageous implementations of this application are the subject of the dependent claims. Preferred examples of this application are described below with reference to the figures.
Brief Description of the Drawings
[0030]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7a
Figure 7b
Figure 7c
Figure 7d
Figure 8
Figure 9a
Figure 9b
Figure 10
Figure 11-1
Figure 11-2
Figure 12
Figure 13a
Figure 13b
Mode for Carrying Out the Invention
[0031] Hereinafter, various examples that help achieve more effective compression when using intra prediction will be described. Some examples achieve an improvement in compression efficiency by using a series of intra prediction modes based on neural networks. The latter can be added to, or alternatively provided exclusively in addition to, other intra prediction modes designed heuristically, for example. Other examples use neural networks to perform a selection from among multiple intra prediction modes. And even other examples utilize both aspects of the art described herein.
[0032] To facilitate understanding of the following examples of the present application, the description begins with the presentation of possible encoders and decoders that are compatible with and can construct the examples outlined later in the present application. FIG. 1 shows an apparatus for encoding an image 10 into a data stream 12 in block units. The apparatus is denoted using reference numeral 14 and can be a still image encoder or a video encoder. In other words, the image 10 can be the current image from the video 16 if the encoder 14 is configured to encode the video 16 including the image 10 into the data stream 12, or if the encoder 14 can exclusively encode the image 10 into the data stream 12.
[0033] As described above, encoder 14 performs encoding in a block-by-block manner or on a block basis. For this reason, encoder 14 subdivides image 10 into blocks, and the units of encoder 14 encode image 10 into data stream 12. Examples of possible subdivisions of image 10 into blocks 18 are shown in more detail below. In general, the subdivision can start with a hierarchical multi-tree subdivision into blocks 18 of a fixed size, such as an array of blocks arranged in rows and columns, or into an array of tree blocks from the entire image area of image 10 or from a pre-partition of image 10, and can end with blocks 18 of different block sizes, such as by using a multi-tree subdivision. These examples should not be treated as excluding other possible ways of subdividing image 10 into blocks 18.
[0034] Furthermore, encoder 14 is a predictive encoder configured to predictively encode image 10 into data stream 12. For a particular block 18, this means that encoder 14 determines a prediction signal for block 18 and encodes the prediction residual, i.e., the prediction error by which the prediction signal deviates from the actual image content within block 18, into data stream 12.
[0035] Encoder 14 can support different prediction modes to derive the prediction signal for a specific block 18. The prediction mode that is important in the following example is the intra prediction mode in which the inside of block 18 is spatially predicted from samples of an adjacent, already encoded image 10. The encoding of the data stream 12 of image 10 and thus the corresponding decoding procedure can be based on a specific encoding order 20 defined among blocks 18. For example, the encoding order 20 can traverse block 18 in a raster scan order such as row by row from top to bottom while traversing each row from left to right. In the case of hierarchical multi-tree based subdivision, the raster scan order can be applied within each hierarchical level and a depth-first traversal order can be applied. That is, the leaf notes within a block of a specific hierarchical level precede blocks of the same hierarchical level having the same parent block according to the encoding order 20. Depending on the encoding order 20, adjacent, already encoded samples of block 18 can usually be arranged on one or more sides of block 18. In the case of the example presented herein, for example, adjacent, already encoded samples of block 18 are arranged above and to the left of block 18.
[0036] Encoder 14 may support not only the intra prediction mode. For example, if encoder 14 is a video encoder, encoder 14 can also support an intra prediction mode in which block 18 is temporarily predicted from an image of a previously encoded video 16. Such an intra prediction mode can be a motion compensated prediction mode in which a motion vector is signaled for such a block 18 indicating the relative spatial offset of the portion where the prediction signal of block 18 is derived as a copy. Additionally or alternatively, other non-intra prediction modes such as an inter-view prediction mode when encoder 14 is a multi-view encoder, or a non-prediction mode in which the inside of block 18 is encoded as it is, i.e., without prediction, can also be made available.
[0037] Before beginning to focus the description of the present application on the intra prediction mode, a more specific example of a possible block-based encoder will be described, namely a possible implementation of encoder 14 such as presenting two corresponding examples of decoders that conform to FIGS. 1 and 2, respectively, as described with respect to FIG. 2.
[0038] Figure 2 shows a possible implementation of the encoder 14 of FIG. 1, i.e., an encoder configured to use transform coding to encode prediction residuals, which is merely an example, and the present application is not limited to that kind of prediction residual coding. According to FIG. 2, the encoder 14 is configured to subtract the corresponding prediction signal 24 from the inbound signal, i.e., the image 10, or the current block 18 on a block-by-block basis, to obtain a prediction residual signal 26 that is later encoded into the data stream 12 by the prediction residual encoder 28, and includes a subtractor 22. The prediction residual encoder 28 is composed of an irreversible coding stage 28a and a reversible coding stage 28b. The irreversible stage 28a receives the prediction residual signal 26 and includes a quantizer 30 that quantizes the samples of the prediction residual signal 26. As already described above, in this example, transform coding of the prediction residual signal 26 is used, and thus the irreversible coding stage 28a includes a transform stage 32 connected between the subtractor 22 and the quantizer 30 to transform such a prediction residual 26 that has been spectrally decomposed by the quantization of the quantizer 30 that presents the residual signal 26 with the transformed coefficients. The transform can be a DCT, DST, FFT, Hadamard transform, etc. Next, the transformed and quantized prediction residual signal 34 undergoes reversible coding by the reversible coding stage 28b, which is an entropy encoder that entropy-encodes the quantized prediction residual signal 34 into the data stream 12. The encoder 14 further includes a prediction residual signal reconstruction stage 36 connected to the output of the quantizer 30 to reconstruct the prediction residual signal from the transformed and quantized prediction residual signal 34 in a manner available to the decoder as well. That is, it is the quantizer 30 that takes into account the coding loss. For this purpose, the prediction residual reconstruction stage 36 includes an inverse quantizer 38 that performs the inverse of the quantization of the quantizer 30, followed by an inverse transformer 40 that performs an inverse transformation with respect to the transformation performed by a transformer 32 such as the inverse of the spectral decomposition of any of the specific transform examples described above. The encoder 14 includes an adder 42 that adds the reconstructed prediction residual signal output by the inverse transformer 40 and the prediction signal 24 to output the reconstructed signal, i.e., the reconstructed samples.This output is supplied to the predictor 44 of the encoder 14, and the encoder 14 determines the prediction signal 24 based thereon. It is the predictor 44 that supports all the prediction modes already described above with respect to FIG. 1. FIG. 2 also shows that when the encoder 14 is a video encoder, the encoder 14 can also include a loop filter 46 that filters a fully reconstructed image that forms the reference image of the predictor 44 with respect to the inter prediction block after filtering.
[0039] As already described above, the encoder 14 operates on a block basis. In the following description, the target block basis is obtained by finely dividing the image 10 into blocks, and for each of these blocks, an intra prediction mode is selected from a set or plurality of intra prediction modes respectively supported by the predictor 44 or the encoder 14, and the selected intra prediction mode is executed individually. However, there may also be other types of blocks into which the image 10 is finely divided in the same way. For example, the above determination as to whether the image 10 is inter-coded or intra-coded can be made at a granularity or in units of blocks deviating from the block 18. For example, the between-mode / within-mode determination can be executed at the level of the coding blocks into which the image 10 is finely divided and each coding block is further finely divided into prediction blocks. Prediction blocks having coding blocks for which it has been determined that intra prediction is to be used are each finely divided for intra prediction mode determination. On the other hand, for each of these prediction blocks, it is determined which of the supported intra prediction modes is to be used for each prediction block. These prediction blocks form the block 18 of interest here. Prediction blocks within the coding blocks related to inter prediction will be treated differently by the predictor 44. They will be inter-predicted from the reference image by determining a motion vector and copying the prediction signal of this block from the position in the reference image indicated by the motion vector. Another block subdivision relates to the subdivision into transform blocks in the unit in which the conversion by the transformer 32 and the inverse transformer 40 is executed. The transformed block can be, for example, the result of further subdividing the coding block. Of course, the examples described here should not be treated as limiting, and there are other examples. For the sake of completeness, it should be noted that the subdivision into coding blocks can use, for example, multi-tree subdivision, and similarly, the prediction blocks and / or the transform blocks can be obtained by further subdividing the coding blocks using multi-tree subdivision.
[0040] A decoder or apparatus for block-based decoding adapted to the encoder 14 of FIG. 1 is shown in FIG. 3. This decoder 54 performs the reverse of what encoder 14 does. That is, it decodes the image 10 from the data stream 12 in block units and, for this purpose, supports a plurality of intra prediction modes. The decoder 54 can include, for example, a residual provider 156. All other possibilities described above with respect to FIG. 1 are also valid for the decoder 54. For this reason, the decoder 54 can be a still image decoder or a video decoder, and all prediction modes and predictabilities are also supported by the decoder 54. The difference between the encoder 14 and the decoder 54 is mainly in the fact that the encoder 14 selects or chooses encoding decisions according to some optimization, such as minimizing some cost functions that can depend, for example, on the encoding speed and / or the encoding distortion. One of these encoding options or encoding parameters can include the selection of the intra prediction mode to be used for the current block 18 from among the available or supported intra prediction modes. Next, the selected intra prediction mode is signaled by the encoder 14 for the current block 18 in the data stream 12, and the decoder 54 uses this signaling of the data stream 12 of block 18 to re-do the selection. Similarly, the sub-division of the image 10 into blocks 18 can be subject to optimization within the encoder 14, and the corresponding sub-division information can be transmitted within the data stream 12, and the decoder 54 restores the sub-division of the image 10 into blocks 18 based on the sub-division information. Summarizing the above, the decoder 54 can be a prediction decoder operating on a block basis, and in addition to the intra prediction mode, the decoder 54 can support other prediction modes, such as the inter prediction mode, for example, if the decoder 54 is a video decoder. In decoding, the decoder 54 can also use the encoding order 20 described with respect to FIG. 1, and since this encoding order 20 is followed by both the encoder 14 and the decoder 54, the same adjacent samples are available for the current block 18 in both the encoder 14 and the decoder 54.Therefore, in order to avoid unnecessary repetitions, the description of the operation mode of the encoder 14 must also apply to the decoder 54, for example, as far as prediction is concerned and as far as the coding of the prediction residue is concerned, as far as the subdivision of the image 10 into blocks is concerned. The difference lies in the fact that the encoder 14, by optimization, selects or inserts some coding options or coding parameters and signals within the data stream 12, and these are derived from the data stream 12 by the decoder 54 in order to redo the prediction, such as subdivision.
[0041] Figure 4 shows a possible implementation of the decoder 54 of FIG. 3, i.e., an implementation that conforms to the implementation of the encoder 14 of FIG. 1 as shown in FIG. 2. Since many elements of the encoder 54 of FIG. 4 are the same as those occurring in the corresponding encoder of FIG. 2, the same reference signs with apostrophes are used in FIG. 4 to denote these elements. In particular, the adder 42’, the optional in-loop filter 46’ and the predictor 44’ are connected to the prediction loop in the same way as they are in the encoder of FIG. 2. The reconfigured, i.e., inverse quantized and inverse transformed, prediction residue signal applied to the added 42’ is derived by a sequence of an entropy decoder 56 that reverses the entropy coding of the entropy encoder 28b, followed by a residue signal reconstruction stage 36’ composed of an inverse quantizer 38’ and an inverse transformer 40’ in the same way as on the coding side. The output of the decoder is the reconstruction of the image 10. The reconstruction of the image 10 can be available directly at the output of the adder 42’ or, alternatively, at the output of the in-loop filter 46’. In order to improve the image quality, some post-filters can be arranged at the output of the decoder in order to subject the reconstruction of the image 10 to some post-filtering, but this option is not shown in FIG. 4.
[0042] Repeating, with respect to FIG. 4, the description given above with respect to FIG. 2 is valid for FIG. 4 except that the encoder only makes relevant decisions regarding the optimization task and the encoding options. However, all descriptions regarding block subdivision, prediction, inverse quantization, and re-conversion are also valid for the decoder 54 of FIG. 4.
[0043] Before proceeding with the description of possible examples of the present application, some notes regarding the above examples must be made. Although not explicitly mentioned above, it is clear that block 18 can have any shape. It can be, for example, rectangular or quadratic. Further, the above description of the operating modes of encoder 14 and decoder 54 often refers to the "current block" 18, but it is clear that encoder 14 and decoder 54 act accordingly for each block for which the intra prediction mode is selected. As described above, there may be other blocks, but the following description focuses on block 18 where image 10 is subdivided and the intra prediction mode is selected.
[0044] To summarize the situation of a particular block 18 for which the intra prediction mode is selected, refer to FIG. 5. FIG. 5 shows the current block 18, i.e., the block that is currently being encoded or decoded. FIG. 5 shows a set 60 of adjacent samples 62, i.e., samples 62 that have spatially adjacent blocks 18. Samples 64 within block 18 are the prediction targets. Thus, the derived prediction signal is a prediction for each sample 64 within block 18. As already described above, a plurality of 66 prediction modes are available for each block 18, and when block 18 is intra predicted, this plurality of 66 modes simply includes mutual prediction modes. A selection 68 is performed on the encoder side and the decoder side to determine one of the intra prediction modes from the plurality of 66 that are used to predict (71) the prediction signal of block 18 based on the adjacent sample set 60. The examples described further below differ with respect to the operation modes regarding the available intra prediction modes 66 and the selection 68, e.g., whether side information is set in the data stream 12 regarding the selection 68 for block 18. However, the description of these examples starts with a specific description that provides mathematical details. According to this first example, the selection of a particular block 18 that is intra predicted is associated with the corresponding side information signaling 70 and the data stream, and the plurality of 66 intra prediction modes includes a set 72 of neural network-based intra prediction modes and a further set 74 of intra prediction modes of heuristic design. One of the intra prediction modes of set 74 can be, for example, an average value determined based on the adjacent sample set 60, and this average value can be the DC prediction mode assigned to all samples 64 within block 18. Additionally or alternatively, set 74 can include a mutual prediction mode that can be called an angular mutual prediction mode in which the sample values of the adjacent sample set 60 are copied to block 18 along a particular prediction in-direction in which such angular intra prediction modes differ from each other.FIG. 5 shows that data stream 12 includes a portion 76 in which a prediction residual that can include transform coding with quantization in a transform domain as needed for encoding, in addition to side information 70 that exists as needed for a selected 68 out of a plurality of 66 intra prediction modes, is encoded.
[0045] In particular, for ease of understanding the following description of a particular example of the present application, FIG. 6 shows a general operation mode of an intra prediction block in an encoder and a decoder. FIG. 6 shows a block 18 and adjacent samples 60 set based on the execution of intra prediction. Note that this set 60 can vary among intra prediction modes of a plurality of 66 intra prediction modes with respect to cardinality. That is, the number of samples in set 60 is actually used according to each intra prediction mode for determining the prediction signal of block 18. However, this is for ease of understanding and is not shown in FIG. 6. FIG. 6 shows that the encoder and decoder have one neural network 800 to 80 KB -1 for each of the neural network-based intra prediction modes in set 72. Set 60 is applied to each neural network to derive the corresponding intra prediction mode between sets 72. In addition to this, FIG. 6 fairly typically shows one block 82 provided based on one or more prediction signals of one or more intra prediction modes in set 74, such as the input, that is, set 60 of adjacent samples, for example, a DC mode prediction signal and / or an angular intra prediction mode prediction signal. The following description is for i = 0 ··· K B -1 having neural network 80 ishows how the parameters of can be advantageously determined. The specific examples shown below also provide another neural network 84 dedicated to providing probability values for each neural network-based intra prediction mode within set 72 to the encoder and decoder, based on a set 86 of adjacent samples that may or may not match set 60. Thus, the probability values are provided when neural network 84 helps to more effectively render side information 70 for mode selection. For example, in the example described below, a variable length code is used to indicate one of the intra prediction modes, and at least as far as set 72 is concerned, the probability values provided by neural network 84 are used as indices into an ordered list of intra prediction modes ordered according to the probability values output by neural network 84 for the neural network-based intra prediction modes within set 72, thereby optimizing or reducing the code rate of side information 70. For this reason, as shown in FIG. 6, mode selection 68 is effectively performed in response to both the probability values provided by additional neural network 84 and side information 70 within data stream 12. 1. An algorithm for training the parameters of a neural network that performs intra prediction A block of a video frame, i.e., block 18
Number
Number
Number
Number
[0046] Next to be described is an algorithm for designing the intra prediction function of several [Number] that may occur in a typical hybrid video coding standard, i.e., set 72, via a data-driven optimization approach. To achieve that goal, the following main design features were considered.
[0047] 1. In the optimization algorithm we implement, we want to use an appropriate approximation of the cost function that includes the number of bits that can be expected to be spent, especially for notifying the prediction residual.
[0048] 2. We want to jointly train several intra predictions to be able to handle various signal characteristics.
[0049] 3. When training the intra prediction, it is necessary to consider the number of bits required to notify which intra mode to use.
[0050] 4. Retain the set of intra-predictions already defined, e.g., HEVC intra-predictions, and train our predictions as complementary predictions.
[0051] 5. Typical hybrid video coding standards usually support several block shapes that can partition a particular block
Number
[0052] In the next four sections, we can explain how to address each of these requirements. More precisely, in Section 1.1, we explain how to handle the first item. In Section 1.2, we explain how to handle items 2 through 3. In Section 1.4, we explain how to take item 4 into account. Finally, in Section 1.5, we explain how to handle the last item. 1.1 Algorithm for Training a Loss Function to Approximate the Rate Function of a Video Codec Data-driven approaches for determining unknown parameters used in a video codec are typically set as optimization algorithms that attempt to minimize a pre-defined loss function over a particular set of training examples. Usually, for a numerical optimization algorithm to actually function, the latter loss function needs to satisfy some smoothness requirements.
[0053] On the other hand, a video encoder such as HEVC exhibits the best performance when making decisions to minimize the rate-distortion cost
Number
Number
Number
Number
[0054] True function
Number
Number
Number
[0055] More precisely, as before,
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
[0056] Function
Number
Number
Number
Number
Number
[0057] weight
Number
Number
Number
Number
Number
Number
[0058] In that task, usually, (stochastic) gradient descent is used. 1.2 Training of Prediction with Fixed Block Shape In this section, a specific block
Number
Number
[0059] Assume that a pre-defined "architecture" of our prediction is given. As a result, several fixed
Number
Number
Number
Number
Number
Number
[0060] The following section will explain this in detail. The function in (2) defines the neural network 800 - 80 in Figure 6 KB -1.
[0061] Next, we model the signaling cost of the intramode to be designed by using the second parameter-dependent function
Number
[0062] for
Number
Number
Number
[0063] Similarly, an example using the function of (4) representing the neural network 84 in FIG. 6 is shown in Section 1.3.
[0064] Function
Number
[0065] This function defines, for example, the VLC code length distribution used for the side information 70, that is, the code length associated with the side information 70 having more sets 72 of cad ponites.
[0066] Next,
Number
Number
[0067] For the time being,
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Mathematics
Mathematics
[0068] In the gradient update, the ratio of the residuals and the
Mathematics
Mathematics
Mathematics
Mathematics
[0069] When the loss function in (5) is given, data-driven optimization is used to
Mathematics
Mathematics
Mathematics
Mathematics
Mathematics
Number
Number
Number
Number
Number
[0070]
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
[0071] Nonlinear activation function
Number
Number
[0072] Here,
Number
Number
Number
[0073] function [Number] appears as follows here. Fixed [Number] In the case of [Number] as [Number] is assumed to be given.
[0074] Here, [Number] is the same as that in (1). Next, [Number] having [Number] For [Number] define as follows.
[0075] Therefore, [Number] parameterized neural network 80 using i is described. This is [Number] a sequence of , and in this example, it is applied alternately within the sequence, [Number] includes. [Number] In the sequence of , [Number] is, for example, [Number] the number of preceding nodes before this neuron layer j in the feed - forward direction of the neural network determined by the dimension m of [Number] the number of columns of and the number of rows of it, [Number] represents a neuron layer such as the j - th layer having the number of neurons of neuron layer j itself determined by the dimension n of . [Number] Each row contains weights that control the strength of the transfer of the activation of each signal intensity of the m previous neurons to each neuron in the neuron layer j corresponding to that row.
Number
Number
Number
[0076] Similarly, the function
Number
Number
Number
Number
[0077] Here,
Number
Number
Number
Number
Number
Number
Number
[0078] That is,
Number
Number
Number
Number
Number
[0079] Next, extend the loss function from (5) to the loss function
Number
[0080] For a large set of training examples, maintain the notation from the end of the previous section and
Number
Number
[0081] To do so, usually first find the weights by optimization (6), then initialize with those weights to find the weights (10) to optimize. 1.5 Co - training of predictions for several block shapes In this section, in the training of predictions, in a general video coding standard, it is considered how to divide blocks into small sub - blocks in various ways and perform intra - prediction on the small sub - blocks.
[0082] That is, some
Number
Number
Number
Number
Number
[0083]
Number
Number
Number
Number
Number
Number
Number
Number
[0084] For the given color component,
Number
Number
Number
Number
[0085] While maintaining the notation of Section 1.2,
Number
Number
Number
Number
Number
Number
Number
[0086] Furthermore,
Number
Number
Number
[0087] Similar to Section 1.4,
Number
Number
Number
[0088] Next,
Number
Number
Number
Number
Number
Number
[0089] Next,
Number
[0090] Next,
Number
[0091] Finally,
Number
Number
Number
Number
[0092] Typically, first
Number
Number
Mathematics
Mathematics
Mathematics
Mathematics
Mathematics
Mathematics
Mathematics
Mathematics
Mathematics
Mathematics
Mathematics
[0093] 2. The decoder reconstructs a flag that is part of the side information 70 from the bitstream and indicates whether any of the following options are true: [label=)
[0094] (i)
Number
[0095] (ii)
Number
Number
[0096] 3. If option 2 in step 2 is true, the decoder proceeds to the specified block 10 as in the case of the underlying hybrid video coding standard.
[0097] 4. If option 1 in step 2 is true, the decoder applies to the
Number
Number
Number
[0098] (i) The decoder [Number] by [Number] defines, and the latter [Number] is used with the underlying standard from data stream 12 and [Number] the index, which is also part of side information 70, through an entropy encoding engine that defines [Number] is analyzed.
[0099] (ii) The decoder [Number] inductively permutes by placing [Number] Here, [Number] for and [Number] is the smallest number having [Number] where
Number
[0100] Next, the decoder extracts from the bit stream 12 a unique
Number
[0101] the latter index
Number
Number
Number
Number
[0102] 5. If option 1 of step 2 is true and the decoder determines the index
Number
Number
Number
[0103] Integration into an existing hybrid video codec of an intra prediction function designed based on a data-driven learning approach. The description had two main parts. In the first part, a specific algorithm for the offline training of the intra prediction function was described. In the second part, how the video decoder uses the latter prediction function to generate a prediction signal for a specific block was described.
[0104] Therefore, what was described in Sections 1.1 to 2 above is, among other things, a device for decoding the image 10 from the data stream 12 in block units. The device 54 supports a plurality of intra prediction modes including a set 72 of intra prediction modes in which the intra prediction signal of the current block 18 of the image 10 is determined by applying a first set 60 of adjacent samples of the current block 18 to the neural network 80 i The device 54 selects (68) one intra prediction mode for the current block 18 from the plurality of intra prediction modes 66 and is configured to predict (71) the current block 18 using one intra prediction mode, i.e., using the corresponding neural network 80 that is selected m The decoder presented in Section 2 had an intra prediction mode 74 within a plurality of the supported intra prediction modes 66 in addition to those based on the neural network of the set 72, but this is merely an example and need not be so. Further, the above description of Sections 1 and 2 may be modified in that the decoder 54 does not use and does not include a further neural network 84. Regarding the above optimization, this is the finding
Number
[0105] A further alternative to the description presented in section 2 above is that decoder 54 selects the intra prediction mode finally used from an ordered list of intra prediction modes according to a second part other than the first part of the data stream, and derives a ranking among a set 72 of neural network-based intra prediction modes according to the first part of the data stream related to the neighborhood of the current block 18 in order to obtain an ordered list of intra prediction modes. The "first part" can be related to, for example, coded parameters or prediction parameters related to one or more blocks adjacent to the current block 18. And the "second part" can be, for example, an index indicating the set 72 of neural network-based intra prediction modes or an index thereof. When interpreted in line with section 2 outlined above, decoder 54 sorts these probability values to determine the rank of each intra prediction mode in set 72, and thereby, for each intra prediction mode in set 72 of intra prediction modes, applies a set 86 of adjacent samples thereto to determine probability values by means of a further neural network 84 for obtaining an ordered list of intra prediction modes. Next, an index within data stream 12 as part of side information 70 is used as an index to the ordered list. Here, this index is M BIt can be encoded using a variable - length code that indicates the code length. And, as described above in Section 2, in item 4i, according to a further alternative, the decoder 54 can use the above - mentioned probability values determined by a further neural network 84 for each neural - network - based intra - prediction mode of the set 72 in order to efficiently perform entropy coding of the index to the set 72. In particular, the symbol alphabet of this index, which is part of the side information 70 and is used as an index to the set 72, includes the symbols or values of each mode within the set 72, and the probability values provided by the neural network 84, in the case of the design of the neural network 84 according to the above description, provide probability values that lead to efficient entropy coding in that these probability values accurately represent the actual symbol statistics. For this entropy coding, for example, arithmetic coding or probability interval partitioning entropy (PIPE) coding can be used.
[0106] Advantageously, no additional information is required for any of the intra - prediction modes of the set 72. Each neural network 80 iWhen, for example, it is advantageously parameterized for the encoder and decoder according to the above description of Sections 1 and 2, the prediction signal of the current block 18 is derived from the data stream without additional guidance. As already shown above, the presence of other intra prediction modes other than the neural network-based mode of set 72 is optional. They are shown above by set 74. In this regard, it should be noted that one possible way to select set 60, i.e., the set of adjacent samples that form the input of prediction within 71, can be such that this set 60 is the same for the intra prediction mode of set 74, i.e., heuristic. The set 60 of the neural network-based intra prediction mode is larger in terms of the number of adjacent samples included in set 60 and affecting intra prediction 71. In other words, the cardinality of set 60 can be larger for the neural network-based intra prediction mode 72 compared to other modes of set 74. For example, the set 60 of any intra prediction mode of set 74 can simply include adjacent samples along a one-dimensional line extending along the side of block 18 such as the left one and the upper one. The set 60 of the neural network-based intra prediction mode extends along the just-mentioned side of block 18, but can cover an L-shaped portion wider than one sample width like the set 60 of the intra prediction mode of set 74. The L-shaped portion can extend further beyond the just-mentioned side of block 18. In this way, the neural network-based intra prediction mode can result in better intra prediction with correspondingly low prediction residuals.
[0107] As described in section 2 above, the side information 70 transmitted to the intra prediction block 18 in the data stream 12 can include a flag generally indicating whether the intra prediction mode selected for the block 18 is a member of set 72 or a member of set 74. However, this flag is merely an option accompanied by side information 70 indicating an index to the entire plurality of intra prediction modes 66 including both sets 72 and 74, for example.
[0108] The following alternative, just described, will be briefly described with respect to FIGS. 7a through 7d. The figures define both the decoder and the encoder simultaneously, that is, from the perspective of their functions with respect to the intra prediction block 18. The difference between the encoder operation mode and the decoder operation mode with respect to the intra coding block 18 is that, on the one hand, the encoder executes all or at least some of the available intra prediction modes 66, for example, determines the optimal one 90 from the perspective of a cost function that minimizes meaning, and the encoder forms the data stream 12, that is, the code enters the date there, and the fact that the decoder derives data therefrom by decoding and reading respectively. FIG. 7a shows whether the flag 70a in the side information 70 of block 18 is the best mode for block 18 as determined by the encoder in step 90, which is within set 72, that is, a neural network-based intra prediction mode, or within set 74, that is, one of the non-neural network-based intra prediction modes. The encoder inserts the flag 70a into the data stream 12 accordingly, while the decoder searches for the flag 70a therefrom. FIG. 7a assumes that the determined intra prediction mode 92 is within set 72. Next, a separate neural network 84 determines the probability values of each neural network-based intra prediction mode in set 72, and uses these probability value sets 72, or more precisely, the neural network-based intra prediction modes therein are ordered according to the probability values, such as in descending order of probability values, thereby resulting in an ordered list 94 of intra prediction modes. Next, an index 70b, which is part of the side information 70, is encoded into the data stream 12 by the encoder and decoded therefrom by the decoder. Thus, the decoder can determine which of sets 72 and 74 to use. The intra prediction mode used for block 18 is located such that if the intra prediction mode used is located in set 72, the ordering 96 of set 72 is performed.When the determined intra prediction mode is located in set 74, the index can also be transmitted in data stream 12. Thus, the decoder can generate a prediction signal for block 18 using the determined intra prediction mode by controlling selection 68 accordingly.
[0109] Figure 7b shows an alternative where flag 70a is not present in data stream 12. Instead, ordered list 94 will include not only the intra prediction modes of set 72 but also the intra prediction modes of set 74. The index within side information 70 is an index to this larger ordered list, indicating the determined intra prediction mode, i.e., that the one determined is the optimization 90. In the case of neural network 84 that provides probability values for neural network-based intra prediction modes only within 72, the ranking between the intra prediction modes of set 72 and those of set 74 for the intra prediction modes of set 74 can be determined by other means such as arranging the neural network-based intra prediction modes of set 72 to precede the modes of set 74 in ordered list 94 or arranging them alternately with each other. That is, the decoder can derive an index from data stream 12 and use index 70 as an index to ordered list 94 by deriving ordered list 94 from a plurality of intra prediction modes 66 using the probability values output by neural network 84. Figure 7c shows a further variation. Figure 7c shows the case where flag 70a is not used, but the flag can be used instead. The problem addressed by Figure 7c relates to the possibility that neither the encoder nor the decoder uses neural network 84. Rather, ordering 96 is derived by other means such as encoded parameters transmitted within data stream 12 with respect to one or more adjacent blocks 18, i.e., portion 98 of data stream 12 related to such one or more adjacent blocks.
[0110] FIG. 7d shows a further variation of FIG. 7a, namely, where index 70b is encoded using entropy encoding and decoded from data stream 12 using entropy decoding, generally denoted by reference numeral 100. The sample statistics or probability distribution used for entropy encoding 100 is controlled by the probability values output by neural network 84 as described above, which makes the entropy encoding of index 70b highly efficient.
[0111] For all of Examples 7a through 7d, it is a fact that the modes of set 74 may not exist. Thus, each module 82 may be missing, and flag 70a is unnecessary anyway.
[0112] Furthermore, although not shown in any figure, it is clear that mode selection 68 at the encoder and decoder can be synchronized with each other without explicit signaling 70, i.e., without consuming side information. Rather, the selection can be derived from other means such as necessarily taking the first one in ordered list 94 or deriving an index into ordered list 94 based on encoding parameters associated with one or more adjacent blocks. FIG. 8 shows an apparatus for designing a set of intra prediction modes of set 72 used for block-based image encoding. Apparatus 108 comprises a parameterizable version of neural network 800 from 80 KB-1 to 80, as well as a parameterizable network 109 that inherits or includes neural network 84. Here, in FIG. 8, as individual units, namely, from neural network 840 for providing probability values for neural network-based intra prediction mode 0, to neural network 84 B-1 for providing probability values associated within neural network-based intra prediction mode K KB-1 are shown. Parameters 111 for parameterizing neural network 84 and from neural network 800 to 80 KB-1Parameter 113 for parameterizing is input or applied to the respective parameter inputs of these neural networks by update data 110. Device 108 has access to a reservoir or a plurality of image test blocks 114 together with corresponding adjacent sample sets 116. Pairs of these blocks 114 and their associated adjacent sample sets 116 are used sequentially by device 108. In particular, the current image test block 114 is applied to the parameterizable neural network 109, and neural network 80 provides prediction signal 118 to each neural network-based intra-prediction mode of set 72, and each neural network 80 provides a probability value for each of these modes. For this purpose, these neural networks use the current parameters 111 and 113.
[0113] In the above description, rec has been used to indicate the image test block 114,
Number
Number
Number
[0114] Therefore, the update data 110 attempts to update the parameters 111 and 113 so as to reduce the encoding cost function. Next, these updated parameters 111 and 113 are used by the parameterizable neural network 109 to process the next image test blocks of the plurality of 112. As described above with respect to Section 1.5, there can be a mechanism that controls that mainly pairs of those image test blocks 114 and their associated adjacent sample sets 116 are applied to the recursive update process in which intra prediction is performed, preferably without block subdivision, thereby preventing the parameters 111 and 113 from being overly optimized based on image test blocks where encoding in units of those sub-blocks is more cost-effective anyway.
[0115] So far, the above examples have mainly related to the case where the encoder and decoder have a set of neural network-based intra prediction modes within the supported intra prediction mode 66. According to the examples described with respect to FIGS. 9a and 9b, this is not necessarily the case. FIG. 9a attempts to outline the operating modes of the encoder and decoder according to an example in which the description is provided in a way that focuses on the differences from the description presented above with respect to FIG. 7a. The multiple supported intra prediction modes 66 may or may not include neural network-based intra prediction modes, and may or may not include non-neural network-based intra prediction modes. Thus, for each of the supported modes 66 provided, module 170 of FIG. 9a, which is constituted by the encoder and decoder respectively, the corresponding prediction signal is not necessarily a neural network. As already shown above, such intra prediction modes can be neural network-based or heuristically motivated and can calculate prediction signals based on the DC intra prediction mode or the angular intra prediction mode or any other. Thus, these modules 170 can be represented as prediction signal computers. However, the encoder and decoder according to the example of FIG. 9a comprise a neural network 84. The neural network 84 can calculate probability values of the supported intra prediction modes 66 based on an adjacent sample set 86, and as a result, change the multiple intra prediction modes 66 into an ordered list 94. The index 70 in the data stream 12 of block 18 points to this ordered list 94. Thus, the neural network 84 helps to reduce the side information rate spent on signaling of the intra prediction mode.
[0116] Figure 9b shows an alternative to Figure 9a in that, instead of ordering, the entropy decoding / encoding 100 of index 70 controls its probability or its simple statistics, i.e., by controlling the entropy probability distribution of the entropy decoding / encoding in the encoder / decoder according to the probability values determined for the neural network 84 for each of the plurality 66 of modes.
[0117] Figure 10 shows an apparatus for designing or parameterizing the neural network 84. Thus, it is an apparatus 108 for designing a neural network for assisting in selecting from among the set 66 of intra prediction modes. Here, for each mode of the set 66, corresponding neural network blocks are integrated to form the neural network 84, and the parameterizable neural network 109 of the apparatus 108 is merely parameterizable with respect to these blocks. For each mode, there is also a prediction signal computer 170, but this does not necessarily need to be parameterizable according to Figure 10. Thus, the apparatus 108 of Figure 10 calculates a cost estimate value for each mode based on the prediction signal 118 calculated by the corresponding prediction signal computer 170 and, optionally, based on the corresponding probability value determined by the corresponding neural network block for this mode. Based on the resulting cost estimate value 124, the minimum cost selector 126 selects the mode with the minimum cost estimate value, and the update data 110 updates the parameters 111 of the neural 84.
[0118] Regarding the descriptions of FIGS. 7a through 7d and FIGS. 9a and 9b, note the following. A common feature of the examples of FIGS. 9a and 9b, also used by some of the examples of FIGS. 7a through 7d, was the probability value of the neural network values for improving or reducing the overhead associated with side information 70 for notifying the decoder of the mode determined on the encoder side in the optimization process 90. However, as shown above for the examples of FIGS. 7a through 7d, it should be clear that the examples of FIGS. 9a and 9b can be changed such that no side information 70 is expended in the data stream 12 regarding mode selection. Rather, the probability values output by the neural network 84 for each mode can be used to necessarily synchronize the mode selection between the encoder and the decoder. In that case, there would be no optimization decision 90 on the encoder side regarding mode selection. Rather, the modes used between the sets 66 would be determined in the same way on the encoder side and the decoder side. A similar statement would apply to the corresponding examples of FIGS. 7a through 7d if changed such that no secondary information 70 within the data stream 12 is used. However, returning to the examples of FIGS. 9a and 9b, it is interesting that the selection process 68 on the decoder side depends on the probability values output by the neural network in terms of changing the interpretation of the side information insofar as the ordering to probability values or the probability distribution estimation dependency regarding the encoder is concerned. The dependency on the probability values not only affects the encoding of the side information 70 into the data stream 12, which uses, for example, variable length coding of each index in an ordered list or entropy coding / decoding with probability distribution estimation according to the probability values of the neural network, but also the optimization step 90: here, the code rate for transmitting the side information 70 can be taken into account and thus affects the decision 90. Example of FIG. 11-1 FIG. 11-1 shows a possible implementation of encoder 14-1, i.e., an encoder configured to use transform coding to encode prediction residuals, which is merely an example, and the present application is not limited to that kind of prediction residual coding. According to FIG. 11-1, encoder 14-1 is configured to subtract a prediction signal 24-1 corresponding to an inbound signal, i.e., image 10, or a current block 18 on a block-by-block basis from the current block 18, to obtain a spatial domain prediction residual signal 26 that is later encoded into data stream 12 by prediction residual encoder 28. The prediction residual encoder 28 includes an irreversible coding stage 28a and a reversible coding stage 28b. The irreversible coding stage 28a includes a quantizer 30 that receives the prediction residual signal 26 and quantizes samples of the prediction residual signal 26. This example uses transform coding of the prediction residual signal 26, and thus, the irreversible coding stage 28a includes a transform stage 32 connected between the subtractor 22 and the quantizer 30 to transform such a prediction residual 27 that has been spectrally decomposed by quantization of the quantizer 30 performed on the transformed coefficients presenting the residual signal 26. The transform can be DCT, DST, FFT, Hadamard transform, etc. Next, the transformed and transform domain quantized prediction residual signal 34 undergoes reversible coding by a reversible coding stage 28b, which is an entropy encoder that entropy encodes the quantized prediction residual signal 34 into data stream 12.
[0119] Encoder 14-1 further includes a transform domain prediction residual signal reconstruction stage 36-1 connected to the transform domain output of quantizer 30 to reconstruct the prediction residual signal from the transformed and quantized prediction residual signal 34 (in the transform domain) in a manner available also to the decoder, i.e., taking into account the encoding loss of quantizer 30. For this purpose, the prediction residual reconstruction stage 36-1 includes an inverse quantizer 38-1 that performs the inverse of the quantization of quantizer 30 to obtain an inverse quantized version 39-1 of the prediction residual signal 34, followed by an inverse transformer 40-1 that performs an inverse transform with respect to the transform performed by a transformer 32 such as the inverse of a spectral decomposition such as any of the specific transforms described above. Downstream of the inverse transformer 40-1, there is a spatial domain output 60 that can include a template useful for obtaining the prediction signal 24-1. In particular, predictor 44-1 can provide a transform domain output 45-1, which, when inverse-transformed by inverse transformer 51-1, provides the prediction signal 24-1 in the spatial domain (the prediction signal 24-1 is subtracted from the inbound signal 10 to obtain the prediction residual 26 in the time domain). In the inter-frame mode, the in-loop filter 46-1 can filter the fully reconstructed image 60 and, after filtering, also form the reference image 47-1 of predictor 44-1 with respect to the inter-prediction block (thus, in these cases, an adder 57-1 input from elements 44-1 and 36-1 is required, but there is no need for an inverse transformer 51-1 to provide the prediction signal 24-1 to subtractor 22 as shown by the dashed line 53-1).
[0120] However, unlike the encoder 14 of FIG. 2, the encoder 14-1 (in the prediction residual reconstruction stage 36-1) includes a transform domain adder 42-1 disposed between the inverse quantizer 38-1 and the inverse transformer 40-1. The transform domain adder 42-1 uses the transform domain prediction signal 45-1 as provided by the transform predictor 44-1 to provide the inverse transformer 40-1 with the sum 43-1 (in the transform domain) of the inverse quantization version 39-1 of the prediction residual signal 34 (as provided by the inverse quantizer 38-1). The predictor 44-1 can obtain the output from the inverse transformer 40-1 as a feedback input.
[0121] Accordingly, the spatial domain prediction signal 24-1 is obtained from the transform domain prediction signal 45-1. Also, the transform domain predictor 44-1, which can operate with a neural network according to the above example, is input by a spatial domain signal but outputs a transform domain signal. Example of FIG. 11-2 FIG. 11-2 shows a possible implementation of decoder 54-2, i.e., one that is adapted to the implementation of encoder 14-1. Since many elements of encoder 54-2 are the same as those occurring in the corresponding encoder of FIG. 11-1, the same reference numerals with “-2” are used in FIG. 11-2 to denote these elements. In particular, adder 42-2, any in-loop filter 46-2, and predictor 44-2 are connected to the prediction loop in the same manner as the encoder of FIG. 11-1. The reconstructed, i.e., inverse quantized and inverse transformed, prediction residual signal 24-2 (e.g., 60) is derived by a residual signal reconstruction stage 36-2 that consists of a sequence of entropy decoder 56 that reverses the entropy coding of entropy encoder 28b, followed by inverse quantizer 38-2 and inverse transformer 40-2 in the same manner as on the encoding side. The output of the decoder is the reconstruction of image 10. To subject the reconstruction of image 10 to some post-filtering to improve the image quality, some post-filters 46-2 can be placed at the output of the decoder. Similarly, the description presented above with respect to FIG. 11-1 is valid for FIG. 11-2, except that the encoder only performs relevant decisions regarding optimization tasks and coding options. However, all explanations regarding block subdivision, prediction, inverse quantization, and inverse transformation are also valid for decoder 54 of FIG. 11-2. The reconstructed signal 24-2 is provided to predictor 44-2, and predictor 44-2 can operate with a neural network according to the examples of FIGS. 5-10. Predictor 44-2 can provide a transform domain prediction value 45-2.
[0122] Contrary to the example of FIG. 4, but similar to the example of FIG. 11-1, the inverse quantizer 38-2 provides an inverse quantized version 39-2 of the prediction residual signal 34 (within the transform domain) that is not directly provided to the inverse transform 40-2. Instead, the inverse quantized version 39-2 of the prediction residual signal 34 is input to the adder 42-2 and is composed of the transform domain prediction value 45-2. Thus, the transform domain reconstruction signal 43-2 is obtained, which, when inverse transformed by the inverse transform 40-2, becomes the reconstruction signal 24-2 in the spatial domain that is used to display the image 10. Example of FIG. 12 Here, refer to FIG. 12. Both the decoder and the encoder are considered simultaneously, that is, from the perspective of their functions regarding the intra prediction block 18. The difference between the encoder operation mode and the decoder operation mode regarding the intra coding block 18 is that, on the one hand, the encoder executes all or at least some of the available intra prediction modes 66, for example, determines the optimal one as 90 from the perspective of a cost function that minimizes the meaning, and the encoder forms the data stream 12, that is, the code fills in the date there, and the fact that the decoder derives data therefrom by decoding and reading respectively. FIG. 12 shows whether the flag 70a in the side information 70 of the block 18 is in the set 72, that is, the neural network-based intra prediction mode, or in the set 74, that is, one of the non-neural network-based intra prediction modes, which is the intra prediction mode determined by the encoder to be the best mode for the block 18 in step 90, showing the alternative operation mode outlined above. The encoder inserts the flag 70a into the data stream 12 accordingly, while the decoder searches for the flag 70a therefrom. FIG. 12 assumes that the determined intra prediction mode 92 is in the set 72. Next, a separate neural network 84 determines the probability values of each neural network-based intra prediction mode in the set 72, and uses these probability value sets 72, or more precisely, the neural network-based intra prediction modes therein are ordered according to the probability values such as in descending order of the probability values, thereby resulting in an ordered list 94 of intra prediction modes. Next, an index 70b, which is part of the side information 70, is encoded by the encoder into the data stream 12 and decoded by the decoder therefrom. Therefore, the decoder can determine which of the sets 72 and 74. The intra prediction mode used for the block 18 is located to execute the ordering 96 of the set 72 when the intra prediction mode used is located in the set 72. When the determined intra prediction mode is located in the set 74, the index can also be transmitted in the data stream 12.Accordingly, the decoder can generate a prediction signal for block 18 using the determined intra prediction mode by controlling selection 68 accordingly.
[0123] As can be seen from FIG. 12, the prediction residual signal 34 (in the transform domain) is encoded into the data stream 12. The inverse quantizers 38-1, 38-2 derive the inverse quantized prediction residual signals 39-1, 39-2 in the transform domain. From the predictors 44-1, 44-2, the transform domain prediction signals 45-1, 45-2 are obtained. Next, the adder 42-1 sums the values 39-1 and 45-1 with each other (or the adder 42-2 sums the values 39-2 and 45-2) to obtain the transform domain reconstruction signal 43-1 (or 43-2). Downstream of the inverse transformers 40-1, 40-2, the spatial domain prediction signals 24-1, 24-2 (e.g., template 60) are obtained and can be used to reconstruct block 18 (which can be, for example, displayed).
[0124] All of the variations of FIGS. 7b-7d can be used to embody the examples of FIGS. 11-1, 11-2, and 12. Discussion A method for generating an intra prediction signal via a neural network is defined, and how this method is included in a video or still image codec is described. In these examples, instead of predicting in the spatial domain, the predictors 44-1, 44-2 can predict in the transform domain of a predefined image transform that may already be available in the underlying codec such as, for example, the discrete cosine transform. Second, each intra prediction mode defined for an image on a block of a particular shape induces an intra prediction mode for an image on a larger block.
[0125] Let B be a block of M rows and N columns of pixels where the image im exists. An adjacent B of B (block 18) for which the already reconstructed image rec is available recAssume that (template 60 or 86) exists. Next, in the examples of FIGS. 5 to 10, a new intra prediction mode defined by a neural network is introduced. Each of these intra prediction modes uses the reconstructed sample rec(24-1, 24-2) to similarly generate a prediction signal pred(45-1, 45-2) that is an image of B rec as well.
[0126] Let T be the image transformation defined on the image of B rec (e.g., the prediction residual signal 34 output by element 30), and let S be the inverse transformation of T (e.g., 43-1 or 43-2). Next, the prediction signal pred(45-1, 45-2) is regarded as a prediction of T(im). This means that in the reconstruction stage, after calculating pred(45-1, 45-2), it is necessary to calculate the image S(pred)(24-1, 24-2) to obtain the actual prediction of the image im(10).
[0127] Note that the transformation T that operates has some energy compression characteristics for natural images. This is exploited as follows. For each of the intra modes defined by the neural network, according to a pre-defined rule, the value of pred(45-1, 45-2) at a specific position in the transform domain is set to zero regardless of the input rec(24-1, 24-2). This reduces the computational complexity for obtaining the prediction signal pred(45-1, 45-2) in the transform domain.
[0128] (Referring to FIGS. 5 to 10, assume that the transformation T(32) and the inverse transformation S(40) are used in the transform residual coding of the underlying codec. In the reconstruction signal (24, 24’) of B, the prediction residual res(34) is inverse-transformed by the inverse transformation S(40) to obtain S(res), and S(res) is added to the underlying prediction signal (24) to obtain the final reconstruction signal (24).) In contrast, FIGS. 11 and 12 refer to the following procedure: When the prediction signals pred(45-1, 45-2) are generated by the neural network intra prediction method as described above, the final reconstructed signals (24-1, 24-2) are obtained by inverse transformation (40-1, 40-2) of pred+res (where pred is 45-1 or 45-2, and res is 39-1 or 39-2), and their sum is 43-1 or 43-2, which is the transformed domain version of the final reconstructed signals 24-1, 24-2.
[0129] Finally, it should be noted that the above changes to the intra prediction performed by the neural network as described above are optional and not necessarily mutually related to each other. This means that for a specific transformation T(32) using the inverse transformation S(40-1, 40-2) and one of the intra prediction modes defined by the above neural network, whether the mode is regarded as a prediction into the transformed domain corresponding to T may be extracted from the bitstream or from predefined settings. FIGS. 13a and 13b Referring to FIGS. 13a and 13b, for example, strategies that can be applied to a spatial domain-based method (e.g., FIGS. 11a and 11b) and / or a transform domain-based method (e.g., FIGS. 1-4) are shown.
[0130] In some cases, a neural network adapted to a block of a specific size can be freely used (e.g., M×N, where M is the number of rows and N is the number of columns), but the actual blocks 18 of the image to be reconstructed have different sizes (e.g., M1×N1). It should be noted that it is possible to perform an operation that enables the use of a neural network adapted to a specific size (e.g., M×N) without the need to use an ad-hoc trained neural network.
[0131] In particular, apparatus 14 or 54 can be enabled to decode an image (e.g., 10) from a data stream (e.g., 12) in block units. Apparatus 14, 54 natively supports at least one intra prediction mode, according to which the intra prediction signal of a block (e.g., 136, 172) of a predetermined size (e.g., M×N) of an image is determined by applying a first template (e.g., 130, 170) of samples adjacent to the current block (e.g., 136, 176) on a neural network (e.g., 80). The apparatus can be configured as follows for a current block (e.g., 18) that is different from the predetermined size (e.g., M1×N1): - Resample a second template (e.g., 60) of samples adjacent to the current block (e.g., 18) to obtain a resampled template (e.g., 130, 170) that conforms to the first template (e.g., 130, 170), - Apply the resampled template (e.g., 130, 170) of samples on the neural network (e.g., 80) to obtain a preliminary intra prediction signal (e.g., 138), - Resample the preliminary intra prediction signal (138) (e.g., U, V, 182) to match the current block (18, B1) to obtain the intra prediction signal of the current block.
[0132] FIG. 13a shows an example in the spatial domain. The spatial domain block 18 (also shown as B1) can be an M1xN1 block in which the image im1 is reconstructed (even if the image im1 is not yet available at this time). Template B 1,rec (e.g., set 60) has the already reconstructed image rec1, where rec1 is adjacent to im1 (and it should be noted that B 1,rec is adjacent to B1). Block 18 and template 60 ("the second template") can form element 132.
[0133] Due to the dimensions of B1, there may be no neural network that can be freely used to reconstruct B1. However, if the neural network can be freely used with blocks of different dimensions (such as "the first template"), the following procedure can be executed.
[0134] The conversion operation (here shown as D or 134) can be applied to, for example, element 130. However, since B1 is still unknown, note that it is easily possible to apply the conversion D(130) only to B 1,rec . The conversion 130 can provide an element 136 formed from the converted (resampled) template 130 and block 138.
[0135] For example, an M1xN1 block B1(18) (with unknown coefficients) can theoretically be converted into an M×N block B(138) (with further unknown coefficients). However, since the coefficients of block B(138) are unknown, it is not necessary to actually perform the conversion.
[0136] Similarly, the conversion D(134) converts the template B 1,rec (60) into a different template B rec (130) with different dimensions. The template 130 has a vertical thickness L (i.e., L columns of the vertical part) and a horizontal thickness K (i.e., K rows of the horizontal part), and B rec = D(B 1,rec ) can be in an L shape. It can be understood that the template 130 can include the following: - A K×N block on B rec (130), - An M×L block on the left side of B rec (130), and, - A K×L block on the left side of the K×N block on B rec (130) and on the M×L block on the left side of B rec (130).
[0137] In some cases, the conversion operation D(134) can be a downsampling operation when M1>M and N1>N (especially when M is a multiple of M1 and N is a multiple of N1). For example, when M1 = 2M and N1 = 2N, the conversion operation D can be based on making some bins invisible in a chess-like manner (e.g., B 1,rec deleting the diagonal from 60 to obtain the value of B rec 130).
[0138] At this point, B rec (B rec = D(rec1)) is an image reconstructed in M×N. In path 138a, devices 14, 54 can use a neural network that is natively trained for MxN blocks (e.g., with predictors 44, 44') (e.g., by operating as shown in FIGS. 5 to 10). By applying the above path (138a), the image im1 of block B is obtained. (In some examples, path 138a does not use a neural network but uses other techniques known in the art).
[0139] At this point, the size of the image im1 of block B(138) is M×N, but the size of the displayed image needs to be M1×N1. However, note that it is simply possible to perform a conversion (e.g., U) 140 that converts the image im1 within block B(138) to M1xN1.
[0140] Note that if D executed at 134 is a downsampling operation, U at 140 may be an upsampling operation. Therefore, U(140) can be obtained by introducing coefficients into the M1xN1 block in addition to the coefficients of the M×N block 138 obtained by the operation 138a using a neural network.
[0141] For example, when M1 = 2M and N1 = 2N, it is easily possible to perform interpolation (e.g., bilinear interpolation) in order to approximate (``infer'') the coefficients of im1 discarded by transformation D. Thus, the M1xN1 image im1 is obtained as element 142 and can be used to display a block image as part of image 10.
[0142] In particular, it is also theoretically possible to obtain block 144, and nevertheless, it is the same as template 60 (except for the errors due to transformations D and U). Thus, advantageously, it is not necessary to transform B 1,rec in order to obtain a new version of rec B that can already be freely used as template 60.
[0143] The operations shown in FIG. 13a can be performed, for example, by predictor 44 or 44'. Thus, the M1xN1 image im1 (142) can be understood as the prediction signal 24 (FIG. 2) or 24' (FIG. 4) that is summed with the prediction residual signal output by the inverse transformer 40 or 40' in order to obtain the reconstructed signal.
[0144] FIG. 13b shows an example in the transform domain (e.g., in the examples of FIGS. 11-1, 11-2). Element 162 is represented as being formed by the spatial domain template 60 (already decoded) and the spatial domain block 18 (having unknown coefficients). Block 18 can have size M1xN1 and can have unknown coefficients, which should be determined, for example, by predictors 44-1 or 44-2.
[0145] While a neural network of size M×N that has been determined can be freely used, there may be no neural network that directly operates on the M1×N1 blocks within the transform domain.
[0146] However, it should be noted that in predictors 44-1 and 44-2, it is possible to obtain a spatial domain template 170 having different dimensions (e.g., reduced dimensions) using the transformation D(166) applied to the template 60 (the "second template"). The template 170 (the "first template") can have an L-shaped shape such as the shape of the template 130 (see above), for example.
[0147] At this point, in passage 170a, a neural network (e.g., 800-80 N ) can be applied according to any of the above examples (see FIGS. 5 to 10). Thus, at the end of passage 170a, the known coefficients of version 172 of block 18 can be obtained.
[0148] However, it should be noted that the dimensions MxN of 172 do not match the dimensions M1xN1 of block 18 that must be visualized. Thus, the transformation into the transform domain (e.g., at 180) can be manipulated. For example, an MxN transform domain block T(176) can be obtained. A technique called zero-padding can be used, for example, by introducing a value "0" corresponding to a frequency value associated with a frequency that does not exist in the M×N transform T(176) in order to increase the number of rows and columns to M1 and N1, respectively. Thus, a zero-padding region 178 can be used (which can have, for example, an L-shape). In particular, the zero-padding region 178 includes a plurality of bins (all zeros) that are inserted into block 176 to obtain block 182. This can be obtained by a transformation V from T (transformed from 172) to T1(182). The dimensions of T(176) do not match the dimensions of block 18, but the dimensions of T1(182) actually match the dimensions of block 18 due to the insertion of the zero-padding region 178. Furthermore, zero-padding is obtained by inserting higher frequency bins (having zero values), which results in a result similar to interpolation.
[0149] Therefore, in the adders 42-1 and 42-2, the conversion T1(182) which is a version of 45-1 and 45-2 can be added. Subsequently, the inverse transform T -1 can be executed to obtain the reconfigured value 60 in the spatial domain used for visualizing the image 10.
[0150] The encoder can encode information regarding resampling (and the use of a neural network for blocks of a size different from the size of block 18) into the data stream 12, and as a result, the decoder has that knowledge. Discussion Assume B1 (e.g., 18) is a block of M1 rows and N1 columns, with M1≧M and N1≧N. Let B1,rec be the neighborhood of B1 (e.g., the adjacent template 60), and assume the region B 1,rec which is regarded as a subset of B rec (e.g., 130). Let im1 (e.g., 138) be the image of B1, and rec1 (e.g., the coefficients of B 1,rec ) be the already reconstructed image of B 1,rec The above solution is based on a pre-defined downsampling operation D (e.g., 134, 166) that maps the image of B1,rec to the image of B1. For example, when M1 = 2M and N1 = 2N, if B rec is composed of K rows above B and L columns to the left of B, and the corner of size K×L in the upper left of B, and B1,rec is composed of 2K rows above B1 and 2L columns to the left of B, and the corner of size 2K×2L in the upper left of B1, then D can be an operation that performs a downsampling operation by a factor of 2 in each direction after applying a smoothing filter. Therefore, D(rec1) can be regarded as the image reconstructed at B rec Using the above neural network-based intra prediction mode, a prediction signal pred(45-1) which is the image at B can be formed from D(rec1).
[0151] Here, two cases are distinguished: First, as in FIGS. 2, 4, and 13a, assume that in B, the neural network-based intra prediction predicts in the sample (spatial) domain. Let U(140) be a fixed upsampling filter that maps the image of B (e.g., 138) to the image of B1 (e.g., 142). For example, when M1 = 2M and N1 = 2N, U can be a bilinear interpolation operation. Next, U(pred) can be formed to obtain an image on B1 (e.g., 45-1) that is regarded as the prediction signal of im1 (e.g., 10).
[0152] Second, as in FIGS. 11-1, 11-2, and 13b, assume that in B, the prediction signal pred (e.g., 45-2) should be regarded as the prediction signal in the transform domain regarding the image transform T on B using the inverse transform S. Let T1 be the image transform on B1 using the inverse transform S1. Assume that a predefined mapping V is given that maps the image from the transform domain of T to the transform domain of T1. For example, when T is the discrete cosine transform of an M×N block using the inverse transform S and T1 is the discrete cosine transform of M1×N1 using the inverse transform S1, the block of transform coefficients of B can be mapped to the block of transform coefficients of B1 by zero-padding and scaling (see, for example, 178). This means that when the position in the frequency space is larger than M or N in the horizontal response vertical direction, all transform coefficients of B1 are set to zero, and the appropriately scaled transform coefficients of B are copied to the remaining M*N transform coefficients of B1. Next, V(pred) can be formed to obtain the element in the transform domain of T1 that is regarded as the prediction signal of T1(im1). The signal V(pred) can be further processed as described above.
[0153] As described above with respect to FIGS. 1 to 10, a method of ranking some intra prediction modes at a particular block B by generating a conditional probability distribution between these modes using neural network-based operations, and whether this ranking can be used to notify which intra prediction mode to apply at the current block was also described. Using a downsampling operation (e.g., 166) in the input of a neural network that generates the latter ranking in the same way as the actual prediction mode produces a ranking for extending to a larger block B1 than just described for the prediction mode, and thus is used to notify which extension mode to use at block B1. Whether to generate a prediction signal using a neural network-based intra prediction mode from a smaller block B on a given block B1 can be predefined or signaled as side information of the underlying video codec. Other examples Generally speaking, a decoder as described above can comprise an encoder as described above, and / or vice versa. For example, encoder 14 can be decoder 54, or include decoder 54 (or vice versa). Encoder 14-1 can be decoder 54-2 (or vice versa), etc. Further, it can also be understood that encoder 14 or 14-1 itself includes a decoder since the quantized prediction residual signal 34 forms a stream that is decoded to obtain the prediction signal 24 or 24-1.
[0154] Although some aspects have been described in the context of an apparatus, it will be apparent that these aspects also represent corresponding method descriptions, and that a block or apparatus corresponds to a method step or a feature of a method step. Similarly, aspects described in the context of a method step also represent corresponding descriptions of corresponding blocks or items or functions of a corresponding apparatus. Some or all of the method steps can be performed by (or using) a hardware device such as, for example, a microprocessor, a programmable computer, or an electronic circuit. In some examples, one or more of the most important method steps can be performed by such a device.
[0155] The encoded data stream of the present invention can be stored on a digital storage medium or transmitted on a transmission medium such as a wireless transmission medium or a wired transmission medium such as the Internet.
[0156] Depending on specific implementation requirements, examples of the present invention can be implemented in hardware or software. The implementation can be performed using a digital storage medium such as a floppy disk, a DVD, a Blu-ray, a CD, a ROM, a PROM, an EPROM, an EEPROM, a flash memory, etc., which stores electronically readable control signals and cooperates (or can cooperate) with a programmable computer system such that each method is executed. Thus, the digital storage medium can be made computer-readable.
[0157] Some examples according to the present invention include a data carrier having electronically readable control signals that can cooperate with a programmable computer system such that one of the methods described herein is executed.
[0158] In general, an example of the present invention can be implemented as a computer program product comprising program code, which functions to execute one of the methods when the computer program product is run on a computer. The program code may be stored, for example, on a machine-readable carrier.
[0159] Another example includes a computer program stored on a machine-readable carrier for executing one of the methods described herein.
[0160] Accordingly, an example of the method of the present invention is a computer program having program code for executing one of the methods described herein when the computer program is run on a computer.
[0161] Accordingly, a further example of the method of the present invention is a data carrier (or digital storage medium, or computer-readable medium) having recorded thereon a computer program for executing one of the methods described herein. The data carrier, digital storage medium, or recorded medium is typically tangible and / or non-transitory.
[0162] Accordingly, a further example of the method of the present invention is a data stream or sequence of signals representing a computer program for executing one of the methods described herein. The data stream or sequence of signals may be configured to be transferred via a data communication connection such as the Internet.
[0163] A further example includes processing means, such as a computer, or a programmable logic device, configured or adapted to execute one of the methods described herein.
[0164] A further example includes a computer having installed thereon a computer program for executing one of the methods described herein.
[0165] Further examples according to the present invention include an apparatus or system configured to transfer (e.g., electronically or optically) a computer program for performing one of the methods described herein to a receiver. The receiver can be, for example, a computer, a mobile device, a memory device, etc. The apparatus or system can include, for example, a file server for transferring the computer program to the receiver.
[0166] In some examples, a programmable logic device (e.g., a field programmable gate array) can be used to perform some or all of the functions of the methods described herein. In some examples, a field programmable gate array can cooperate with a microprocessor to perform one of the methods described herein. Generally, the methods are preferably performed by any hardware device.
[0167] The apparatuses described herein can be implemented using a hardware device, or using a computer, or using a combination of a hardware device and a computer.
[0168] The apparatuses described herein, or any component of the apparatuses described herein, can be implemented at least partially in hardware and / or software.
[0169] The methods described herein can be performed using a hardware device, or using a computer, or using a combination of a hardware device and a computer.
[0170] The methods described herein, or any component of the apparatuses described herein, can be performed at least partially by hardware and / or software.
[0171] The above embodiments merely illustrate the principles of the present invention. It is understood that changes and modifications to the configurations and details described herein will be apparent to those skilled in the art. Therefore, it is intended to be limited only by the impending claims, rather than by the examples and specific details presented as explanations and descriptions herein.
Claims
Claim 1 An apparatus (54-2) for decoding an image (10) from a data stream (12) in block units, wherein an intra prediction signal of a block (136, 172) of a predetermined size of the image is determined by applying a first template (130, 170) of samples adjacent to the current block to a neural network (80), the apparatus supporting at least one intra prediction mode, for a current block (18) different from the predetermined size, to obtain a resampled template (130, 170), resampling (134, 166) a second template (60) of samples adjacent to the current block (18) so as to match the first template (130, 170), to obtain preliminary intra prediction signals (138, 172, 176), applying (138a, 170a, 44-1, 44-2) the resampled template (130, 170) of the samples to the neural network (80), an apparatus configured to resample (140, 180) the preliminary intra prediction signals (138, 172, 176) so as to match the current block (18) to obtain the intra prediction signal (142, 24-1, 24-2) of the current block (18).
Citation Information
Patent Citations
Video signal processing method and apparatus
JP2010538520A
Method and Apparatus for Intra Prediction for Rru
US20080130745A1
Method and an apparatus for processing a video signal
US20100239002A1
Image coding method, image decoding method, image coding device and image decoding device
WO2016199330A1
Coding device, decoding device, coding method, and decoding method
WO2018199051A1