Scalable video coding that uses coding based on subblocks of transformation coefficient blocks within the enhancement layer.

By subdividing enhancement layer blocks based on base layer signals and forming weighted averages, the method improves scalable video coding efficiency through optimized prediction and reduced signaling, addressing inefficiencies in existing techniques.

JP2026090372APending Publication Date: 2026-06-02DOLBY VIDEO COMPRESSION LLC

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
DOLBY VIDEO COMPRESSION LLC
Filing Date
2026-02-09
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing scalable video coding techniques lack efficiency in encoding enhancement layers, particularly in utilizing base layer information for improved prediction and subdivision of transformation coefficient blocks.

Method used

The method involves controlling the subdivision of transformation coefficient blocks in the enhancement layer based on base layer residual signals, forming a weighted average of inter-layer and enhancement layer prediction signals, and using base layer hints for more efficient encoding and motion compensation.

Benefits of technology

This approach enhances coding efficiency by optimizing prediction signals and reducing signaling overhead, leading to improved compression ratios and encoding performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026090372000001_ABST
    Figure 2026090372000001_ABST
Patent Text Reader

Abstract

The present invention provides a scalable video encoding method and encoder, as well as a scalable video decoding method and decoder, that achieve higher encoding efficiency. [Solution] The method is designed to more efficiently encode 414 based on subblocks 412 of the conversion coefficient block 402 of the enhancement layer signal 400, by controlling the subdivision of each subblock 412 of the conversion coefficient block 402 based on the base layer residual signal or base layer signal 200. In particular, by utilizing each base layer hint, the subblocks become longer along the spatial frequency axis horizontal to the edge extension observable from the base layer residual signal or base layer signal.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This invention relates to scalable video coding. [Background technology]

[0002] In non-scalable coding, intra coding refers to a coding technique that utilizes only the data of the already coded portion of the current image (e.g., reconstruction samples, coding mode, or symbol statistics), rather than the reference data of the already coded image. For example, intra-coded images (intra images) are used in broadcast bitstreams to synchronize the decoder with the bitstream at so-called random access points. Intra images are also used to limit error propagation in error-prone environments. Generally, since the image used as a reference image is not available here, the first image in a coded video sequence must be coded as an intra image. Intra images are also used in scene cuts where time prediction typically cannot provide a suitable predictive signal.

[0003] Furthermore, intra-encoding modes are used for specific regions / blocks within so-called inter-images, where they perform better than inter-encoding modes in terms of ratio distortion efficiency. This is a common case in flat regions as well as regions where time prediction is considerably inadequate (occlusion, partial dissolve, or fading objects).

[0004] In scalable coding, the concept of intra coding (coding of intra images and intra blocks within inter-images) is extended to all images belonging to the same access unit or time instant. Therefore, the intra coding mode for spatial or quality enhancement layers increases coding efficiency while simultaneously enabling the use of inter-layer predictions from underlying images instantaneously. This means that not only are already coded portions of the current enhancement layer image available for intra-prediction, but already coded underlying images are also available instantaneously. The latter concept is also referred to as inter-layer intra-prediction.

[0005] In state-of-the-art hybrid video coding standards (such as H.264 / AVC or HEVC), the image in a video sequence is divided into blocks of samples. The size of the blocks is fixed, or the coding technique provides a hierarchical structure that allows blocks to be subdivided into even smaller blocks. Typically, the reconstruction of a block is obtained by generating a prediction signal for the block and adding the transmitted residual signal. Typically, the residual signal is transmitted using transform coding. This means that quantization indices for the transform coefficients (also called transform coefficient levels) are transmitted using entropy coding techniques. Then, on the decoder side, these transmitted transform coefficient levels are scaled and inversely transformed to obtain the residual signal that is added to the prediction signal. The residual signal is generated by intra-prediction (using only data already transmitted for the current moment of time) or by inter-prediction (using data already transmitted for different moments of time).

[0006] If inter prediction is used, the prediction block is obtained by motion-compensated prediction using samples from already reconstructed frames. This can be done by uni-directional prediction (using one reference picture and a set of motion parameters). Alternatively, the prediction signal can be generated by multi-hypothesis prediction. In the latter case, two or more prediction signals are superimposed. That is, for each sample, a weighted average is constructed to form the final prediction signal. The multiple prediction signals (that are superimposed) are generated using different motion parameters for different hypotheses (e.g., different reference pictures or motion vectors). Also, for uni-directional prediction, it is possible to multiply the samples of the motion-compensated prediction signal by a constant coefficient and add a constant offset to form the final prediction signal. Also, such scaling and offset correction are used for all hypotheses or for selected hypotheses in multi-hypothesis prediction.

[0007] In current state-of-the-art video coding techniques, the intra prediction signal for a block is obtained by predicting samples from the spatial neighborhood of the current block (which are the blocks reconstructed before the current block according to the order of blocks being processed). In the latest standards, various prediction methods for performing prediction in the spatial domain are utilized. Precise granular directional prediction modes, with or without filtering the samples of the neighboring blocks, are extended to specific angles to generate the prediction signal. Further, there are plane-based and DC-based prediction modes that use the samples of neighboring blocks to generate a flat prediction plane or a DC prediction block.

[0008] In older video coding standards (e.g., H.263, MPEG-4), intra prediction was performed within the transform domain. In this case, the transfer coefficients were inverse quantized. And for a subset of the transform coefficients, the transform coefficient values were predicted using the corresponding reconstructed transform coefficients of adjacent blocks. The inverse quantized transform coefficients were added to the predicted transform coefficient values, and the reconstructed transform coefficients were used as input to the inverse transform. The output of the inverse transform formed the final reconstructed signal for the block.

[0009] In scalable video coding as well, base layer information is utilized to support prediction processing for the enhancement layer. In the state-of-the-art video coding standard for scalable coding, the SVC extension of H.264 / AVC, there is one additional mode for improving the coding efficiency of intra prediction processing within the enhancement layer. This mode is indicated at the macroblock level (a block of 16×16 luma samples). This mode is supported only when the co-located samples in the lower layer are encoded using an intra prediction mode. If this mode is selected for a macroblock in the quality enhancement layer, the prediction signal is assembled by the co-located samples of the reconstructed lower layer signal before the non-blocking filter operation. If the inter-layer intra prediction mode is selected within the spatial enhancement layer, the prediction signal is generated by upsampling the co-located reconstructed base layer signal (after the non-blocking filter operation). An FIR filter is used for upsampling. Generally, for the inter-layer intra prediction mode, an additional residual signal is transmitted by transform coding. Also, if it is correspondingly indicated in the bitstream, the transmission of the residual signal can be omitted (inferred to be equal to zero). The final reconstructed signal is obtained by adding the reconstructed residual signal (obtained by scaling the transmitted transform coefficient levels and applying the inverse spatial transform) to the prediction signal. SUMMARY OF THE INVENTION [Problems that the invention aims to solve]

[0010] However, in scalable video coding, it is preferable to achieve higher coding efficiency.

[0011] Therefore, an object of the present invention is to provide a concept for scalable video coding that achieves higher coding efficiency. [Means for solving the problem]

[0012] This objective is achieved by the subject matter of the enclosed independent claims.

[0013] One embodiment of the present invention is that if the subdivision of each transformation coefficient block's subblocks is controlled based on the base layer residual signal or base layer signal, then encoding based on the subblocks of the enhancement layer's transformation coefficient block can be performed more efficiently. In particular, by utilizing each base layer hint, the subblocks become longer along the horizontal spatial frequency axis with respect to the edge extension observable from the base layer residual signal or base layer signal. Thus, the shape of the subblocks can be fitted to the estimated energy distribution of the transformation coefficients of the enhancement layer's transformation coefficient block such that, with increasing probability, each subblock is filled with either nearly entirely important transformation coefficients (i.e., transformation coefficients not quantized to zero) or non-important transformation coefficients (i.e., only transformation coefficients quantized to zero), while with decreasing probability, any subblock has an equal number of important transformation coefficients on one hand and non-important transformation coefficients on the other. However, the encoding efficiency for encoding transformation coefficient blocks in the enhancement layer increases due to the fact that subblocks without important transformation coefficients can be efficiently indicated in the data stream simply by using a single flag, and that subblocks almost entirely filled with important transformation coefficients do not require the wasted signaling to encode the non-important transformation coefficients scattered therein.

[0014] One embodiment of the present invention provides a better predictor for predictively encoding an enhancement layer signal in scalable video coding, which is achieved by forming an enhancement layer prediction signal from an inter-layer prediction signal and an enhancement layer intra-prediction signal in a manner with different weightings with respect to different spatial frequency components, that is, by forming a weighted average of the inter-layer prediction signal and the enhancement layer intra-prediction signal in the portion to be reconstructed, so that the weightings to which the inter-layer prediction signal and the enhancement layer intra-prediction signal contribute to the enhancement layer prediction signal vary with respect to different spatial frequency components. Therefore, it is possible to interpret the enhancement layer prediction signal from the inter-layer prediction signal and the enhancement layer intra-prediction signal in a manner optimized with respect to the spectral characteristics of the individual contributing components (i.e., the inter-layer prediction signal on the one hand and the enhancement layer intra-prediction signal on the other hand). For example, the inter-layer prediction signal is obtained from the reconstructed base layer signal based on improvements in resolution or quality. The interlayer prediction signal is more accurate at lower frequencies compared to higher frequencies. As far as the enhancement layer intra-prediction signal is concerned, the opposite is true: its accuracy increases with higher frequencies compared to lower frequencies. In this example, at lower frequencies, the contribution of the interlayer prediction signal to the enhancement layer prediction signal exceeds, for each weighting, the contribution of the enhancement layer intra-prediction signal to the enhancement layer prediction signal. However, at higher frequencies, it does not exceed the contribution of the enhancement layer intra-prediction signal to the enhancement layer prediction signal. Therefore, a more accurate enhancement layer prediction signal is achieved. As a result, coding efficiency increases, leading to a higher compression ratio.

[0015] Different possibilities are described by various embodiments to incorporate the concepts outlined above into any scalable video coding based on the concept. For example, the formation of a weighted average is formed in either the spatial domain or the transformation domain. Performing a spectral weighted average requires transformations to be performed on individual contributions, namely the inter-layer prediction signal and the enhancement layer intra-prediction. However, it is preferable to avoid spectrally filtering either the inter-layer prediction signal or the enhancement layer intra-prediction signal in the spatial domain, including, for example, FIR or IIR filtering. However, performing the formation of a spectral weighted average in the spatial domain avoids the detour of individual contributions to the weighted average via the transformation domain. The decision of which domain is actually chosen to perform the formation of the spectral weighted average depends on whether the scalable video data stream contains residual signals in the form of transformation coefficients with respect to the portion that should now be composed in the enhancement layer signal. If it does not, the detour via the transformation domain is avoided. On the other hand, if residual signals are present, the detour via the conversion region is even more advantageous because it allows for the direct addition of the transmitted residual signals within the conversion region to the spectrally weighted average within the conversion region.

[0016] One embodiment of the present invention is that information available from the encoding / decoding of the base layer, i.e., base layer hints, is used to make the motion compensation prediction of the enhancement layer more efficient by more efficiently encoding the enhancement layer motion parameters. In particular, the set of motion parameter candidates collected from adjacent, already reconstructed blocks of the enhancement layer signal frame is expanded by one or more base layer motion parameter sets of blocks of the base layer signal, co-located in the block of the enhancement layer frame. As a result, the available quality of the motion parameter candidate set is improved on the basis that the motion compensation prediction of the block of the enhancement layer signal is performed by selecting one of the motion parameter candidates from the expanded motion parameter candidate set and using that selected motion parameter candidate for prediction. Additionally, or alternatively, the list of motion parameter candidates for the enhancement layer signal is ordered depending on the base layer motion parameters involved in the encoding / decoding of the base layer. Therefore, the probability distribution for selecting enhancement layer motion parameters from an ordered list of motion parameter candidates is compressed so that, for example, explicit index syntax elements are encoded using fewer bits, for example, using entropy coding. Furthermore, additionally, or alternatively, the index used in the base layer coding / decoding supports as the basis for determining the index in the motion parameter candidate list for the enhancement layer. Thus, any signaling of the index for the enhancement layer is completely avoided. Or, simply with respect to the index, the deviation of the prediction thus determined is transmitted in the enhancement layer substream. As a result, coding efficiency is improved.

[0017] One embodiment of the present invention is that scalable video coding is made more efficient by evaluating the spatial variation of base layer coding parameters on the base layer signal, thereby deriving / selecting subblock subdivisions to be used for enhancement layer prediction from a set of possible subblock subdivisions of the enhancement layer block. For this reason, if so, less signaling overhead must be spent in the enhancement layer data stream to represent these subblock subdivisions. The subblock subdivisions thus selected are used when predictively coding / decoding the enhancement layer signal.

[0018] One embodiment of the present invention is that the coding efficiency of scalable video coding is increased by substituting lost spatial intra-prediction parameter candidates in the spatial adjacency of the current block of the enhancement layer by using intra-prediction parameters of co-located blocks of the base layer signal. Thus, the coding efficiency for coding spatial intra-prediction parameters is increased, or more precisely, expected to increase, due to the improved predictive quality of the set of intra-prediction parameters of the enhancement layer. A suitable predictor for the intra-prediction parameters for the intra-predicted blocks of the enhancement layer is useful and increases the likelihood that the signaling of the intra-prediction parameters of each enhancement layer block will be performed with fewer bits on average.

[0019] Further advantages are described in the dependent claims.

[0020] A preferred embodiment is described in detail below with reference to the drawings. [Brief explanation of the drawing]

[0021] [Figure 1]Figure 1 is a block diagram showing one embodiment of a scalable video encoder in which the embodiments and aspects described herein can be implemented. [Figure 2] Figure 2 is a block diagram showing one embodiment of a scalable video decoder that fits the scalable video encoder of Figure 1, and the embodiments and aspects described herein can be similarly implemented. [Figure 3] Figure 3 is a block diagram showing a more specific embodiment of a scalable video encoder capable of implementing the embodiments and aspects described herein. [Figure 4] Figure 4 is a block diagram of a scalable video decoder that fits the scalable video encoder of Figure 3, and the embodiments and aspects described herein can be similarly implemented. [Figure 5] Figure 5 is a schematic diagram showing versions of video, its base layer, and enhancement layer, further illustrating the encoding / decoding ordering. [Figure 6] Figure 6 is a schematic diagram showing a portion of a layered video signal to illustrate possible prediction modes for the enhancement layer. [Figure 7] Figure 7 is a diagram illustrating the formation of the enhancement layer prediction signal by spectrally varying the weighting between the enhancement layer intra-prediction signal and the inter-layer prediction signal, according to the embodiment. [Figure 8] Figure 8 is a schematic diagram of syntactic elements likely included within the enhancement layer's subsystem. [Figure 9] Figure 9 is a schematic diagram illustrating a possible realization of the formation shown in Figure 7, according to an embodiment in which the formation / bonding is performed within a spatial domain. [Figure 10] Figure 10 is a schematic diagram illustrating the realization of the formation shown in Figure 7, according to an embodiment in which the formation / binding is performed within the spectral region. [Figure 11]Figure 11 is a schematic diagram showing a portion extracted from a layered video signal to illustrate the derivation of spatial intra-prediction parameters from the base layer to the enhancement layer signal according to the embodiment. [Figure 12] Figure 12 is a schematic diagram illustrating an extension of the derivation shown in Figure 11 according to an embodiment. [Figure 13] Figure 13 is a schematic diagram showing a set of spatial intra-prediction parameter candidates into which a set of spatial intra-prediction parameter candidates obtained from the base layer is inserted according to the embodiment. [Figure 14] Figure 14 is a schematic diagram showing a portion extracted from a layered video signal to illustrate the granular derivation of prediction parameters from the base layer according to the embodiment. [Figure 15] Figures 15a and 15b are schematic diagrams illustrating how to select the appropriate subdivision for the current block using spatial variations in base layer motion parameters, following two different examples within the base layer. [Figure 15c] Figure 15c is a schematic diagram illustrating the first possible coarse selection method among the possible sub-block subdivisions for the current enhancement layer block. [Figure 15d] Figure 15d is a schematic diagram illustrating a second possibility for the coarsest selection method among the possible sub-block subdivisions for the current enhancement layer block. [Figure 16] Figure 16 is a schematic diagram showing a portion extracted from a layered video signal to illustrate the use of sub-block subdivision derivation for the current enhancement layer block. [Figure 17] Figure 17 is a schematic diagram showing a portion extracted from a layered video signal to illustrate the use of base layer hints for efficiently encoding enhancement layer motion parameter data. [Figure 18] Figure 18 is a schematic diagram illustrating the first possibility for increasing the efficiency of enhancement layer motion parameter signaling. [Figure 19a]Figure 19a is a schematic diagram illustrating a second possibility of using base layer hints to make enhancement layer motion parameter signaling more efficient. [Figure 19b] Figure 19b is a schematic diagram illustrating the first possibility of transporting the base layer in the order listed in the enhancement layer motion parameter candidates. [Figure 19c] Figure 19c is a schematic diagram illustrating a second possibility of transporting the base layer in the order listed in the enhancement layer motion parameter candidates. [Figure 20] Figure 20 is a schematic diagram illustrating another possibility for utilizing base layer hints to make enhancement layer motion parameter signaling more efficient. [Figure 21] Figure 21 is a schematic diagram illustrating a portion extracted from a layered video signal to illustrate an embodiment in which the sub-blocks of the conversion coefficient block are appropriately adjusted to hints obtained from the base layer. [Figure 22] Figure 22 is a schematic diagram illustrating different possible methods for obtaining appropriate sub-block subdivisions of the conversion coefficient block from the base layer. [Figure 23] Figure 23 is a block diagram showing a more detailed embodiment for a scalable video decoder. The embodiments and aspects described herein may be implemented. [Figure 24a] Figure 24a is a block diagram showing a scalable video encoder that fits the embodiment of Figure 23. The embodiments and aspects described herein can be implemented. [Figure 24b] Figure 24b is a block diagram showing a scalable video encoder conforming to the embodiment of Figure 23. The embodiments and aspects described herein can be implemented. [Figure 25]Figure 25 is a schematic diagram illustrating the generation of the interlayer intra-prediction signal by summing the (upsampled / filtered) base layer reconstructed signal (BL Reco) with the spatial intra-prediction using the already encoded adjacent block difference (EH Diff) signal. [Figure 26] Figure 26 is a schematic diagram illustrating the generation of the interlayer intra-prediction signal by summing the (upsampled / filtered) base layer residual signal (BL Resi) with the spatial intra-prediction using the reconfigured enhancement layer samples (EH Reco) of already encoded adjacent blocks. [Figure 27] Figure 27 is a schematic diagram illustrating the generation of the interlayer intra-prediction signal by frequency-weighted summation of the (upsampled / filtered) base layer reconstructed signal (BL Reco) and the spatial intra-prediction using the reconstructed enhancement layer samples (EH Reco) of already encoded adjacent blocks. [Figure 28] Figure 28 is a schematic diagram illustrating the base layer signal and enhancement layer signal used in the specification. [Figure 29] Figure 29 is a schematic diagram illustrating the motion compensation prediction of the enhancement layer. [Figure 30] Figure 30 is a schematic diagram illustrating the prediction using base layer residuals and enhancement layer reconstruction. [Figure 31] Figure 31 is a schematic diagram illustrating prediction using BL reconstruction and EL difference signals. [Figure 32] Figure 32 is a schematic diagram illustrating the prediction using BL reconstruction and a second hypothesis for the EL difference signal. [Figure 33] Figure 33 is a schematic diagram illustrating predictions using BL reconstruction and EL reconstruction. [Figure 34] Figure 34 is a schematic diagram illustrating, as an example, the decomposition of an image into square blocks and the corresponding quadtree structure. [Figure 35] Figure 35 is a schematic diagram illustrating the allowable disassembly of a square block within a subblock in a preferred embodiment. [Figure 36] Figure 36 is a schematic diagram illustrating the position of the motion vector prediction. (a) represents the position of the spatial candidate, and (b) represents the position of the temporal candidate. [Figure 37] Figure 37 is a schematic diagram illustrating the block (a) that integrates the algorithm and the redundancy check (b) performed for the spatial candidate. [Figure 38] Figure 38 is a schematic diagram illustrating the block (a) that integrates the algorithm and the redundancy check (b) performed for the spatial candidate. [Figure 39] Figure 39 is a schematic diagram illustrating the scanning directions (diagonal, vertical, horizontal) for a 4x4 transformation block. [Figure 40] Figure 40 is a schematic diagram illustrating the scanning directions (diagonal, vertical, horizontal) for an 8x8 transformation block. The shaded areas define important subgroups. [Figure 41] Figure 41 is a 16x16 transformation diagram in which only diagonal scanning is defined. [Figure 42] Figure 42 is a schematic diagram illustrating vertical scanning for 16x16 conversion, as proposed in JCTVC-G703. [Figure 43] Figure 43 is a schematic diagram illustrating the implementation of vertical and horizontal scanning for a 16x16 transformation block. Each coefficient subgroup is defined as either a row or a column. [Figure 44] Figure 44 is a schematic diagram illustrating the vertical and horizontal scanning for a 16x16 transformation block. [Figure 45] Figure 45 is a schematic diagram illustrating backward-adapted enhancement layer intra-prediction using adjacent reconstructed enhancement layer samples and reconstructed base layer samples. [Figure 46]Figure 46 is a schematic diagram showing an image / frame of the enhancement layer to illustrate the spatial insertion of the difference signal according to the embodiment. [Modes for carrying out the invention]

[0022] Figure 1 shows a general embodiment for a scalable video encoder incorporating the embodiments outlined below. The scalable video encoder in Figure 1 is generally denoted by reference numeral 2 and receives and encodes video 4. The scalable video encoder 2 is configured to encode video 4 into a data stream 6 in a scalable manner. That is, the data stream 6 includes a first portion 6a having video 4 encoded therein with a first information content, and another portion 6b having video 4 encoded therein with a larger information content than the first portion 6a. For example, the information content of portions 6a and 6b differ in quality or fidelity, i.e., in the amount of pixelary deviation from the original video 4 and / or spatial resolution. However, other forms of different information content may also apply, for example, color fidelity. Portion 6a is called the base layer data stream or base layer substream. Portion 6b, on the other hand, is called the enhancement layer data stream or enhancement layer substream.

[0023] The scalable video encoder 2 is configured to utilize redundancy between versions 8a and 8b of the reconfigurable video 4, on the one hand from the base layer substream 6a without the enhancement layer substream 6b, and on the other hand from both substreams 6a and 6b. To do so, the scalable video encoder 2 can use interlayer prediction.

[0024] As shown in Figure 1, the scalable video encoder 2 either accepts two versions of video 4, 4a and 4b. Both versions 4a and 4b have different amounts of information content, just as the base layer substream 6a and the enhancement layer substream 6b do. Thus, for example, the scalable video encoder 2 is configured to generate substreams 6a and 6b. As a result, the base layer substream 6a has version 4a encoded within it. On the other hand, the enhancement layer data stream (substream) 6b has version 4b encoded within it, using an inter-layer prediction based on the base layer substream 6b. Encoding of both substreams 6a and 6b is lossy.

[0025] Even if the scalable video encoder 2 simply receives the original version of video 4, it is configured to obtain two versions 4a and 4b intra-from it by obtaining a base layer version 4a, for example, through spatial downscaling and / or tone mapping from a higher bit depth to a lower bit depth.

[0026] Figure 2 shows a scalable video decoder that fits the scalable video encoder 2 of Figure 1 in a similar manner to that which is suitable for incorporating the embodiments outlined below. The scalable video decoder of Figure 2 is generally denoted by reference numeral 10. The scalable video decoder is generally configured to decode the encoded data stream 6 so as to reconstruct the enhancement layer version 8b of the video from both portions 6a and 6b of the data stream 6 if they reach the scalable video decoder 10 in a complete manner, or so as to reconstruct the base layer version 8a of the video from them if, for example, portion 6b is unavailable due to transmission loss. That is, the scalable video decoder 10 is configured to be able to reconstruct version 8a from only the base layer substream 6a, and to be able to reconstruct version 8a from both portions 6a and 6b using interlayer predictions.

[0027] Before a more detailed description of the following embodiments of the present invention (i.e., embodiments show how the embodiments of Figures 1 and 2 are specifically realized), a more detailed realization of the scalable video encoder and decoder of Figures 1 and 2 will be described with respect to Figures 3 and 4. Figure 3 shows a scalable video encoder 2 comprising a base layer encoder 12, an enhancement layer encoder 14, and a multiplexer 16. The base layer encoder 12 is configured to encode a base layer version 4a of the input video. The enhancement layer encoder 14 is configured to encode an enhancement layer version 4b of the video. Thus, the multiplexer 16 receives a base layer substream 6a from the base layer encoder 12 and an enhancement layer substream 6b from the enhancement layer encoder 14 and multiplexes both into an encoded data stream 6 for output.

[0028] As shown in Figure 3, both encoders 12 and 14 are predictive encoders that use, for example, spatial and / or temporal prediction to encode their respective input versions 4a and 4b into their respective substreams 6a and 6b. In particular, encoders 12 and 14 are each hybrid video block encoders. That is, each of encoders 12 and 14 is configured to encode their respective input versions of the video on a block-by-block basis, while, for example, the images or frames of video versions 4a and 4b are selected from among different prediction modes for each block of the subdivided blocks. The different prediction modes of the base layer encoder 12 include spatial and / or temporal prediction modes. Enhancement layer encoder 14, on the other hand, additionally supports an inter-layer prediction mode. The subdivision within the blocks differs between the base layer and the enhancement layer. The prediction mode, prediction parameters for the prediction mode selected for each block, prediction residuals, and optionally, block subdivisions for each video version are described by respective encoders 12 and 14 using their respective syntax, which includes syntactic elements encoded sequentially into their respective substreams 6a and 6b using entropy coding. Interlayer prediction is used, for example, one or more times to predict the enhancement layer video, prediction mode, prediction parameters, and / or samples of block subdivisions, as mentioned in a couple of examples. Thus, both the base layer encoder 12 and the enhancement layer encoder 14 include prediction encoders 18a and 18b, respectively, followed by entropy encoders 19a and 19b. Meanwhile, the prediction encoders 18a and 18b form syntactic element streams from the inbound versions 4a and 4b, respectively, using prediction coding. The entropy encoders 19a and 19b entropy encode the syntactic elements output by their respective prediction encoders 18a and 18b. As mentioned above, the interlayer prediction of encoder 2 is relevant at different times in the encoding procedure of the enhancement layer. Thus, the predictive encoder 18b is shown to be connected to one or more of the predictive encoder 18a, its output, and the entropy encoder 19a.Similarly, the entropy encoder 19b optionally utilizes interlayer predictions, for example, by predicting the context used for entropy coding from the base layer. Thus, the entropy encoder 19b is optionally shown to be connected to any of the elements of the base layer encoder 12.

[0029] In the same manner as in Figure 2 for Figure 1, Figure 4 shows a possible implementation of a scalable video decoder 10 that fits the scalable video encoder of Figure 3. Thus, the scalable video decoder 10 of Figure 4 comprises a demultiplexer 40 that receives a data stream 6 to obtain substreams 6a and 6b, a base layer decoder 80 configured to decode the base layer substream 6a, and an enhancement layer decoder 60 configured to decode the enhancement layer substream 6b. As shown, the decoder 60 is connected to the base layer decoder 80 to receive information from there in order to utilize interlayer predictions. This allows the base layer decoder 80 to reconstruct a base layer version 8a from the base layer substream 6a. The enhancement layer decoder 60 is then configured to reconstruct an enhancement layer version 8b of the video using the enhancement layer substream 6b. Similar to the scalable video encoder in Figure 3, the enhancement layer decoder 60 and the base layer decoder 80 each contain an entropy decoder 100,320, followed by a prediction decoder 102,322, respectively.

[0030] To simplify understanding of the following embodiments, Figure 5 illustrates different versions of video 4, namely base layer versions 4a and 8a, which are deviated from each other by coding loss. Similarly, enhancement layer versions 4b and 8b are deviated from each other by coding loss, respectively. The base layer signal and the enhancement layer signal consist of sequences of images 22a and 22b, respectively. They are shown in Figure 5 as being registered to each other along the time axis 24, i.e., to the base layer version of image 22a, as well as the temporally corresponding image 22b of the enhancement layer signal. As previously mentioned, image 22b represents video 4 with higher spatial resolution and / or higher fidelity, etc. (e.g., higher bit depth of the image sample values). Solid and dotted lines are used to indicate the coding / decoding order defined between images 22a and 22b. As shown in the example in Figure 5, the encoding / decoding order traverses images 22a and 22b in such a way that the base layer image 22a at a given timestamp / moment is traversed before the enhancement layer image 22b at the same timestamp of the enhancement layer signal. With respect to the time axis 24, images 22a and 22b are traversed by the encoding / decoding order 26 in the order of their provided times. However, an order that deviates from the order of the provided times of images 22a and 22b is also possible. Neither the encoder 2 nor the decoder 10 needs to encode / decode sequentially along the encoding / decoding order 26. Rather, encoding / decoding is used in parallel. The encoding / decoding order 26 defines the availability between adjacent base layer and enhancement layer signal portions in a spatial, temporal, and / or inter-layer sense. As a result, when encoding / decoding the current portion of the enhancement layer, the available portion of that current enhancement layer portion is defined throughout the encoding / decoding order. Therefore, since the simply adjacent portions available according to this encoding / decoding order 26 are used by the encoder for prediction, the decoder accesses the same source of information to refine the prediction.

[0031] With respect to the following figures, it will be explained how the scalable video encoder or decoder described in Figures 1 to 4 forms an embodiment of the present invention according to one embodiment of this application. Possible examples of the embodiments described below will be discussed using the designation “Embodiment C”.

[0032] In particular, Figure 6 illustrates image 22b of the enhancement layer signal indicated using reference numeral 360 and image 22a of the base layer signal indicated using reference numeral 200. Temporally corresponding images of different layers are shown in the manner in which they were recorded relative to each other with respect to the time axis 24. Diagonal lines are used to distinguish the portion of the base layer signal 200 and the enhancement layer signal 36 that has already been encoded / decoded according to the encoding / decoding order from the portion 36 that has not yet been encoded or decoded according to the encoding / decoding order shown in Figure 5. Figure 6 also shows a portion 28 of the enhancement layer signal 360 that is currently being encoded / decoded.

[0033] According to the embodiment described, the prediction of portion 28 uses both intra-layer predictions within the enhancement layer itself and inter-layer predictions from the base layer to predict portion 28. However, the predictions are coupled so that these predictions contribute to the final prediction of portion 28 in a spectrally varying manner. As a result, in particular, the ratio between the two contributions varies spectrally.

[0034] In particular, portion 28 is predicted spatially or temporally from an already reconstructed portion of the enhancement layer signal 400, i.e., any portion in the enhancement layer signal 400 indicated by the shaded area in Figure 6. Spatial prediction is illustrated using arrow 30. Temporal prediction, on the other hand, is illustrated using arrow 32. Temporal prediction includes, for example, motion compensation prediction as motion vector information is transmitted in the enhancement layer substream for the current portion 28. The motion vector indicates the replacement of a portion of the reference image of the enhancement layer signal 400 to be copied in order to obtain the temporal prediction of the current portion 28. Spatial prediction 30 includes, within the current portion 28, a spatially adjacent portion to be estimated, an already encoded / decoded portion of image 22b, and the spatially adjacent current portion 28. For this purpose, intra-predictive information, such as the estimated (or angular) direction, is provided in the enhancement layer substream for the current portion 28. Combinations of spatial prediction 30 and temporal prediction 32 are also used. In any case, the result is that the intra-prediction signal 34 of the enhancement layer is obtained as shown in Figure 7.

[0035] Interlayer prediction is used to obtain another prediction for the current portion 28. For this purpose, the base layer signal 200 is modified in portion 36 of the enhancement layer signal 400, which is spatially and temporally corresponding to the current portion 28, so that the interlayer prediction signal for the current portion 28 undergoes resolution or quality improvement to obtain increased potential resolution. The improvement procedure is illustrated using arrow 38 in Figure 6 and results in the interlayer prediction signal 39 as shown in Figure 7.

[0036] Therefore, two prediction contributions 34 and 39 exist for the current portion 28. The weighted average of both contributions is then formed with respect to the current portion 28, so that the weights of the inter-layer prediction signal and the enhancement layer intra-prediction signal contribute to the enhancement layer prediction signal 42, in a way that varies the spatial frequency components differently, as schematically shown in Figure 7 by 44. Figure 7 exemplifies the case where, with respect to all spatial frequency components, the weights that the prediction signals 34 and 38 contribute to the final prediction signal are added together by adding the same value 46 for all spectral components, however, the ratio between the weight applied to prediction signal 34 and the weight applied to prediction signal 39 is spectrally varied.

[0037] On the other hand, the prediction signal 42 is used directly in the current portion 28 by the enhancement layer signal 400. Alternatively, the residual signal is provided in the enhancement layer substream 6b of the current portion 28, which is brought into the reconfigured version 54 of the current portion 28 by coupling 50 with the prediction signal 42, for example, as shown in the summation in Figure 7. As an intermediate note, it should be noted that both the scalable video encoder and decoder are hybrid video decoders / encoders that use predictive coding and transform coding to encode / decode the predictive residual.

[0038] To summarize the descriptions in Figures 6 and 7, the enhancement layer substream 6b includes an intra-prediction parameter 56 to control spatial and / or temporal predictions 30, 32 with respect to the current portion 28, an optional weighting parameter 58 to control the formation of a spectrally weighted average 41, and residual information 59 to show the residual signal 48. The scalable video encoder, on the other hand, determines all of these parameters 56, 58, 59 accordingly and inserts the parameters 56, 58, 59 into the enhancement layer substream 6b. The scalable video decoder uses the parameters 56, 58, 59 to reconstruct the current portion 28 as outlined above. All of these elements 56, 58, 59 undergo some quantization, i.e., using a ratio / distortion cost function as the quantization. The scalable video encoder then determines these parameters / elements accordingly. Interestingly, encoder 2 uses the parameters / elements 56, 58, 59 thus determined to serve as the basis for any prediction for the portion of the enhancement layer signal 400, which follows, for example, in the order of encoding / decoding, in order to obtain a reconstructed version 54 with respect to the current portion 28.

[0039] Different possibilities exist for the weighting parameters 58 and how they control the formation of the spectral weighted average 41. For example, the weighting parameters 58 may indicate only one of two states with respect to the current portion 28: one state that activates the formation of the spectral weighted average as described above, and the other state that deactivates the contribution of the interlayer prediction signal 38. As a result, the final enhancement layer prediction signal 42 is then created solely by the enhancement layer intra-prediction signal 34. The weighting parameters 58 for the current portion 28 may switch between activating spectral weighted average formation and the interlayer prediction signal 39 forming the enhancement layer prediction signal 42 on its own. Alternatively, the weighting parameters 58 may be designed to indicate one of the three states / binaries mentioned. Or, the weighting parameters 58 may further control the spectral weighted average formation 41 with respect to the current portion 28 with respect to the spectral change in the ratio between the weights that the prediction signals 34 and 39 contribute to the final prediction signal 42. Later, it will be explained that spectral weighted averaging 41 involves filtering one or both of the predicted signals 34 and 39 before adding them together, for example, using a high-pass filter and / or a low-pass filter. In this case, the weighting parameter 58 indicates the filter characteristics or filter for the filter to be used with respect to the prediction of the current portion 28. Alternatively, it will be explained below that the weighting parameter 58 may indicate / set the values ​​of the individual weights of these spectral components, where the spectral weighting in spectral weighted averaging 41 is achieved by the individual weighting of the spectral components in the transformation region.

[0040] Additionally, or alternatively, the weighting parameters for the current portion 28 can indicate whether the spectral weighting in step 41 is performed in the transformation domain or the spatial domain.

[0041] Figure 9 illustrates an embodiment for performing a spectrally weighted averaging construction in the spatial domain. Predicted signals 39 and 34 are illustrated as being obtained in the form of their respective pixel arrays, which correspond to the current portion 28 pixel raster. To perform a spectrally weighted averaging construction, both pixel arrays of predicted signals 34 and 39 are shown to be subjected to filtering. Figure 9 illustrates filtering as an example by showing filter kernels 62 and 64 moving through the pixel arrays of predicted signals 34 and 39 to perform, for example, FIR filtering. However, IIR filtering is also possible. Furthermore, only one of the predicted signals 34 and 39 may be subjected to filtering. Since the transfer functions of both filters 62 and 64 are different, the sum 66 of the results of filtering the pixel arrays of predicted signals 39 and 34 yields the result of the spectrally weighted averaging construction, i.e., the enhancement layer predicted signal 42. In other words, the sum 66 easily adds the co-located samples in the filtered predicted signals 39 and 34 using filters 62 and 64, respectively. As a result, 62-66 yield a spectrally weighted average configuration 41. Figure 9 illustrates that, in the case of residual information 59 existing in the form of transformation coefficients, the residual signal 48 in the transformation region is shown, and the inverse transformation 68 is used to yield the spatial region in the form of a pixel array 70, resulting in a concatenation 52 that yields a reconstructed version 55, which is achieved by a simple pixel-wise addition of the residual signal array 70 and the enhancement layer prediction signal 42.

[0042] Again, recall that the prediction is performed by a scalable video encoder and decoder, using the prediction for reconstruction within the decoder and encoder, respectively.

[0043] Figure 10 illustrates how spectral weighted averaging construction is performed within the transformation region. Here, the pixel arrays of the predicted signals 39 and 34 undergo transformations 72 and 74, respectively, resulting in spectral decompositions 76 and 78, respectively. Each spectral decomposition 76 and 78 creates a transformation coefficient array with one transformation coefficient per spectral component. Each transformation coefficient block 76 and 78 is multiplied by the corresponding blocks of weighting, namely blocks 82 and 84. As a result, for each spectral component, the transformation coefficients in blocks 76 and 78 are individually weighted. For each spectral component, the weights in blocks 82 and 84 add a common value to all spectral components. However, this is not mandatory. In effect, the multiplier 86 between blocks 76 and 82, and the multiplier 88 between blocks 78 and 84, represent spectral filtering within the transformation region, respectively. Then, the conversion coefficient / spectral component addition 90 completes the spectral weighted average configuration 41 to yield a converted region version of the enhancement layer prediction signal 42 in the form of a block of conversion coefficients. As shown in Figure 10, in the case of the residual signal 59 showing the residual signal 48 in the form of a block of conversion coefficients, the residual signal 59 is easily converted coefficient additive concatenation (or another concatenation) 52 with the conversion coefficient block representing the enhancement layer prediction signal 42 to yield a reconstructed version of the current portion 28 in the converted region. Thus, the inverse transformation 84 applied to the result of the concatenation 52 yields a pixel array that reconstructs the current portion 28, i.e., the reconstructed version 54.

[0044] As mentioned above, the current parameters in the enhancement layer substream 6b for the current portion 28, for example, residual information 59, or weighting parameters 58, indicate whether the average configuration 41 is performed in the transformation region shown in Figure 10, or in the spatial region according to Figure 9. For example, the residual information 59 also indicates the absence of any transformation coefficient blocks for the current portion 28. Alternatively, the spatial region is used. Or, the weighting parameters 58 switch between both regions, regardless of whether the residual information 59 includes transformation coefficients or does not.

[0045] It is then explained that, in order to obtain an interlayer enhancement layer prediction signal, the difference signal is calculated and managed between the already reconstructed portion of the enhancement layer signal and the interlayer prediction signal. The spatial prediction for the first portion of the difference signal is co-located with a portion of the enhancement layer signal, and the spatial prediction of the difference signal is now reconstructed from the second portion of the difference signal, which is spatially adjacent to the first portion of the enhancement layer signal and belongs to the already reconstructed portion, and is used to spatially predict the difference signal at that time. Alternatively, the temporal prediction of the first portion of the difference signal is co-located with a portion of the enhancement layer signal, which is now reconstructed from the second portion of the difference signal, which belongs to a previously reconstructed frame of the enhancement layer signal, and is used to obtain a temporally predicted difference signal. The interlayer prediction signal and the predicted difference signal are used to obtain an interlayer enhancement layer prediction signal, which is then coupled with the interlayer prediction signal.

[0046] The following diagrams illustrate how the scalable video encoder or decoder described above in relation to Figures 1 to 4 can be realized to form an embodiment of the present invention according to another aspect of this application.

[0047] To illustrate this point, refer to Figure 11. Figure 11 shows the possibility of performing the current spatial prediction 30 for portion 28. Consequently, the following description of Figure 11 is combined with the descriptions relating to Figures 6-10. In particular, the following description will be explained later in relation to the illustrated examples of implementation by referring to “Examples” X and Y.

[0048] The situation shown in Figure 11 corresponds to that shown in Figure 6, namely, the base layer signal 200 and the enhancement layer signal 400. Already encoded / decoded portions are indicated using diagonal lines. Within the enhancement layer signal 400, the portions currently to be encoded / decoded have adjacent blocks 92 and 94. Here, illustratively, with respect to both blocks 92 and 94 having the same size as the current block 28, block 92 is drawn above the current portion 28 and block 94 is drawn to the left. However, size matching is not mandatory. Rather, the portions of the blocks into which the image 22b of the enhancement layer signal 400 is subdivided have different sizes. They are not even limited to squares. They may be rectangles or other shapes. Furthermore, the current block 28 has adjacent blocks that are not clearly represented in Figure 11. However, the adjacent blocks have not yet been decoded / encoded. That is, the adjacent blocks follow the encoding / decoding order and are therefore not available for prediction. Beyond this, there are other blocks (such as block 96, which is diagonally adjacent to the current block 28, for example, block 96, which is adjacent to the current block 28 in the upper left corner) that have already been encoded / decoded according to the encoding / decoding order. However, blocks 92 and 94 are predetermined adjacent blocks that play a role in predicting the intra-prediction parameters for the current block 28, which receives the intra-prediction 30 in the example considered here. The number of such predetermined adjacent blocks is not limited to two; it may be more, or even just one.

[0049] The scalable video encoder and scalable video decoder determine a predetermined set of adjacent blocks, here blocks 92 and 94, from a set of already encoded adjacent blocks. Here, blocks 92-96 depend on a predetermined sample position 98 in the current section 28, for example, the upper left sample. For example, only those already encoded adjacent blocks of the current section 28 form a set of "predetermined adjacent blocks" that include a sample position directly adjacent to the predetermined sample position 98. In any case, adjacent already encoded / decoded blocks include a sample 102 adjacent to the current block 28 based on the sample value that should be spatially predicted for the region of the current block 28. For this purpose, spatial prediction parameters such as 56 are shown in the enhancement layer substream 6b. For example, the spatial prediction parameter for the current block 28 indicates the spatial direction in which the sample value of sample 102 should be copied into the region of the current block 28.

[0050] In any case, at least with respect to the spatially corresponding region of the temporally corresponding image 22a, when spatially predicting the current block 28 using block prediction as described above, for example using block prediction selection between spatial prediction mode and temporal prediction mode, the scalable video decoder / encoder has already reconstructed (in the case of an encoder, encoded) the base layer 200 using the base layer substream 6a.

[0051] In Figure 11, several blocks 104 into which the image 22a, arranged in time with the base layer signal 200, is subdivided lie within and around the region locally corresponding to the exemplarily represented current portion 28. This is just the case for spatially predicted blocks in the enhancement layer signal 400. The spatial prediction parameters, and the selection of the spatial prediction mode, are included in or shown within the base layer substream for those blocks 104 in the base layer signal 200, which is shown with respect to the base layer signal.

[0052] Here, as an example, the intra-prediction parameters are used and encoded in the following bitstream to enable the reconstruction of the enhancement layer signal from the encoded data stream relating to the selected block 28 for spatial intra-layer prediction 30.

[0053] Intra-prediction parameters are often coded using the concept of "most likely intra-prediction parameters," which is a fairly small subset of all possible intra-prediction parameters. For example, the "most likely intra-prediction parameter" set might contain one, two, or three intra-prediction parameters. On the other hand, the set of all possible intra-prediction parameters might contain 35 intra-prediction parameters. If an intra-prediction parameter is included in the most likely intra-prediction parameter set, it is represented by fewer bits in the bitstream. If an intra-prediction parameter is not included in the most likely intra-prediction parameter set, its signaling in the bitstream requires more bits. Therefore, the amount of bits that should be spent on syntactic elements to signal the intra-prediction parameter for the current intra-predicted block depends on the quality of the most likely, or perhaps favorable, set of intra-prediction parameters. Using this concept, on average, a low number of bits are needed to code an intra-prediction parameter. The evaluation is based on obtaining a suitable set of the most likely intra-prediction parameters.

[0054] Typically, the most likely set of intra-prediction parameters is selected in such a way that it includes the intra-prediction parameters of directly adjacent blocks and / or, additionally, often uses intra-prediction parameters, for example, in the form of initial setting parameters. For example, since adjacent blocks have the same main gradient direction, it is generally advantageous to include the intra-prediction parameters of adjacent blocks in the most likely set of intra-prediction parameters.

[0055] However, if adjacent blocks are not encoded in spatial intra-prediction mode, those parameters will not be available to the decoder.

[0056] In scalable coding, however, it is possible to use intra-prediction parameters of co-located base layer blocks. Thus, according to the embodiment outlined below, this situation is utilized using intra-prediction parameters of co-located base layer blocks in the case of uncoded adjacent blocks in spatial intra-prediction mode.

[0057] As a result, according to Figure 11, a set of intra-prediction parameters that is likely favorable for the current enhancement layer block is constructed by checking the intra-prediction parameters of predetermined adjacent blocks, and, for example, by reclassifying each predetermined adjacent block into a co-located block in the base layer if the predetermined adjacent block does not have appropriate intra-prediction parameters related to it, since each predetermined adjacent block is not encoded in intra-prediction mode.

[0058] First, it is checked whether a predetermined adjacent block, such as block 92 or 94 of the current block 28, was predicted using a spatial intra-prediction mode. That is, it is checked whether a spatial intra-prediction mode was selected for that adjacent block. As a result, the intra-prediction parameters of that adjacent block are included in the set of intra-prediction parameters that are likely favorable for the current block 28, or, if any, alternatively, in the intra-prediction parameters of the co-located block 108 in the base layer. This process can be performed for each of the predetermined adjacent blocks 92 and 94.

[0059] For example, if each predetermined adjacent block is not a spatial intra-prediction block, then, using something like an initial prediction, the intra-prediction parameters of block 108 in the base layer signal 200 are instead included in a set of presumably favorable prediction parameters for the current block 28, which are co-located in the current block 28. For example, the co-located block 108 is determined using a predetermined sample position 98 of the current block 28. That is, block 108 covers a position 106 that locally corresponds to a predetermined sample position 98 in the temporally arranged image 22a of the base layer signal 200. Naturally, a further check is performed to determine whether this co-located block 108 in the base layer signal 200 is actually a spatial intra-prediction block. This is illustrated in the case of Figure 11. However, if the co-located block is also not encoded in the intra-prediction mode, then the presumably favorable set of intra-prediction parameters is left with no contribution for its predetermined adjacent block. Alternatively, the initial intra-prediction parameters are used instead. That is, the initial intra-prediction parameters are inserted into a set of intra-prediction parameters that are likely to be favorable.

[0060] Therefore, if block 108, which is co-located with the current block 28, is spatially intrapredictive, then its intrapredictive parameters shown in the base layer substream 6a are used as a kind of alternative for predetermined adjacent blocks 92 or 94 of the current block 28, which do not have any intrapredictive parameters, because the intrapredictive parameters are encoded using another prediction mode, such as a time prediction mode.

[0061] In another embodiment, in a given case, even if each predetermined adjacent block is in an intra-prediction mode, the intra-prediction parameters of the predetermined adjacent block are replaced by the intra-prediction parameters of the co-located base layer block. For example, further checks are performed for any predetermined adjacent block in an intra-prediction mode, such as whether the intra-prediction parameters meet a predetermined criterion. If the predetermined criterion is not met by the intra-prediction parameters of the adjacent block, but the same criterion is met by the intra-prediction parameters of the co-located base layer block, the replacement is performed even for very adjacent intra-encoded blocks. For example, if the intra-prediction parameters of the adjacent block do not represent an angular intra-prediction mode (but, for example, a DC or planar intra-prediction mode), but the intra-prediction parameters of the co-located base layer block represent an angular intra-prediction mode, the intra-prediction parameters of the adjacent block are replaced by the intra-prediction parameters of the base layer block.

[0062] The interprediction parameters for the current block 28 are determined based on syntactic elements present in the encoded data stream, such as the enhancement layer substream 6b for the current block 28 and a set of potentially favorable intraprediction parameters. That is, syntactic elements are encoded using fewer bits in the case of the interprediction parameters for the current block 28 that are members of the potentially favorable intraprediction parameters than in the case of the remaining members of the set of possible intraprediction parameters that do not lead to the potentially favorable intraprediction parameters.

[0063] The possible set of intra-predictive parameters includes several angular modes, which are satisfied by copying from already encoded / decoded adjacent samples by copying along the angular direction of each mode / parameter; one DC mode, which is satisfied by setting the samples of the current block to a constant value determined based on already encoded / decoded adjacent samples, for example by some mean; and a planar mode, which is satisfied by setting the samples of the current block to a value distribution that follows a linear function of x and y slope and cutoff, for example, based on already encoded / decoded adjacent samples.

[0064] Figure 12 illustrates the potential use of alternative spatial prediction parameters obtained from the co-located block 108 of the base layer, along with syntactic elements shown in the enhancement layer substream. Figure 12 shows the current block 28 in an enlarged manner, along with the adjacent already encoded / decoded sample 102 and predetermined adjacent blocks 92 and 94. Figure 12 also exemplifies the angular direction 112 indicated by the spatial prediction parameters of the co-located block 108.

[0065] For the current block 28, the syntactic element 114 shown in the enhancement layer substream 6b can, for example, as shown in Figure 13, show a conditionally encoded index 118 in a list 122 which is the result of possible favorable intra-prediction parameters, illustrated here exemplarily as angular direction 124. Alternatively, if the actual intra-prediction parameter 116 is not in the most likely set 122 and is index 123 in a list 125 of possible intra-prediction modes that are excluded as possible, as shown in 127, then the candidates in list 122 will consequently identify the actual intra-prediction parameter 116. Encoding the syntactic element consumes fewer bits for the actual intra-prediction parameter that belongs in list 122. For example, the syntactic element includes a flag and an index part. The flag indicates whether to include or exclude members of list 122, or whether the index points to either list 122 or list 125, i.e., whether to include or exclude members of list 122. A syntactic element includes a field that identifies one of the members 124 of List 122 or an escape code. The syntactic element also includes a second field that identifies a member from List 125, which may include or exclude a member of List 122, in the case of an escape code. The order of members 124 in List 122 is determined, for example, based on initialization rules.

[0066] Therefore, the scalable video decoder obtains or retrieves the syntactic element 114 from the enhancement layer substream 6b. The scalable video encoder then inserts the syntactic element 114 into the enhancement layer substream 6b. Next, for example, the syntactic element 114 is used to index a spatial prediction parameter from list 122. When forming list 122, the aforementioned alternative is performed by checking whether predetermined adjacent blocks 92 and 94 are of a spatial prediction coding mode type. Otherwise, as described above, the co-located block 108 is checked in turn to see if it is a spatially predicted block, and if so, the same spatial prediction parameters, such as the angular direction 112 used to spatially predict this co-located block 108, are included in list 122. Also, if the base layer block 108 does not contain a suitable intra-prediction parameter, list 122 is excluded without contribution from the respective predetermined adjacent block 92 or 94. Because, to avoid List 122 being empty, for example, both predetermined adjacent blocks 92 and 98 are intra-predicted, and therefore at least one of the members 124 is unconditionally determined using the initial intra-prediction parameters, as is the case with co-located block 108 which lacks suitable intra-prediction parameters. Alternatively, List 122 may also be empty.

[0067] Naturally, the embodiments described with respect to Figures 11 to 13 can be linked to the embodiments outlined with respect to Figures 6 to 10. In particular, the intra-prediction obtained using the spatial intra-prediction parameters derived by bypassing the base layer according to Figures 11 to 13 represents the enhancement layer intra-prediction signal 34 of the embodiments in Figures 6 to 10, which is coupled to the inter-layer prediction signal 38 as described above in a spectrally weighted manner.

[0068] With respect to the following drawings, as described with respect to Figures 1 to 4, it will be explained how a scalable video encoder or decoder can be realized to form an embodiment of the present application according to another embodiment of the present application. Later, additional realizations for the embodiments described below will be presented with reference to Examples T and U.

[0069] Figure 14 shows images 22b and 22a of the enhancement layer signal 400 and the base layer signal 200, respectively, in a temporal registration method. The portion to be encoded / decoded is currently indicated by 28. According to the current embodiment, the base layer signal 200 is predictively encoded by a scalable video encoder and predictively reconstructed by a scalable video decoder using base layer coding parameters that spatially vary the base layer signal. The spatial variation is shown in Figure 14 using a shaded portion 132 enclosed by an unshaded region. Within the shaded portion 132, the base layer coding parameters used to predictively encode / reconstruct the base layer signal 200 are constant. When transitioning from the shaded portion 132 to the unshaded region, the base layer coding parameters change. According to the embodiment outlined above, the enhancement layer signal 400 is encoded / reconstructed within units of blocks. The current portion 28 is such a block. According to the embodiment outlined above, the subblock subdivision for the current portion 28 is selected from a set of possible subblock subdivisions based on the spatial variation of the base layer coding parameters within the co-located portion 134 of the base layer signal 200, i.e., within the spatially co-located portion of the temporally corresponding image 22a of the base layer signal 200.

[0070] In particular, for the current section 28, instead of showing it in the subdivision information of the enhancement layer substream 6b, the above description suggests selecting a subblock subdivision from the set of possible subblock subdivisions of the current section 28 such that the selected subblock subdivision is the coarsest of the set of possible subblock subdivisions. There, when the base layer signal is moved over to the co-located section 134, the base layer coding parameters subdivide the base layer signal 200 such that within each subblock of each subblock subdivision they are sufficiently similar to one another. For ease of understanding, refer to Figure 15a. Figure 15a shows section 28 in the co-located section 134, with the spatial variation of the base layer coding parameters indicated using diagonal lines. In particular, section 28 shows three different subblock subdivisions applied to block 28. In particular, quadtree subdivision is used illustratively in the case of Figure 15a. That is, the set of possible subblock subdivisions is (or is defined by) quadtree subdivisions. The three specific examples of subblock subdivision of section 28 shown in Figure 15a belong to different levels of hierarchy in the subdivision of the quadtree of block 28. From bottom to top, the level or coarseness of subdivision of block 28 within the subblocks increases. At the highest level, section 28 remains intact. At the next lower level, block 28 is subdivided into four subblocks. And at least one of the latter is further subdivided into four more subblocks at the next lower level, and so on. In Figure 15a, at each level, the quadtree subdivision is selected to be the one with the fewest number of subblocks, yet not one that overlaps with a base layer coding parameter change boundary. That is, in the case of Figure 15a, the quadtree subdivision of block 28 that should be selected to subdivide block 28 is the lowest one shown in Figure 15a. Here, the base layer coding parameters of the base layer are constant within each part co-located in each subblock of the subdivision.

[0071] Therefore, the subdivision information for block 28 does not need to be shown in the enhancement layer substream 6b. As a result, coding efficiency is increased. Moreover, as outlined, the method for obtaining subdivision is adaptable regardless of the current position of part 28 with respect to any grid, or any registration of the sample array of the base layer signal 200. Furthermore, the subdivision derivation works in particular for fractional spatial resolution ratios between the base layer and the enhancement layer.

[0072] Based on the sub-block subdivisions of section 28 determined in this way, section 28 is predictively reconstructed / encoded. It should be noted that, with respect to the above description, different possibilities exist to "measure" the coarseness of different available sub-block subdivisions of the current block 28. For example, the magnitude of coarseness is determined based on the number of sub-blocks. The more sub-blocks each subdivision has, the lower the level. This definition clearly does not apply in the case of Figure 15a, where the "magnitude of coarseness" is determined by a combination of the number of sub-blocks in each subdivision and the minimum size of all sub-blocks in each subdivision.

[0073] For completeness, Figure 15b illustrates the case where, using the subdivisions of Figure 35 as an example of available sets, one possible subdivision is selected from one set of available subdivisions for the current block 28. Different shaded (and unshaded) areas indicate regions co-located in the base layer signal that have the same base layer coding parameters associated with them.

[0074] As outlined above, the selection is performed by traversing the possible subblock subdivisions in a certain sequential order, such as the order of increasing or decreasing levels of roughness, and within each subblock subdivision, selecting that possible subblock subdivision from the possible subdivisions, provided that the base layer coding parameters are sufficiently similar to one another. (When using a traverse that follows increasing levels of roughness) it is no longer applied. Or (When using a traverse that follows decreasing levels of roughness) it happens to be applied first. In a binary manner, all possible subdivisions are tested.

[0075] In the descriptions of Figures 14 and 15a and 15b, the broad term "base layer coding parameters" is used in preferred embodiments, but these base layer coding parameters represent base layer prediction parameters, i.e., parameters that relate to the formation of predictions of the base layer signal but not to the formation of prediction residuals. Thus, for example, base layer coding parameters include prediction modes that distinguish between spatial predictions (prediction parameters for blocks / parts of the base layer signal assigned to spatial predictions such as angular direction) and temporal predictions (prediction parameters for blocks / parts of the base layer signal assigned to temporal predictions such as motion parameters).

[0076] However, within a given subblock, the definition of "sufficient" similarity of base layer coding parameters merely determines / defines a subset of base layer coding parameters. For example, similarity is determined based solely on the prediction mode. Alternatively, prediction parameters that further adjust spatial and / or temporal predictions form the parameters on which the similarity of base layer coding parameters within a given subblock depends.

[0077] Furthermore, as already outlined, in order for them to be sufficiently similar to each other, the base layer coding parameters within a given subblock must be exactly equal to each other within each subblock. Alternatively, the magnitude of the similarity used must be within a given interval in order to satisfy the "similarity" criterion.

[0078] As outlined above, the subdivision of the selected subblock is not merely the amount predicted or transferred from the base layer signal. Rather, the base layer coding parameters themselves are transferred to the enhancement layer signal, based on them, in order to obtain enhancement layer coding parameters for the subblocks of subblock subdivision obtained by transferring the selected subblock subdivision from the base layer signal to the enhancement layer signal. As far as motion parameters are concerned, for example, scaling is used to account for the transfer from the base layer to the enhancement layer. Preferably, only those parts or syntactic elements of the base layer prediction parameters are used to set the subblocks of the current portion of subblock subdivision obtained from the base layer, which affect the magnitude of similarity. The fact that these syntactic elements of the prediction parameters in each subblock of the selected subblock subdivision are somehow similar to one another ensures that the syntactic elements of the base layer prediction parameters used to predict the corresponding prediction parameters of the subblocks of the current portion 308 are similar, or even equal to one another. As a result, in the first case, which allows for some variation, some important "meanings" of the syntactic elements of the base layer prediction parameters corresponding to the portion of the base layer signal covered by each subblock are used as predictors for the corresponding subblocks. However, sometimes only the portion of the syntactic elements that contribute to the magnitude of similarity are also used to predict the prediction parameters of the subblocks of the subdivision in the enhancement layer, by adding only the transcription of the subdivision itself so as to infer or pre-set the mode of the current portion 28 subblock, although the mode-specific base layer prediction parameters participate in determining the magnitude of similarity.

[0079] One such possibility, which does not use only subdivided interlayer prediction from the base layer to the enhancement layer, is now illustrated with respect to the following figure (Figure 16). Figure 16 shows image 22b of the enhancement layer signal 400 and image 22a of the base layer signal 200 in the manner registered along the provided time axis 24.

[0080] According to the embodiment in Figure 16, the base layer signal 200 is predictively reconstructed by a scalable video decoder by subdividing the image 22a of the base layer signal 200 into intra-blocks and inter-blocks, and then predictively encoded using a scalable video encoder. According to the example in Figure 16, the latter subdivision is performed in a two-stage method. First, the frame 22a is normally subdivided into maximum blocks or maximum coding units, indicated by reference numeral 302 in Figure 16, using double lines along its periphery. Then, each maximum block 302 is subordinated to a subdivision of a hierarchical quadtree within the coding units that form the aforementioned intra-blocks and inter-blocks. As a result, they are the leaves of the quadtree subdivision of the maximum blocks 302. In Figure 16, reference numeral 304 is used to indicate these leaf blocks or coding units. Typically, solid lines are used to indicate the periphery of these coding units. Spatial intra-prediction is used for intra-blocks, while temporal inter-prediction is used for inter-blocks. The prediction parameters related to spatial intra-prediction and temporal inter-prediction are set within units of smaller blocks, respectively. However, intra and inter-blocks or coding units 304 are subdivided. Such subdivisions are illustrated in Figure 16 with respect to one of the coding units 304, using reference numeral 306 to indicate smaller blocks. The outlines of the smaller blocks 304 are formed using dotted lines. That is, in the embodiment of Figure 16, the spatial video encoder has the opportunity to choose between spatial prediction and temporal prediction with respect to each coding unit 304 of the base layer. However, with respect to the enhancement layer signal, the degrees of freedom increase. Here, in particular, frame 22b of the enhancement layer signal 400 is assigned to one of each of a set of prediction modes, which include not only spatial intra-prediction and temporal inter-prediction but also inter-layer prediction, as outlined below in detail, within the coding unit into which frame 22b of the enhancement layer signal 400 is subdivided.The subdivision within these coding units is done in the same manner as described for the base layer signals. First, frame 22b is normally subdivided into rows and columns of the largest block whose outline is formed, using regular solid lines within the coding units whose outline is formed, and subdivision using double lines within the coding units whose outline is formed.

[0081] One coding unit 308 in the current image 22b of the enhancement layer signal 400 is inferred to be assigned to the interlayer prediction mode, as exemplary, and is indicated using diagonal lines. In a manner similar to Figures 14, 15a and 15b, Figure 16 shows how the subdivision of coding unit 308 is predictively obtained by local transport from the base layer signal at 312. In particular, the local region superimposed by coding unit 308 is shown in 312. Within this region, dotted lines indicate boundaries between adjacent blocks of the base layer signal, or more generally, boundaries that are likely to change via the base layer coding parameters of the base layer. Consequently, these boundaries are the boundaries of the prediction block 306 of the base layer signal 200, and partially coincide with the boundaries between adjacent coding units 304 or equally adjacent maximum coding units 302 of the base layer signal 200, respectively. The dotted lines in 312 indicate the subdivision of the current coding unit 308 in the prediction block derived / selected by local transport from the base layer signal 200. Details regarding local transport were mentioned above.

[0082] As previously described, according to the embodiment in Figure 16, only the subdivisions within the prediction block are not adopted from the base layer. Rather, the prediction parameters of the base layer signal used in region 312 are used to obtain prediction parameters that should be used to perform predictions with respect to the prediction block of the encoding unit 308 of the enhancement layer signal 400.

[0083] In particular, according to the embodiment of Figure 16, the subdivisions within the prediction blocks are not only obtained from the base layer signal, but the prediction modes are also used in the base layer signal 200 to encode / reconstruct each region locally covered by each subblock of the obtained subdivision. One example is as follows: To obtain the subdivision of the encoding unit 308 as described above, the prediction modes are used in conjunction with the relevant base layer signal 200. Mode-specific prediction parameters are used to determine the "similarity" discussed above. Thus, the different shaded areas shown in Figure 16 correspond to different prediction blocks 306 in the base layer. Each of the different prediction blocks 306 has an intra or inter-prediction mode, i.e., a spatial or temporal prediction mode related to them. As described above, in order to be "sufficiently similar", the prediction modes used in the regions jointly arranged in each subblock of the subdivision of the encoding unit 308, and the specific prediction parameters for each prediction mode in the sub-area, must be exactly equal to each other. Alternatively, some variation may be tolerated.

[0084] In particular, according to the embodiment of Figure 16, all blocks indicated by diagonal lines extending from the upper left to the lower right are set as intra-prediction blocks of the coding unit 308 because the locally corresponding portions of the base layer signal are covered by prediction blocks 306 having spatial intra-prediction modes related to them. On the other hand, the other blocks (i.e., blocks indicated by diagonal lines extending from the lower left to the upper right) are set as inter-prediction blocks because the locally corresponding portions of the base layer signal are covered by prediction blocks 306 having temporal inter-prediction modes related to them.

[0085] On the other hand, with respect to an alternative embodiment, the derivation of the predictions is limited to the details for performing the predictions within the coding unit 308. That is, the derivation of the subdivision of coding unit 308 within the prediction block and the allocation of these prediction blocks within prediction blocks coded using non-temporal or spatial predictions and prediction blocks coded using temporal predictions are limited. It does not follow the embodiment of Figure 16.

[0086] In the latter embodiment, all prediction blocks of the coding unit 308 having assigned non-time prediction modes undergo non-time predictions, such as spatial intra-predictions, while using prediction parameters derived from the prediction parameters of locally matching intra-blocks of the base layer signal 200 as enhancement layer prediction parameters for these non-time mode blocks. As a result, such derivations relate to the spatial prediction parameters of locally co-located intra-blocks of the base layer signal 200. For example, such spatial prediction parameters may be angular directions along which the spatial prediction is performed. As outlined above, a definition of similarity by itself is required, either by superimposing that the spatial base layer prediction parameters are the same for each non-time prediction block of the coding unit 308, or by superimposing that, with respect to each non-time prediction block of the coding unit 308, the average of the spatial base layer prediction parameters is used by each non-time prediction block to derive its own prediction parameters.

[0087] Alternatively, all prediction blocks of coding unit 308 that have an assigned non-time prediction mode perform inter-layer prediction in the following manner: First, the base layer signal is decomposed or quality-enhanced to obtain the inter-layer prediction signal in the region where at least the non-time prediction mode prediction blocks of coding unit 308 are spatially co-located. Then, these prediction blocks of coding unit 308 are predicted using the inter-layer prediction signal.

[0088] The scalable video decoder and encoder, by default, cause all encoding units 308 to undergo either spatial prediction or interlayer prediction. Alternatively, the scalable video encoder / decoder supports both alternatives and indicates them in the encoded video data stream signal. The version used is limited to the non-temporal prediction mode prediction block of encoding unit 308. In particular, the decision in either alternative is indicated in the data stream, for example, individually for any size of encoding unit 308.

[0089] As far as another prediction block of the coding unit 308 is concerned, the coding unit 308 undergoes time inter-prediction using prediction parameters derived from locally matching inter-block prediction parameters, just as it is a non-time prediction mode prediction block. The resulting derivation relates, in turn, to motion vectors assigned to the corresponding portions of the base layer signal.

[0090] For all other coding units that have both a spatial intra-prediction mode and a temporal inter-prediction mode assigned to them, each coding unit receives spatial or temporal predictions in the following manner: In particular, each coding unit is further subdivided into prediction blocks that have the prediction mode assigned to it. The prediction mode is common to all prediction blocks within a coding unit, and in particular, is the same prediction mode assigned to each coding unit. That is, coding units that are different from coding units such as coding unit 308 and have a spatial intra-prediction mode related to it, or a temporal inter-prediction mode related to it, are subdivided into prediction blocks of the same prediction mode. In other words, the prediction modes are inherited only from each coding unit derived from the subdivision of each coding unit.

[0091] The subdivision of all coding units, including 308, is the subdivision of a quadtree within a prediction block.

[0092] A further difference between an interlayer prediction mode coding unit, such as coding unit 308, and a spatial intra-prediction mode or temporal inter-prediction mode coding unit is when the prediction blocks of the spatial intra-prediction mode coding unit or the temporal inter-prediction mode coding unit are subjected to spatial and temporal predictions, respectively. The prediction parameters are set independently of the base layer signal 200, for example, by being shown in the enhancement layer substream 6b. Even the subdivisions of other coding units that have an interlayer prediction mode related to them, such as coding unit 308, are shown in the enhancement layer signal 6b. That is, interlayer prediction mode coding units such as 308 have the advantage of being suitable for the need for low bit transmission rate signaling. According to the embodiment, the mode index of coding unit 308 itself does not need to be shown in the enhancement layer substream. Optionally, another parameter is transmitted for coding unit 308 for individual prediction blocks, such as the prediction parameter residual. Additionally, or alternatively, predicted residuals for encoding unit 308 are transmitted / indicated in the enhancement layer substream 6b. Meanwhile, the scalable video decoder retrieves this information from the enhancement layer substream, and the scalable video encoder according to the present embodiment determines these parameters and inserts them into the enhancement layer substream 6b.

[0093] In other words, the prediction of the base layer signal 200 is made using base layer coding parameters, such that the base layer coding parameters spatially vary the base layer signal 200 within a unit of base layer block 304. Available prediction modes for the base layer include, for example, spatial and temporal predictions. The base layer coding parameters further include individual prediction parameters for prediction modes such as angular direction, insofar as it relates to the spatially predicted block 304, and motion vector, insofar as it relates to the temporally predicted block 304. The individual prediction parameters for the latter prediction modes vary the base layer signal within a unit smaller than base layer block 304, i.e., within the aforementioned prediction block 306. To satisfy the requirement outlined earlier regarding sufficient similarity, it is necessary that the prediction modes of all base layer blocks 304 that overlap in the area of ​​each possible subblock subdivision are equal to one another. Then, only each subblock subdivision is included in the selection candidate list to obtain the selected subblock subdivision. However, the requirements are even stricter. This means that the individual prediction parameters of the prediction mode of the prediction block, which overlap the common region of the subdivisions of each subblock, must also be equal to each other. In the base layer signal, for each subdivision of this subblock and each subblock of the corresponding region, only the subdivisions that satisfy this requirement are included in the selection candidate list to obtain the final selected subdivision.

[0094] In particular, as briefly outlined above, there can be differences in how the selection is performed within the set of possible subblock divisions. Refer to Figures 15c and 15d for a more detailed outline of this. Assume that set 352 encompasses all possible subblock divisions 354 of the current block 28. Naturally, Figure 15c is merely an example. The set 352 of possible or available subblock divisions of the current block 28 is known to the scalable video decoder and scalable video encoder by the initial setup, or is shown in the encoded data stream, such as a sequence of images or similar. Following the example in Figure 15c, each member of set 352, i.e., each available subblock division 354, undergoes a check 356 to check whether the region within the co-located portion 108 of the base layer signal being subdivided is simply superimposed by the prediction block 306 and the encoding unit 304 by transferring each subblock division 354 from the enhancement layer to the base layer. Then, we check whether the base layer coding parameters satisfy the requirement of sufficient similarity. See, for example, the exemplary subdivision with reference number 354. According to this exemplary available subdivision of subblocks, the current block 28 is subdivided into four quadrants / subblocks 358. The upper left subblock corresponds to region 362 in the base layer. Clearly, this region 362 overlaps with the four blocks of the base layer, namely two prediction blocks 306 and two coding units 304, which represent the prediction block itself, rather than another subdivision within the prediction block. Therefore, if all base layer coding parameters of these prediction blocks overlapping region 362 satisfy the similarity criterion, and this is also true for all subblocks / quaternions of a possible subblock subdivision 354 and their corresponding regions overlapping base layer coding parameters, then this possible subblock subdivision 354 belongs to the set of subblock subdivisions 364, satisfying the requirements for all regions covered by the subblocks of each subblock subdivision.Then, from this set 364, the coarsest subdivision is selected as indicated by arrow 366, and as a result, from set 352, a subdivision 368 of the selected subblock is obtained.

[0095] Clearly, it is preferable to try to avoid performing check 356 on all members of set 352. Thus, as shown in Figure 15d and as previously mentioned, possible subdivisions 354 are traversed to increase or decrease in size. The traversal is indicated using double-headed arrows 372. Figure 15d shows that, with respect to at least some of the available subdivisions of the subblock, the level of size or size is equal to one another. In other words, the ordering according to the level of increasing or decreasing size is ambiguous. However, this does not prevent the search for the “largest subdivision of the subblock” belonging to set 364, since only one of such equally large possible subdivisions of the subblock belongs to set 364. Thus, the largest possible subdivision of the subblock 368 is found as soon as the result of criterion check 356 changes from filled to unfilled, when traversing in the direction of increasing levels of size in the possible subdivisions of the subblock 354 traversed from second to last, which is the subdivision of the subblock to be selected. Alternatively, if the sub-block subdivision 368 is mostly a recently traversed sub-block subdivision and is traversed in the direction of decreasing levels, the result of the criterion check 356 switches from not filled to filled.

[0096] With respect to the following figures, a scalable video encoder or decoder as described above with respect to Figures 1 to 4 is implemented to form an embodiment of the present application according to another aspect of the application. Possible embodiments of the embodiments described below are presented with reference to Examples K, A, and M.

[0097] To illustrate the embodiment, refer to Figure 17. Figure 17 shows the possibilities for the current portion 28 time prediction 32. As a result, the following description of Figure 17 is combined with the description of Figures 6 to 10 as far as it relates to the combination with the interlayer prediction signal, or with the description of Figures 11 to 13 as far as it relates to the combination with the time interlayer prediction mode.

[0098] The situation shown in Figure 17 corresponds to the situation shown in Figure 6. That is, the base layer signal 200 and the enhancement layer signal 400 are shown together with the already encoded / decoded portion, indicated by the diagonal lines. Within the enhancement layer signal 400, the portion to be encoded / decoded now has adjacent blocks 92 and 94, which are described here exemplary as block 92 above and 94 to the left of the current portion 28. Both blocks 92 and 94 are exemplary the same size as the current block 28. However, size matching is not mandatory. Rather, the portion of the block in image 22b of the subdivided enhancement layer signal 400 can have different sizes. They are not even limited to rectangles. They may be rectangular or other shapes. The current block 28 with another adjacent block is not explicitly shown in Figure 17. However, the other adjacent block has not yet been encoded / decoded. That is, they follow the encoding / decoding order and are therefore not available for prediction. In addition to these, there are other blocks besides blocks 92 and 94 that have already been coded / decoded according to the coding / decoding order, for example, block 96 adjacent to the current block 28, diagonally to the upper left of the current block 28. However, blocks 92 and 94 have predetermined adjacent blocks that, in the example considered here, play the role of predicting the inter-prediction parameters for the current block 28 that will receive the inter-prediction 30. The number of such predetermined adjacent blocks is not limited to two; it may be greater than one or simply one. A discussion of possible embodiments is presented with respect to Figures 36-38.

[0099] The scalable video encoder and scalable video decoder determine a set of predetermined adjacent blocks, here, blocks 92, 94, from a set of already encoded adjacent blocks, here, blocks 92-96, within the current portion 28, such as the upper left sample, depending on a predetermined sample position 98. For example, only those already encoded adjacent blocks of the current portion 28 form a set of "predetermined adjacent blocks" that include a sample position directly adjacent to the predetermined sample position 98. Another possibility is described with respect to Figures 36-38.

[0100] In any case, according to the decoding / encoding order, the previously encoded / decoded portion 502 of the enhancement layer signal 400 of the image 22b, which has been replaced by the motion vector 504 from the co-located position of the current block 28, contains sample values ​​reconstructed based on the sample values ​​of portion 28 predicted by mere copying, interpolation, etc. For this purpose, the motion vector 504 is shown in the enhancement layer substream 6b. For example, the time prediction parameter for the current block 28 shows a displacement vector 506 indicating the replacement of portion 502 from the co-located position of portion 28 in the reference image 22b by arbitrary interpolation to be copied over the samples of portion 28.

[0101] In any case, when predicting the current block 28 in time, the scalable video decoder / encoder has already reconstructed (and, in the case of an encoder, encoded) the base layer 200 using the base layer substream 6a. Block prediction is used as described above, as long as the relevant spatially corresponding regions of the temporally corresponding image 22a are thus relevant, and a block selection between, for example, spatial prediction mode and temporal prediction mode is used.

[0102] In Figure 17, the time interval of the image 22a of the base layer signal 200 is subdivided into several blocks 104. Blocks 104 are located around a region that locally corresponds to the exemplarily represented current portion 28, just as it is when the enhancement layer signal 400 has spatially predicted blocks. The spatial prediction parameters are included in or shown within the base layer substream 6a with respect to those blocks 104 in the base layer signal 200. The selection of the spatial prediction mode is shown with respect to the base layer signal 200.

[0103] Here, exemplarily, for a time-intra-layer prediction 32 with respect to a selected block 28, inter-prediction parameters such as motion parameters are determined using one of the following methods to enable the reconstruction of the enhancement layer signal from the encoded data stream.

[0104] The first possibility is explained with respect to Figure 18. Specifically, first, a set 512 of motion parameter candidates 514 is gathered or generated from pre-determined, already reconfigured blocks of a frame, such as blocks 92 and 94. Motion parameters are motion vectors. The motion vectors of blocks 92 and 94 are represented using arrows 516 and 518, marked with 1 and 2 respectively. As shown, these motion parameters 516 and 518 directly form candidate 514. Some candidates are formed by combining motion vectors such as 518 and 516, as shown in Figure 18.

[0105] Furthermore, a set of one or more base layer motion parameters 524 522 of a block 108 of the base layer signal 200, which is jointly arranged in section 28, is collected from or generated from the base layer motion parameters. In other words, motion parameters related to a block 108 jointly arranged in the base layer are used to obtain one or more base layer motion parameters 524.

[0106] At that time, one or more base layer motion parameters 524, or scaled versions thereof, are added 526 to the set of motion parameter candidates 514 512 in order to obtain an extended set of motion parameter candidate sets 528 of motion parameter candidates. This can be done in a variety of ways, such as simply adding the base layer motion parameters 524 at the end of the list of candidates 514, or in different ways, one example of which is outlined with respect to Figure 19a.

[0107] At least one of the motion parameter candidates 532 from the extended motion parameter candidate set 528 is then selected. The motion compensation prediction of part 28 is performed using the selected motion parameter candidate from the extended motion parameter candidate set. The selection 534 is performed by the method of index 536 in the list / set 528, as illustrated in a data stream such as substream 6b for part 28, or by another method described with respect to Figure 19a.

[0108] As mentioned above, it is checked whether the base layer motion parameter 523 has been encoded in an encoded data stream such as the base layer substream 6a using integration. If the base layer motion parameter 523 has been encoded in an encoded data stream using integration, then the summation 526 is suppressed.

[0109] The motion parameters described in Figure 18 relate to the complete set of motion parameters, including the number of motion hypotheses, either for motion vectors (motion vector predictions) only, or for each block, reference index, or dividing information (integration). Thus, the "scaled version" in the case of spatial scalability derives from the scaling of the motion parameters used in the base layer signal according to the spatial resolution ratio between the base layer signal and the enhancement layer signal. Depending on the method of encoding the data stream, the encoding / decoding of the base layer motion parameters of the base layer signal relates to motion vector predictions, such as spatially or temporally, or integration.

[0110] The incorporation 526 of motion parameters 523 used in the co-arranged portion 108 of the base layer signal within the set 528 of integrated / motion vector candidates 532 enables a highly capable index among intra-layer candidates 514 and one or more inter-layer candidates 524. The selection 534 involves an explicit indication of the index in the expanded set / list of motion parameter candidates in the enhancement layer signal 6b for each prediction block, each coding unit or similar. Alternatively, the selection index 536 may be inferred from other information in the enhancement layer signal 6b, or from information in the inter-layer.

[0111] As per the possibility in Figure 19a, the formation of the final motion parameter candidate list 542 for the enhancement layer signal for part 28 is performed only arbitrarily, as outlined with respect to Figure 18. That is, formation 542 may be 528 or 512. However, lists 528 / 512 are ordered 544 depending on base layer motion parameters, such as motion parameters represented by motion vectors 523 of a co-arranged base layer block 108. For example, the rank of a member, i.e., motion parameter candidate 532 or 514 of list 528 / 512, is determined based on the deviation of each member with respect to a potentially scaled version of motion parameter 523. The larger the deviation, the lower the rank of each member 532 / 512 in the ordered list 528 / 512'. Consequently, ordering 544 involves determining the magnitude of the deviation for each member 532 / 514 of list 528 / 512. The selection 534 of one candidate 532 / 512 from the ordered list 528 / 512' is performed and controlled via the explicitly shown index syntax element 536 in the encoded data stream to obtain an enhancement layer motion parameter from the ordered motion parameter candidate list 528 / 512' with respect to a portion 28 of the enhancement layer signal. Then, the time prediction 32 is performed by motion compensation prediction of a portion 28 of the enhancement layer signal using the selected motion parameter, where index 536 points to 534.

[0112] With respect to the motion parameters mentioned in Figure 19a, the motion parameters described above with respect to Figure 18 are applied. Decoding of the base layer motion parameters 520 from the encoded data stream involves (optionally) spatial or temporal motion vector prediction or integration. Ordering is done according to the magnitude of the difference between each enhancement layer motion parameter candidate and the base layer motion parameter of the base layer signal, with respect to the block of base layer signal co-located with the current block of enhancement layer signal. That is, with respect to the current block of enhancement layer signal, a list of enhancement layer motion parameter candidates is first determined. Next, ordering is performed, as is explained below. The selection is performed by obvious signaling.

[0113] Alternatively, the ordering 544 is made according to the magnitude of the difference between the base layer motion parameter 523 of the base layer signal relating to block 108 of the base layer signal co-located in the current block 28 of the enhancement layer signal, and the base layer motion parameter 546 of the spatially and / or temporally adjacent block 548 in the base layer. The determined ordering in the base layer is then transferred to the enhancement layer. As a result, the enhancement layer motion parameter candidates are ordered with respect to the corresponding base layer candidates in the same manner as the determined ordering. In this regard, when the relevant base layer block 548 is spatially / temporally co-located with adjacent enhancement layer blocks 92 and 94 relating to the considered enhancement layer motion parameter, the base layer motion parameter 546 is said to correspond to the enhancement layer motion parameter of the adjacent enhancement layer blocks 92,94. Alternatively, when the adjacency relationship between the associated base layer block 548 and the block 108 co-located in the current enhancement layer block 28 (left neighbor, top neighbor, A1, A2, B1, B2, B0, or see Figures 36-38 for further examples) is the same as the adjacency relationship between the current enhancement layer block 28 and the enhancement layer adjacent blocks 92, 94, respectively, then the base layer motion parameter 546 is said to correspond to the enhancement layer motion parameter of the enhancement layer adjacent blocks 92, 94. Based on the base layer ordering, selection 534 is then performed by explicit signaling.

[0114] To illustrate this in more detail, refer to Figure 19b. Figure 19b outlines the first alternative for obtaining an enhancement layer ordered for a list of motion parameter candidates using base layer hints. Figure 19b shows the current block 28 and the positions of three different predetermined samples, namely, exemplary, upper left sample 581, lower left sample 583, and upper right sample 585. The examples are inserted for illustrative purposes only. Assume that a predetermined set of adjacent blocks includes, exemplary, four types of adjacent. Adjacent block 94a covers adjacent sample position 587 immediately above sample position 581. Adjacent block 94b includes or covers sample position 589, which is adjacent to sample position 585, immediately above it. Similarly, assume that adjacent blocks 92a and 92b include sample positions 591 and 593, which are adjacent to sample positions 581 and 583, immediately to the left of them. Furthermore, as explained with respect to Figures 36 to 38, note that despite the predetermined rules for determining the number, the predetermined number of adjacent blocks changes. Nevertheless, the predetermined adjacent blocks 92a, 92b, 94a, and 94b can be distinguished by their respective rules.

[0115] According to the alternative in Figure 19b, the co-located blocks in the base layer are determined for each predetermined adjacent block 92a, 92b, 94a, 94b. For example, for this purpose, the top-left sample 595 of each adjacent block is used. This is the case when the current block 28 has a top-left sample 581, which is formally mentioned in Figure 19a. This is illustrated in Figure 19b using a dotted arrow. By this means that for each predetermined adjacent block, the corresponding block 597 is found in addition to the co-located block 108, the co-located current block 28. Using the motion parameters m1, m2, m3, m4 of the co-located base layer block 597 and their respective differences to the base layer motion parameter m of the co-located base layer block 108, the enhancement layer motion parameters M1, M2, M3, M4 of the predetermined adjacent blocks 92a, 92b, 94a, 94b are ordered in list 528 or 512. For example, the greater the distance between any of m1-m4, the higher the corresponding enhancement layer motion parameters M1-M4. That is, a higher index is required to index in the same state from list 528 / 512'. The absolute difference is used with respect to the magnitude of the distance. Similarly, motion parameter candidates 532 or 514 are rearranged in the list with respect to their ranks, which are the combination of enhancement layer motion parameters M1-M4.

[0116] Figure 19c shows an alternative in which the corresponding blocks in the base layer are determined in a different way. In particular, Figure 19c shows the predetermined adjacent blocks 92a, 92b, 94a, 94b of the current block 28, and the jointly positioned block 108 of the current block 28. According to the embodiment of Figure 19c, the base layer blocks corresponding to these of the current block 28, i.e., 92a, 92b, 94a, 94b, are determined in a manner such that these base layer blocks are related to the enhancement layer adjacent blocks 92a, 92b, 94a, 94b, using the same adjacent determination rules to determine these base layer adjacent blocks. In particular, Figure 19c shows the predetermined sample positions of the jointly positioned block 108, i.e., the upper left, lower left, and upper right sample positions 601. Based on these sample locations, the four adjacent blocks of block 108 are determined in the same way as described for the enhancement layer adjacent blocks 92a, 92b, 94a, and 94b with respect to the predetermined sample locations 581, 583, and 585 of the current block 28. The four base layer adjacent blocks 603a, 603b, 605a, and 605b are found in this manner. 603a clearly corresponds to the enhancement layer adjacent block 92a. Base layer block 603b corresponds to the enhancement layer adjacent block 92b. Base layer block 605a corresponds to the enhancement layer adjacent block 94a. Base layer block 605b corresponds to the enhancement layer adjacent block 94b. In the same manner as previously described, the base layer motion parameters M1-M4 of base layer blocks 903a, 903b, 905a, 905b, and their distances to the base layer motion parameter m of the co-located base layer block 108, are used to order motion parameter candidates in lists 528 / 512 formed from the motion parameters M1-M4 of enhancement layer blocks 92a, 92b, 94a, 94b.

[0117] According to the possibilities in Figure 20, the formation of the final motion parameter candidate list 562 for the enhancement layer signal for portion 28 is simply performed arbitrarily, as outlined with respect to Figures 18 and / or 19. That is, formation 562 is 528 or 512 or 528 / 512'. Reference numeral 564 is used in Figure 20. According to Figure 20, index 566 pointed out in motion parameter candidate list 564 is determined depending on index 567 in motion parameter candidate list 568, which was used, for example, to encode / decode the base layer signal with respect to co-located block 108. For example, when reconstructing the base layer signal in block 108, the list of motion parameter candidates 568 is determined based on the motion parameters 548 of the adjacent block 548 of block 108, which has the same adjacency relationship (adjacent to the left, adjacent above, A1, A2, B1, B2, B0, or see Figures 36-38 for another example) as block 108 has the same adjacency relationship as the current block 28 with respect to predetermined adjacent enhancement layer blocks 92, 94. Here, the determination 572 of list 567 potentially uses the same configuration rules used in formation 562, such as ordering within the list members of lists 568 and 564. More generally, the index 566 for an enhancement layer is determined in such a way that its adjacent enhancement layer blocks 92, 94 are pointed to by the index 566, which is co-located with the base layer block 548 related to the indexed base layer candidate, i.e., what the index 567 points to. As a result, index 567 serves as a key prediction for index 566. The enhancement layer motion parameters are then determined using index 566 in the motion parameter candidate list 564, and the motion compensation prediction in block 28 is performed using the determined motion parameters.

[0118] The same principles described above apply to the motion parameters mentioned in Figure 20, as they do to Figures 18 and 19.

[0119] With respect to the following figures, as described above with respect to Figures 1 to 4, it is described how a scalable video encoder or decoder can be realized to form an embodiment of the present application according to another embodiment of the present application. A detailed implementation of the embodiment described below is described below with reference to Embodiment V.

[0120] This embodiment relates to residual coding in the enhancement layer. In particular, Figure 21 exemplifies the image 22b of the enhancement layer signal 400 and the image 22a of the base layer signal 200 in a time-registered manner. Figure 21 shows the enhancement layer signal and how it is reconstructed in a scalable video decoder or encoded in a scalable video encoder, focusing on a predetermined conversion coefficient block and a predetermined portion 404 of the conversion coefficient 402 representing the enhancement layer signal 400. In other words, the conversion coefficient block 402 represents the spatial decomposition of portion 404 of the enhancement layer signal 400. As already described above according to the coding / decoding ordering, the corresponding portion 406 of the base layer signal 200 is already coded / coded when the conversion coefficient block 402 is coded / coded. As far as the base layer signal 200 is concerned, the coding / decoding of the prediction is used, including the signaling of the base layer residual signal within the coded data stream, such as the base layer substream 6a.

[0121] As described in the embodiment with respect to Figure 21, the scalable video decoder / encoder takes advantage of the fact that the evaluation 408 of the base layer signal or base layer residual signal results in a favorable selection of the subdivision of the conversion coefficient block 402 in the subblock 412 in a portion 406 jointly arranged in portion 404. In particular, several possible subblock subdivisions for subdividing the conversion coefficient block 402 into subblocks are supported by the scalable video decoder / encoder. These possible subblock subdivisions subdivide the conversion coefficient block 402 in a regularly rectangular subblock 412. That is, the conversion coefficients 414 of the conversion coefficient block 402 are arranged in columns and rows, and according to the possible subblock subdivisions, these conversion coefficients 414 are regularly densely packed within the subblock 412, so that the subblock 412 itself is arranged in rows and columns. Evaluation 408 uses the subdivision of the subblock thus selected to enable setting the ratio of the number of rows to the number of columns in subblock 412, i.e., the ratio of their width to height, in such a way that the encoding of the transformation coefficient block 402 is most efficient. If, for example, evaluation 408 finds that the reconstructed base layer signal 200 in the co-located portion 406, or at least the base layer residual signal in the corresponding portion 406, consists mainly of horizontal edges in the spatial domain, then the transformation coefficient block 402 likely exists having importance, i.e., the transformation coefficient level is non-zero, i.e., the quantized transformation coefficient is near the zero horizontal frequency side of the transformation coefficient block 402. In the case of vertical edges, the transformation coefficient block 402 likely exists with a non-zero transformation coefficient level at a location near the zero vertical frequency side of the transformation coefficient block 402. Therefore, first, subblock 412 is selected such that it is longer along the vertical direction and smaller along the horizontal direction. Secondly, the subblock is made longer horizontally and smaller vertically. The latter case is schematically shown in Figure 40.

[0122] In other words, the scalable video decoder / encoder selects one subblock subdivision from a set of possible subblock subdivisions based on the base layer residual signal or base layer signal. Then, the encoding 414 or decoding of the conversion coefficient block 402 is performed while applying the selected subblock subdivision. In particular, since the positions of the conversion coefficients 414 are traversed within the units of subblock 412, all positions within one subblock are traversed in such a way that they immediately follow the next subblock in the subblock ordering defined within the subblock. For subblocks 412, such as subblock 412, which are currently visited, the syntactic element is shown in a data stream such as the enhancement layer substream 6b, which has a reference code 412 indicating whether the currently visited subblock has important conversion coefficients. In Figure 21, the syntactic element 416 is illustrated with respect to two exemplary subblocks. If each syntactic element of each subblock indicates an insignificant transformation coefficient, then nothing else needs to be transmitted in the data stream or enhancement layer substream 6b. Rather, the scalable video decoder sets the transformation coefficients in that subblock to zero. However, if each subblock's syntactic element 416 indicates that this subblock has an important transformation coefficient, then other information related to the transformation coefficients in that subblock is indicated in the data stream or substream 6b. On the decoding side, the scalable video decoder decodes from the data stream or substream 6b a syntactic element 418 indicating the level of the transformation coefficients in each subblock. The syntactic element 418 indicates the scan order of these transformation coefficients in each subblock, and optionally, the position of the important transformation coefficients in that subblock according to the scan order of the transformation coefficients in each subblock.

[0123] Figure 22 illustrates the different possibilities that exist for performing a selection among the possible sub-block subdivisions in evaluation 408. Figure 22 again describes a portion 404 of the enhancement layer signal, where the transformation coefficient block 402 is related to representing the spectral decomposition of portion 404. For example, the transformation coefficient block 402 represents the spectral decomposition of the enhancement layer residual signal, with a scalable video decoder / encoder predictively encoding / decoding the enhancement layer signal. In particular, the encoding / decoding of the transformation is used by the scalable video decoder / encoder to encode the enhancement layer residual signal. The encoding / decoding of the transformation is performed in a block-like manner, i.e., in blocks into which the image 22b of the enhancement layer signal is subdivided. Figure 22 shows the corresponding or co-arranged portions 406 of the base layer signal. Here, the scalable video decoder / encoder applies predictive encoding / decoding to the base layer signal with respect to the predicted residuals of the base layer signal, i.e., with respect to the base layer residual signal, while using transform encoding / decoding. In particular, block transformation is used with respect to the base layer residual signal. That is, the base layer residual signal is transformed in blocks, which are individually transformed blocks indicated by dotted lines in Figure 22. As shown in Figure 22, the block boundaries of the base layer transformed blocks do not need to coincide with the outline of the co-located portion 406.

[0124] Nevertheless, to perform evaluation 408, one or a combination of the following options A-C can be used.

[0125] In particular, the scalable video decoder / encoder performs a transformation 422 in part 406 to the base layer residual signal or reconstructed base layer signal in order to obtain a transformation coefficient block 424 of transformation coefficients that are sized to match the transformation coefficient block 402 to be encoded / decoded. Checking the distribution of transformation coefficient values ​​in transformation coefficient blocks 424, 426 is used to appropriately set the dimensions of subblock 412 along the horizontal frequency direction i.e. 428 and along the vertical frequency direction i.e. 432.

[0126] In addition, or alternatively, the scalable video decoder / encoder inspects all the conversion coefficient blocks of the base layer conversion block 434, indicated by the different shaded areas in Figure 22, partially overlapping at least the jointly arranged portion 406. In the exemplary case of Figure 22, there are four base layer conversion blocks, and then their conversion coefficient blocks are inspected. In particular, all of these base layer conversion blocks are of different sizes from each other, and further different in size with respect to the conversion coefficient block 412. Scaling 436 is performed with respect to these conversion coefficient blocks overlapping the base layer conversion block 434 to yield an approximation of the conversion coefficient block 438 of the spectral decomposition of the base layer residual signal in portion 406. The distribution of the conversion coefficient values ​​in that conversion coefficient block 438, i.e., 442, is used in evaluation 408 to appropriately set the dimensions 428 and 432 of the subblocks. As a result, the subdivision of the conversion coefficient block 402 into subblocks is selected.

[0127] Further alternatives that can be used additionally or as an alternative to perform evaluation 408 are to examine the base layer residual signal or reconstructed base layer signal using edge detection 444 or determination of the main gradient direction in the spatial domain. For example, determine the extension direction of the detected edge, or within the co-located portion 406, based on the determined gradient, and appropriately set the subblock dimensions 428 and 432.

[0128] Although not explicitly stated above, when traversing the positions of the conversion coefficients and the units of subblock 412, it is preferable to traverse subblock 412 in the order starting from the zero-frequency angle of the conversion coefficient block (upper left corner of Figure 21) and reaching the highest-frequency angle of block 402 (lower right corner of Figure 21). Furthermore, entropy coding is used to represent the syntactic elements in the data stream 6b. That is, syntactic elements 416 and 418 are coded convenient entropy codes, such as arithmetic, variable-length coding, or other forms of entropy coding. The ordering of traversing subblock 412 also depends on the subblock shape selected according to 408. With respect to subblocks selected to be wider than their height, the order of traversing is to traverse the subblock column by column first, then to the next column, and so on. Beyond this, it should be noted again that the base layer information used to select the dimensions of the subblocks is the self-reconstructed base layer residual signal or the base layer signal itself.

[0129] The following describes different embodiments that can be combined with the embodiments described above. The embodiments described below relate to many different embodiments or measures for making scalable video encoding more efficient. Partially, the embodiments described above are described in detail below, retaining the general concepts, in order to present another derived embodiment of it. These descriptions presented below are used to obtain alternatives or extensions of the embodiments / examples described above. However, most of the embodiments described below relate to sub-embodiments that can be optionally combined with embodiments already described above; that is, they can be implemented together with the embodiments above in a single scalable video decoder / encoder. However, this is not required.

[0130] To facilitate understanding of the previous description, more detailed embodiments for realizing a suitable scalable video encoder / decoder incorporating embodiments and combinations of embodiments are presented below. The different embodiments described below are enumerated using alphanumeric symbols. These embodiments generally perform, here according to one embodiment, some of the descriptions of the reference elements of these embodiments in the drawings described above. However, as far as individual embodiments are concerned, the provision of elements in the realization of the scalable video decoder / encoder is not necessary as far as every embodiment is concerned. Depending on the embodiment in question, some elements and some interconnections are omitted in the drawings described below. Only the elements referenced in each embodiment are provided to perform the work or function mentioned in the description of each embodiment. However, here, when some elements are referenced in relation to a function, there are sometimes alternatives in particular.

[0131] However, in order to provide an overview of the functionality of the scalable video decoder / encoder, the following embodiment is performed. The elements shown in the following figure are now briefly explained.

[0132] Figure 23 shows a scalable video decoder for decoding an encoded data stream 6 in which video is encoded, such that the main sub-part of the encoded data stream 6 (i.e., 6a) represents video at a first resolution or quality level. An additional part 6b of the encoded data stream corresponds to the representation of video at an increasing resolution or quality level. To keep the data volume of the encoded data stream 6 low, redundancy in the interlayer between substreams 6a and 6b is utilized when forming substream 6b. Some of the embodiments described below are directed from the base layer to the interlayer prediction, with substream 6a involved, and then to the enhancement layer, with substream 6b involved.

[0133] The scalable video decoder includes two block-based predictive decoders 80 and 60 operating in parallel, receiving substreams 6a and 6b, respectively. As shown in the figure, the demultiplexer 40 provides the decoding stages 80 and 60 separately, along with the corresponding substreams 6a and 6b.

[0134] The intrastructure of the block-based predictive coding stages 80 and 60 is the same, as shown in the diagram. For each input of the decoding stages 80 and 60, the entropy decoding modules 100;320, the inverse converters 560;580, the adders 180;340, and the arbitrary filters 120;300 and 140;280 are connected in series in the order described. As a result, at the end of this series connection, the reconstructed base layer signal 600 and the reconstructed enhancement layer signal 360 are obtained, respectively. Meanwhile, the outputs of the adders 180,340 and the filters 120,140,300,280 provide different versions of the reconstructed base layer and enhancement layer signals, respectively. Then, each predictive provider 160;260 receives a subset or all of these versions and, based on that, provides the predictive signal to the residual input of the adder 180;340. The entropy decoding stage 100;320 decodes from the respective input signals 6a and 6b, and the converted coefficient block enters the inverse transducer 560;580, which encodes parameters including the prediction parameters for the prediction provider 160;260.

[0135] Therefore, prediction providers 160 and 260 predict blocks of video frames at their respective resolution / quality levels. For this purpose, prediction providers 160 and 260 are selected from among predetermined prediction modes, such as spatial intra-prediction mode and temporal intra-prediction mode. Both modes are intra-layer prediction modes, i.e., prediction modes that depend solely on the data in the substream in which each level resides.

[0136] However, to take advantage of the aforementioned interlayer redundancy, the enhancement layer decoding stage 60 additionally includes an encoded parameter interlayer predictor 240, a resolution / quality improver 220, and / or a prediction provider 260 compared with the prediction provider 160. Furthermore, / or also, the enhancement layer decoding stage 80 supports an interlayer prediction mode that can provide an enhancement layer prediction signal 420 based on data obtained from the intra-stage of the base layer decoding stage 80. The resolution / quality improver 220 improves the resolution or quality of one of the reconstructed base layer signals 200a, 200b, 200c or the base layer residual signal 480 in order to obtain an interlayer prediction signal 380. The encoded parameter interlayer predictor 240 is responsible for predicting encoded parameters such as prediction parameters and motion parameters, respectively. The prediction provider 260 further supports the interlayer prediction mode according to the reconstructed portion of the base layer signal, such as 200a, 200b, 200c. Alternatively, a reconfigured portion of the base layer residual signal 640, potentially improved to an increased resolution / quality level, can be used as a reference / foundation.

[0137] As previously mentioned, decoding stages 60 and 80 are operated in a block-based manner. That is, frames of video are subdivided into block-like parts. Different coarseness levels are used to assign prediction modes performed by prediction providers 160, 260, local transformations by inverse converters 560, 580, filter coefficient selection by filters 120, 140, and prediction parameter settings for the prediction modes by prediction providers 160, 260. In other words, subpartitioning a frame into prediction blocks is, in turn, a sequence of subpartitionings of the frame into blocks, e.g., so-called coding units or prediction units, from which a prediction mode has been selected. Subpartitioning a frame into blocks for transformation coding, so-called transformation units, differs from subpartitioning into prediction units. Some of the interlayer prediction modes used by prediction provider 260 are described below with reference to embodiments. The prediction provider 260 is applied based on several intra-layer prediction modes, i.e., prediction modes that intra-obtain each prediction signal input to each adder 180,340, i.e., each based solely on the state related to the current level coding stages 60,80.

[0138] Some further details of the block shown in the figure will become apparent from the descriptions of the individual embodiments below. Note that these descriptions are equally and generally reproducible to other embodiments and figure descriptions unless such descriptions are explicitly related to the embodiments provided.

[0139] In particular, the embodiment for the scalable video decoder in Figure 23 represents a possible realization of the scalable video decoder according to Figures 2 and 4. Although the scalable video decoder according to Figure 23 has been described previously, Figure 23 shows the corresponding scalable video encoder, and the same reference codes are used for the intra elements of the predictive coding / decoding scheme in Figures 23 and 24. The reasons are as previously stated. Also, for the purpose of maintaining a general predictive basis between the encoder and decoder, reconfigurable versions of the base and enhancement layer signals are used in the encoder and reconfigure the already coded portion up to this end to obtain a reconfigurable version of the scalable video. Thus, the only difference from the description in Figure 23 is that, as with the coding parameter inter-layer predictor 240, the predictive providers 160 and 260 determine the predictive parameters in some ratio / distortion optimization process rather than receiving the predictive parameters from the data stream. Rather, the providers transmit the predictive parameters thus determined to the entropy decoders 19a and 19b. The entropy decoders 19a and 19b sequentially transmit their respective base layer substreams 6a and enhancement layer substreams 6b via the multiplexer 16 for inclusion in the data stream 6. In a similar manner, these entropy encoders 19a and 19b receive the reconstructed base layer signal 200 and the reconstructed enhancement layer signal 400 and the predicted residuals between the original base layer and enhancement layer versions 4a and 4b, rather than outputting the entropy decoding result of such residuals, so that the conversion modules 724 and 726 obtain them via the subsequent subtractors 720 and 722. However, the structure of the scalable video encoder in Figure 24 is consistent with the structure of the scalable video decoder in Figure 23. Therefore, with respect to these issues, reference is made with respect to the description above in Figure 23. Here, any part that states any derivation from any data stream, as just outlined, must be changed to the respective decisions of each element having subsequent insertion into each data stream.

[0140] The techniques for intra-coding enhancement layer signals used in the embodiments described below include multiplexing methods for generating intra-prediction signals (using base layer data) for enhancement layer blocks. These methods are provided in addition to methods for generating intra-prediction signals based solely on samples of the reconstructed enhancement layer.

[0141] Intra-prediction is part of the process of reconstructing the intra-encoded block. The final reconstructed block is obtained by adding the residual signal (which may be zero) encoded in the transform to the intra-prediction signal. The residual signal is generated by the inverse quantization (scaling) of the transform coefficient levels transmitted in the bitstream that followed the inverse transform.

[0142] The following description applies to scalable coding with a quality enhancement layer (representing input video with the same resolution as the base layer but with higher quality or fidelity) and scalable coding with a spatial enhancement layer (representing higher resolution, i.e., more samples, than the base layer). In the case of a quality enhancement layer, upsampling of the base layer signal is not required in blocks such as 220, but is applied to filtering of the reconstructed base layer samples, such as 500. In the case of a spatial enhancement layer, upsampling of the base layer signal is generally required in blocks such as 220.

[0143] The following examples support different methods for using reconstructed base layer samples (e.g., 200) or base layer residual samples (e.g., 640) for intra-prediction of enhancement layer blocks. One or more of the methods described below can be supported in addition to intra-layer intra-coding (where only reconstructed enhancement layer samples (e.g., 400) are used for intra-prediction). The use of a particular method is indicated at the level of the largest supported block size (such as the size of a block in H.264 / AVC within HEVC or a macroblock in the coding tree / maximum coding unit). Or it is indicated for all supported block sizes. Or it is indicated with respect to a subset of supported block sizes.

[0144] For all methods described below, the prediction signal is used directly as the reconstruction signal for the block; that is, the residual is not transmitted at all. Alternatively, the selected method for interlayer intra-prediction is combined with residual coding. In certain embodiments, the residual signal is transmitted via transform coding; that is, the quantized transform coefficients (transform coefficient levels) are transmitted using entropy coding techniques (e.g., variable-length coding or arithmetic coding (example 19b)). The residual is then obtained by inverse quantization (scaling) the transmitted transform coefficient levels and applying an inverse transform (example 580). In certain versions, the complete residual block corresponding to the block from which the interlayer intra-prediction signal originates is transformed using a single transform (example 726); that is, the entire block is transformed using a single transform of the same size as the prediction block. In another embodiment, the prediction block is further subdivided into smaller blocks (e.g., using hierarchical decomposition); and a separate transform is applied to each of the smaller blocks (which may also have different block sizes). In another embodiment, the coding unit is divided into smaller prediction blocks. Then, for zero or more prediction blocks, the prediction signal is generated using one of the methods for interlayer intra-prediction. Then, the residual of the entire coding unit is transformed using a single transformation (example: 726). Alternatively, the coding unit is subdivided into different transformation units, where the subdivision to form transformation units (blocks to which a single transformation is applied) is different from the subdivision to decompose the coding unit into prediction blocks.

[0145] In certain embodiments, the reconstructed base layer signal (e.g., 380), which is upsampled / filtered, is used directly as the prediction signal. Multiplexing methods for using the base layer to intra-predict the enhancement layer include the following: The reconstructed base layer signal (e.g., 380), which is upsampled / filtered, is used directly as the enhancement layer prediction signal. This method is similar to the well-known H.264 / SVC inter-layer intra-prediction mode. In this method, the prediction block for the enhancement layer is formed by co-located samples of the reconstructed base layer signal, which is upsampled (e.g., 220) to match the corresponding sample position in the enhancement layer and optionally filtered before or after upsampling. In contrast to the SVC inter-layer intra-prediction mode, this mode is supported at any block size, as well as at the macroblock level (or the maximum supported block size). This means that the mode is not only indicated with respect to the largest supported block size, but that blocks of the largest supported block size (macroblocks in MPEG4 and H.264, and coded tree blocks / largest coded units in HEVC) are hierarchically subdivided into smaller blocks / coded units, and the use of inter-layer intra-predictive mode is indicated for any supported block size (with respect to the corresponding block). In certain embodiments, this mode is supported only for a selected block size. Then, a syntactic element indicating the use of this mode is sent only with respect to the corresponding block size. Or, the value of a syntactic element indicating the use of this mode (in another coding parameter) is correspondingly restricted with respect to a different block size. Another difference from the inter-layer intra-predictive mode in the SVC extension of H.264 / AVC is that the inter-layer intra-predictive mode is supported not only when co-located regions in the base layer are intra-coded, but also when co-located base layer regions are inter-coded or partially inter-coded.

[0146] In a particular embodiment, spatial intra-prediction of the difference signal (see Example A) is performed. The multiplexing method includes the following: A reconstructed base layer signal (e.g., 380) (potentially upsampled / filtered) is combined with a spatial intra-prediction signal, where the spatial intra-prediction (e.g., 420) is obtained based on difference samples for adjacent blocks (e.g., 260). The difference sample represents the difference between the reconstructed enhancement layer signal (e.g., 400) and the reconstructed base layer signal (e.g., 380) (potentially upsampled / filtered).

[0147] Figure 25 illustrates the generation of the interlayer intra-prediction signal by the sum of the (upsampled / filtered) base layer reconstructed signal 380 (BL Reco) and the spatial intra-prediction using the difference signal 734 (EH Diff) of the already encoded adjacent block 736. There, the difference signal (EH Diff) for the already encoded block 736 is generated by subtracting the (upsampled / filtered) base layer reconstructed signal 380 (BL Reco) from the reconstructed enhancement layer signal (EH Reco) (example 400), where the already encoded / decoded portion is shown in diagonal lines. The current encoded / decoded block / region / part is 28. In other words, the interlayer intra-prediction method described in Figure 25 uses two superimposed input signals to generate the prediction block. For this method, the difference signal 734 is required. The difference signal 734 is the difference between the reconstructed enhancement layer signal 400 and the co-located reconstructed base layer signal 200. The base layer signal 200 is upsampled 220 to match the corresponding sample position in the enhancement layer and is optionally filtered before or after upsampling (if upsampling is not applied, it is filtered, in the case of quality scalable coding). In particular for spatial scalable coding, the difference signal 734 typically contains mainly high-frequency components. The difference signal 734 is available for all already reconstructed blocks (i.e., all enhancement layer blocks that have already been coded / decoded). The difference signal 734 for adjacent samples 742 of an already coded / decoded block 736 is used as input to a spatial intra-prediction technique (such as the spatial intra-prediction mode specified in H.264 / AVC or HEVC). The spatial intra-prediction indicated by arrow 744 generates a prediction signal 746 for the different components of the block 28 to be predicted. In certain embodiments, any trimming functionality of spatial intra-prediction processing (as is well known from H.264 / AVC or HEVC) is modified or disabled to match the dynamic range of the difference signal 734.The intra-prediction method actually used (one of several provided methods, which may include planar intra-prediction, DC intra-prediction, or directional intra-prediction 744 having any particular angle) is shown in bitstream 6b. It is possible to use a different spatial intra-prediction technique (a method for generating a prediction signal using samples from already encoded adjacent blocks) than the method provided for H.264 / AVC and HEVC. The resulting prediction block 746 (using difference samples from adjacent blocks) is the first part of the final prediction block 420.

[0148] A second portion of the predicted signal is generated using a co-located region 28 in the reconstructed signal 200 of the base layer. With respect to the quality enhancement layer, the co-located base layer sample is used directly or optionally filtered by, for example, a low-pass filter or a filter 500 that attenuates high-frequency components. With respect to the spatial enhancement layer, the co-located base layer sample is upsampled. For upsampling 220, an FIR filter or a set of FIR filters is used. IIR filters can also be used. Optionally, the reconstructed base layer sample 200 is filtered before upsampling, or the base layer predicted signal (the signal obtained after upsampling the base layer) is filtered after the upsampling step. The base layer reconstruction process may include one or more additional filters, such as a non-blocking filter (e.g., 120) or an adaptive loop filter (e.g., 140). The base layer reconstruction 200 used for upsampling is the reconstructed signal before any loop filter (e.g., 200c). Alternatively, it is the reconstructed signal after a non-blocking filter, but before any other filter (as exemplified by 200b). Alternatively, it is the reconstructed signal after a specific filter, or the reconstructed signal after applying all filters used in the base layer decoding process (as exemplified by 200a).

[0149] The two resulting parts of the predicted signal (the spatially predicted difference signal 746 and the potentially filtered / upsampled base layer reconstruction 380) are added together sample by sample 732 to form the final predicted signal 420.

[0150] To move the embodiment outlined just now to the embodiments of Figures 6-10 is that the possibility of predicting the current block of the enhancement layer signal, as outlined just now, is supported by each scalable video decoder / encoder as an alternative to the prediction scheme outlined with respect to Figures 6-10. The mode used is indicated in the enhancement layer substream 6b via its respective prediction mode identifier, which is not shown in Figure 8.

[0151] In certain embodiments, intra-prediction follows inter-layer residual prediction (see Example B). Multiplexing methods for generating an intra-prediction signal using base layer data include the following: A conventional spatial intra-prediction signal (obtained using samples from adjacent reconstructed enhancement layers) is coupled to a base layer residual signal (inverse transform of base layer transformation coefficients, or the difference between base layer reconstruction and base layer prediction) (upsampled / filtered).

[0152] Figure 26 shows the generation of the interlayer intra-prediction signal 420, which is the sum of 752, consisting of the (upsampled / filtered) base layer residual signal 754 (BL Resi) and the spatial intra-prediction 756, which uses the reconfigured enhancement layer samples 758 (EH Reco) of already encoded adjacent blocks indicated by the dotted line 762.

[0153] The concept shown in Figure 26 superimposes two prediction signals to form a prediction block 420, where one prediction signal 764 originates from an already reconstructed enhancement layer sample 758, and the other prediction signal 754 originates from a base layer residual sample 480. The first portion 764 of the prediction signal 420 is obtained by applying spatial intra-prediction 756 using the reconstructed enhancement layer sample 758. Spatial intra-prediction 756 is one of the methods specified in H.264 / AVC, or one of the methods specified in HEVC, or it is another spatial intra-prediction technique that generates a prediction signal 764 for the current block 18 that forms the sample 758 of the adjacent block 762. The intra-prediction method 756 actually used (one of several provided methods, which may include planar intra-prediction, DC intra-prediction, or directional intra-prediction with any particular angle) is shown in bitstream 6b. It is possible to use a different spatial intra-prediction technique (a method for generating a prediction signal using samples from already encoded adjacent blocks) than the method provided for H.264 / AVC and HEVC. A second portion 754 of the prediction signal 420 is generated using the co-arranged residual signal 480 of the base layer. With respect to the quality enhancement layer, the residual signal can be used to be reconstructed in the base layer, or the residual signal can be further filtered. With respect to the spatial enhancement layer 480, the residual signal is upsampled 220 (to map the base layer sample positions to the enhancement layer sample positions) before it is used as the second portion of the prediction signal. The base layer residual signal 480 can also be filtered before or after the upsampling stage. An FIR filter is applied to upsample the residual signal 220. The upsampling process is configured in such a way that it is not filtered across the transformed block boundaries in the base layer applied for the purpose of upsampling.

[0154] The base layer residual signal 480 used for interlayer prediction may be a residual signal from which the base layer transformation coefficient levels are obtained by scaling and inverse transformation 560. Alternatively, the base layer residual signal 480 is the difference between the reconstructed base layer signal 200 (before or after deblocking and additional filtering, or during any filtering operation) and the prediction signal 660 used in the base layer.

[0155] The two generated signal components (the spatial intra-prediction signal 764 and the inter-layer residual prediction signal 754) are added together to form the final enhancement layer intra-prediction signal 752.

[0156] This means that any scalable video decoder / encoder according to Figures 6-10 will use or support the prediction mode outlined in Figure 26 to form the alternative prediction mode described above in Figures 6-10 with respect to the currently encoded / decoded portion 28.

[0157] In certain embodiments, weighted predictions of spatial intra-prediction and base layer reconstruction (see Example C) are used. This actually represents a description of such weighted predictions, which is interpreted not only as an alternative to the above embodiments but also as a description of possible ways of carrying out the embodiments outlined above with respect to Figures 6 to 10 in a different manner from the given embodiments.

[0158] Multiplexing methods for generating an intra-prediction signal using base layer data include the following: The reconstructed (upsampled / filtered) base layer signal is coupled to the spatial intra-prediction signal, where the spatial intra-prediction is obtained based on samples from the reconstructed enhancement layer of adjacent blocks. The final prediction signal is obtained by weighting the spatial prediction signal and the base layer prediction signal (example 41) in a way that different frequency components use different weightings. This is achieved, for example, by filtering the base layer prediction signal (example 38) with a low-pass filter (example 62), filtering the spatial intra-prediction signal (example 34) with a high-pass filter (example 64), and then adding the resulting filtered signal (example 66). Alternatively, frequency-based weighting is achieved by transforming (examples 72, 74) the base layer prediction signal (example 38) and the enhancement layer prediction signal (example 34), and superimposing the resulting transformed blocks (examples 76, 78). There, different weighting coefficients (examples 82, 84) are used for different frequency positions. The resulting transformed block (example 42 in Figure 10) is then inversely transformed (example 84) and used as the enhancement layer prediction signal (example 54). Alternatively, the resulting transformed coefficients are added (example 52) to a scaled transmitted transformed coefficient level (example 59), and then inversely transformed (example 84) to obtain a reconstructed block (example 54) before deblocking and in-loop processing.

[0159] Figure 27 shows the generation of the interlayer intra-prediction signal by a frequency-weighted sum of the (upsampled / filtered) base layer reconstructed signal (BL Reco) and the spatial intra-prediction using the reconstructed enhancement layer samples (EH Reco) of already encoded adjacent blocks.

[0160] The concept in Figure 27 uses two superimposed signals 772, 774 to form a prediction block 420. The first portion 774 of signal 420 is obtained by applying a spatial intra-prediction 776, corresponding to 30 in Figure 6, using a reconfigured sample 778 of an already configured adjacent block in the enhancement layer. The second portion 772 of the prediction signal 420 is generated using a co-configured reconstructed signal 200 from the base layer. With respect to the quality enhancement layer, the co-configured base layer samples 200 are used directly. Alternatively, they are optionally filtered by, for example, a low-pass filter or a filter that attenuates high-frequency components. With respect to the spatial enhancement layer, the co-configured base layer samples are upsampled 220. For upsampling, an FIR filter or a set of FIR filters is used. It is also possible to use an IIR filter. Optionally, the reconstructed base layer samples are filtered before upsampling. Alternatively, the base layer prediction signal (the signal obtained after upsampling the base layer) is filtered after the upsampling step. The base layer reconstruction process may include one or more additional filters, such as a non-blocking filter 120 or an adaptive loop filter 140. The base layer reconstruction 200 used for upsampling is the reconstructed signal 200c before any of the loop filters 120, 140. Alternatively, it is the reconstructed signal 200b after the non-blocking filter 120, but before another filter. Alternatively, it is the reconstructed signal 200a after a particular filter, or the reconstructed signal after applying all the filters 120, 140 used in the base layer decoding process.

[0161] When the reference numerals used in Figures 23 and 24 are compared with the reference numerals used in relation to Figures 6 to 10, block 220 corresponds to reference numeral 38 used in Figure 6. 39 corresponds to a portion of 380. At least with respect to the portion co-located in the current portion 28, 420 co-located in the current portion 28 corresponds to 42. Spatial prediction 776 corresponds to 32.

[0162] Two prediction signals (a potentially upsampled / filtered base layer reconstruction 386 and an enhancement layer intra-prediction 782) are combined to form a final prediction signal 420. The method for combining these signals may have the characteristic that different weighting coefficients are used for different frequency components. In a particular embodiment, the upsampled base layer reconstruction is filtered with a low-pass filter (e.g., 62) (it is also possible to filter the base layer reconstruction before upsampling 220). The intra-prediction signal (e.g., 34 obtained by 30) is filtered with a high-pass filter (e.g., 64). The signals filtered through both are added together 784 (e.g., 66) to form the final prediction signal 420. The pair of low-pass and high-pass filters represents a pair of orthogonal mirror filters, but this is not necessarily required.

[0163] In another specific embodiment (as illustrated in Figure 10), the combining of two prediction signals 380 and 782 is achieved via a spatial transformation. Both the (potentially upsampled / filtered) base layer reconstructed 380 and the intra-prediction signal 782 are transformed (as illustrated in 72, 74) using the spatial transformation. The transformation coefficients (as illustrated in 76, 78) of both signals are then scaled with appropriate weighting coefficients (as illustrated in 82, 84) and then added (as illustrated in 90) to form a transformation coefficient block (as illustrated in 42) of the final prediction signal. In one version, the weighting coefficients (as illustrated in 82, 84) are selected such that, with respect to each transformation coefficient position, the sum of the weighting coefficients for the components of both signals is equal to 1. In another version, with respect to some or all of the transformation coefficient positions, the sum of the weighting coefficients is not equal to 1. In a particular version, the weighting coefficients are selected such that, with respect to the transformation coefficients representing low-frequency components, the weighting coefficient for base layer reconstruction is greater than the weighting coefficient for the enhancement layer intra-prediction signal, and with respect to the transformation coefficients representing high-frequency components, the weighting coefficient for base layer reconstruction is less than the weighting coefficient for the enhancement layer intra-prediction signal.

[0164] In one embodiment, the transformed coefficient block (represented as 42) obtained (by combining the weighted transformed signals for both components) is inversely transformed (represented as 84) to form the final predicted signal 420 (represented as 54). In another embodiment, the prediction is made directly within the transformed region. That is, the encoded transformed coefficient level (represented as 59) is scaled (i.e., inversely quantized) to the transformed coefficients (represented as 42) of the predicted signal (obtained by combining the weighted transformed signals for both components), and then added (represented as 52) to the resulting block (which is then inversely transformed (represented as 84) to obtain the reconstructed signal 420 for the current block (although not shown in Figure 10, (before the potential deblocking 120 and further in-loop filtering stage 140)). In other words, in the first embodiment, the transformed block obtained by combining the weighted transformed signals for both components is inversely transformed and used as the enhancement layer predicted signal. Alternatively, in the second embodiment, the obtained conversion coefficients are added to the scaled transmitted conversion coefficient levels and then inversely converted to obtain a reconstructed block before deblocking and in-loop processing.

[0165] The selection of base layer reconstruction and residual signal (see Example D) is also used. For the method of using the reconstructed base layer signal (as described above), the following versions are available.

[0166] • Reconstructed base layer samples 200c prior to unblocking 120 and further in-loop processing 140 (such as adaptive offset filters or adaptive loop filters as samples). • After deblocking 120, however, the reconstructed base layer sample 200b is subjected to another in-loop process 140 (such as an adaptive offset filter or adaptive loop filter as a sample). • Reconstructed base layer sample 200a after deblocking 120 and further in-loop processing 140 (such as an adaptive offset filter or adaptive loop filter as a sample), or between multiple in-loop processing steps.

[0167] The selection of the corresponding base layer signals 200a, b, c is fixed for the implementation of a particular decoder (and encoder). Alternatively, it is indicated within bitstream 6. In the latter case, different versions are used. The use of a particular version of the base layer signal is indicated at the sequence level, or at the image level, or at the slice level, or at the maximum coding unit level, or at the coding unit level, or at the prediction block level, or at the transformation block level, or at any other block level. In other versions, the selection may depend on other coding parameters (such as coding mode) or on the characteristics of the base layer signal.

[0168] In another embodiment, multiple versions of the method of using the (upsampled / filtered) base layer signal 200 are used. For example, two different modes are provided that directly use the upsampled base layer signal (i.e., 200a), where the two modes use different interpolation filters. Or, one mode uses additional filtering 500 of the (upsampled) base layer reconstructed signal. Similarly, multiple different versions are provided for other modes as described above. The upsampled / filtered base layer signal 380 employed for the different versions of the mode differs among the interpolation filters used (including interpolation filters that also filter integer sample positions). Or, the upsampled / filtered base layer signal 380 for the second version is obtained by filtering the upsampled / filtered base layer 500 for the first version. One selection of the different versions is shown at the sequence, image, slice, maximum coded unit, coded unit level, prediction block level, or transform block level. Alternatively, it can be inferred from the characteristics of the corresponding reconstructed base layer signal or transmitted coding parameters.

[0169] The same applies to modes that use the reconstructed base layer residual signal via 480. Here, different versions of the interpolation filter or additional filtering steps used are also employed.

[0170] Different filters are used to upsample / filter the reconstructed base layer signal and the base layer residual signal. This means that a different approach is used to upsample the base layer residual signal than to upsample the reconstructed base layer signal.

[0171] For base layer blocks where the residual signal is zero (i.e., no conversion coefficient levels are sent to the block at all), the corresponding base layer residual signal is replaced with another signal obtained from the base layer. For example, this could be a high-pass filtered version of the reconstructed base layer block, or another different signal obtained from a reconstructed sample of an adjacent block or a sample of the reconstructed base layer residual.

[0172] With respect to the samples used for spatial intra-prediction within the enhancement layer (see Example H), the following special processing is provided: In modes using spatial intra-prediction, unavailable adjacent samples within the enhancement layer (adjacent blocks are encoded after the current block, so adjacent samples are unavailable) are replaced with corresponding samples of the upsampled / filtered base layer signal.

[0173] As far as the coding of intra-prediction modes is concerned (see Example X), the following special modes and functionalities are provided. With respect to modes using spatial intra-prediction, such as 30a, the coding of the intra-prediction mode is such that (if available,) information about the intra-prediction mode in the base layer is modified in a way that is used to more efficiently encode the intra-prediction mode in the enhancement layer. This is used, for example, for parameter 56. If a co-located region in the base layer (e.g., 36) is intra-coded using a particular spatial intra-prediction mode, then a similar intra-prediction mode is likely to be used in the enhancement layer block (e.g., 28). Intra-prediction modes are usually indicated in a way that one or more modes are classified as the most likely mode from a set of possible intra-prediction modes, where they are indicated by shorter codewords. (Alternatively, the shorter the arithmetic code, the fewer bits the binary choice yields.) In HEVC intra-prediction, the intra-prediction mode of the block above (if available) and the intra-prediction mode of the block to the left (if available) are included in the set of most likely modes. In addition to these modes, one or more additional modes (often used) are included in the list of most likely modes, where the actual additional modes depend on the usefulness of the intra-prediction modes of the block above the current block and the block to the left of the current block. In HEVC, three modes are precisely classified as most likely modes. In H.264 / AVC, one mode is classified as the most likely mode. This mode is obtained based on the intra-prediction modes used for the block above the current block and the block to the left of the current block. Any other concepts for classifying intra-prediction modes (different from H.264 / AVC and HEVC) are possible and used for the following extensions.

[0174] The concept of using one or more most likely modes in base layer data for efficient coding of intra-predictive modes within an enhancement layer is modified in such a way that the most likely modes include the intra-predictive modes used in the co-located base layer block (assuming the corresponding base layer block is intra-coded). In a particular embodiment, the following approach is used: Given a current enhancement layer block, the co-located base layer block is determined. In a particular version, the co-located base layer block is the base layer block covering the co-located position of the top-left sample of the enhancement block. In another version, the co-located base layer block is the base layer block covering the co-located position of the central sample of the enhancement block. In yet another version, another sample within the enhancement layer block is used to determine the co-located base layer block. If the determined co-configured base layer block is intra-coded, and the base layer intra-prediction mode specifies an angular intra-prediction mode, and the intra-prediction mode obtained from the enhancement layer block to the left of the current enhancement layer block does not use an angular intra-prediction mode, then the intra-prediction mode obtained from the left enhancement layer block is replaced with the corresponding base layer intra-prediction mode. Otherwise, if the determined co-configured base layer block is intra-coded, and the base layer intra-prediction mode specifies an angular intra-prediction mode, and the intra-prediction mode obtained from the enhancement layer block above the current enhancement layer block does not use an angular intra-prediction mode, then the intra-prediction mode obtained from the above enhancement layer block is replaced with the corresponding base layer intra-prediction mode. Other versions use a different approach to modify the list of most likely modes (consisting of a single element) using the base layer intra-prediction mode.

[0175] Intercoding techniques for spatial and quality enhancement layers are then provided.

[0176] In state-of-the-art hybrid video coding standards (such as H.264 / AVC or the upcoming HEVC), an image sequence is divided into blocks of samples. The block size is fixed, or the coding technique provides a hierarchical structure that allows the block to be further subdivided into blocks with smaller block sizes. Block reconstruction is typically achieved by generating a prediction signal for the block and adding the transmitted residual signal. The residual signal is typically transmitted using transform coding, meaning that quantization indices for the transform coefficients (also called transform coefficient levels) are transmitted using entropy coding techniques. On the decoder side, these transmitted transform coefficient levels are then scaled and inversely transformed to obtain the residual signal that is added to the prediction signal. The residual signal is generated either by intra-prediction (using only data already transmitted for the current moment) or by inter-prediction (using data already transmitted for different moments).

[0177] In interpretation, the prediction block is obtained by motion-compensated prediction using samples of already reconstructed frames. This is done by unidirectional prediction (using one reference image and one set of motion parameters). Alternatively, the prediction signal can be generated by multiple-hypothesis prediction. In the latter case, two or more prediction signals are superimposed; that is, for each sample, a weighted average is constructed to form the final prediction signal. The (superimposed) multiple prediction signals are generated using different motion parameters for different hypotheses (e.g., different reference images or motion vectors). It is also possible to add a certain offset to form the final prediction signal by multiplying the samples of the motion-compensated prediction signal with a constant coefficient for unidirectional prediction. Such scaling and offset corrections are also used for all or selected hypotheses in the multiple-hypothesis prediction.

[0178] In scalable video coding, base layer information is used to support inter-prediction processing for the enhancement layer. The SVC extension of H.264 / AVC, the most advanced video coding standard for scalable coding, is an additional mode for improving the coding efficiency of inter-prediction processing within the enhancement layer. This mode is represented at the macroblock level (blocks of 16x16 luma samples). In this mode, reconstructed residual samples from the lower layer are used to improve the motion-compensated prediction signal in the enhancement layer. This mode is also called inter-layer residual prediction. If this mode is selected for a macroblock in the quality enhancement layer, the inter-layer prediction signal is constructed from co-arranged samples of the reconstructed lower-layer residual signal. If the inter-layer residual prediction mode is selected in the spatial enhancement layer, the prediction signal is generated by upsampling the co-arranged, reconstructed base-layer residual signal. FIR filters are used for upsampling; however, filtering is not applied across transform block boundaries. The prediction signal, derived from the reconstructed base layer residual sample, is added to the conventional motion-compensated prediction signal to form the final prediction signal for the enhancement layer block. Generally, for the interlayer residual prediction mode, the additional residual signal is transmitted via transform coding. Transmission of the residual signal is omitted (inferred to be equal to zero) if it is represented correspondingly in the bitstream. The final reconstructed signal is obtained by adding the reconstructed residual signal (obtained by scaling the transmitted transform coefficient levels and applying the inverse space transform) to the prediction signal (where the interlayer residual prediction signal is added to the motion-compensated prediction signal).

[0179] Next, techniques for intercoding enhancement layer signals are described. This section describes a method for employing base layer signals, in addition to already reconstructed enhancement layer signals, to intra-predict enhancement layer signals to be encoded in scalable video coding scenarios. By employing base layer signals for inter-predicting enhancement layer signals to be encoded, prediction errors are sufficiently reduced. This results in an overall bit transmission rate saving with respect to the encoding of the enhancement layer. The main focus of this section is to increase the block-based motion compensation of enhancement layer samples by using already encoded enhancement layer samples with additional signals from the base layer. The following description provides possibilities for using various signals from the encoded base layer. Although a quadtree block partition is generally employed as a preferred embodiment, the examples presented are applicable to general block-based hybrid coding approaches without assuming any particular block partition. The use of base layer reconstruction of the current time index, base layer residuals of the current time index, or even base layer reconstruction of already encoded images for inter-predicting enhancement layer blocks to be encoded is described. Furthermore, a method is described for coupling the base layer signal with the already encoded enhancement layer signal in order to obtain better predictions for the current enhancement layer. One of the most cutting-edge technologies is interlayer residual prediction in H.264 / SVC. Interlayer residual prediction in H.264 / SVC is employed for all inter-encoded macroblocks, whether encoded or not, using the SVC macroblock type, which is indicated by the use of either a base mode flag or a conventional macroblock type. The flag is added to the macroblock syntax for the spatial and quality enhancement layer. When this residual prediction flag, which indicates the use of interlayer residual prediction, is equal to 1, the residual signal of the corresponding region in the reference layer is upsampled block by block using a bilinear filter and used as a prediction for the residual signal of the enhancement layer macroblock. As a result, only the corresponding difference signal needs to be encoded in the enhancement layer. The following notation will be used in the descriptions in this section. t0 := Time index of the current image t1 := Time index of the already reconstructed image EL:=Enhancement Layer BL:=Base layer EL(t0):=Current enhancement layer image to be encoded EL_reco := Enhancement layer reconstruction BL_reco := Base layer reconstruction BL_resi:=Base layer residual signal (inverse transformation of base layer transformation coefficients, or the difference between base layer reconstruction and base layer prediction) EL_diff := Difference between enhancement layer reconstruction and upsampled / filtered base layer reconstruction The different base layer signals and enhancement layer signals are used in the descriptions shown in Figure 28.

[0180] Regarding the description, the following properties of the filter are used. Linearity: Although many of the filters mentioned in the description are linear, nonlinear filters are also used. • Number of output samples: In an upsampling operation, the number of output samples is greater than the number of input samples. Here, filtering the input data produces more samples than the input values. In conventional filtering, the number of output samples is equal to the number of input samples. Such filtering operations are used, for example, in quality-scalable coding. • Phase delay: With respect to filtering samples at integer positions, the phase delay is typically zero (or an integer delay within the sample). With respect to the occurrence of samples at decimal positions (e.g., half-per or quarter-per positions), filters with a decimal delay (within the sample unit) are typically applied to samples on an integer grid.

[0181] Conventional motion-compensated prediction, used in all hybrid video encoding standards (e.g., MPEG-2, H.264 / AVC, or the upcoming HEVC standard), is illustrated in Figure 29. To predict the signal for the current block, an already reconstructed region of the image is replaced and used as the prediction signal. For signaling of the replacement, the motion vector is typically encoded in the bitstream. With respect to integer-sample-precision motion vectors, a reference region in the reference image is directly copied to form the prediction signal. However, it is also possible to transmit fractional-sample-precision motion vectors. In this case, the prediction signal is obtained by filtering the reference signal with a filter having a fractional-sample delay. The reference image used is typically specified by including the reference image index in the bitstream syntax. Generally, it is also possible to superimpose two or more prediction signals to form the final prediction signal. The concept is supported, for example, in a B-slice with two motion hypotheses. In this case, the multiple prediction signals are generated using different motion parameters (e.g., different reference images or motion vectors) for different hypotheses. For unidirectional predictions, it is also possible to multiply a sample of the motion-compensated prediction signal with a constant coefficient and add a certain offset to form the final prediction signal. Such scaling and offset corrections can also be used for all or selected hypotheses in multi-hypothesis predictions.

[0182] The following descriptions apply to scalable coding with a quality enhancement layer (where the enhancement layer represents an input video with the same resolution as the base layer but with higher quality or fidelity) and scalable coding with a spatial enhancement layer (where the enhancement layer has a higher resolution than the base layer, i.e., a larger number of samples). For the quality enhancement layer, upsampling of the base layer signal is not required, but filtering of the reconstructed base layer samples is applied. For the spatial enhancement layer, upsampling of the base layer signal is generally required.

[0183] The embodiments support different methods for using reconstructed base layer samples or base layer residual samples for inter-prediction of enhancement layer blocks. In addition to conventional inter-prediction and intra-prediction, it is possible to support one or more of the methods described below. The usage of a particular method is shown at the level of the largest supported block size (such as macroblocks in H.264 / AVC, or coded tree blocks / maximum coded units in HEVC). It is shown for all supported block sizes. Or it is shown for a subset of supported block sizes.

[0184] For all the methods described below, the prediction signal is used directly as the reconstruction signal for the block. Alternatively, a selected method for interlayer interprediction is combined with residual coding. In a particular embodiment, the residual signal is transmitted via transform coding. That is, the quantization transform coefficients (transformation coefficient levels) are transmitted using entropy coding techniques (e.g., variable-length coding or arithmetic coding), and the residual is obtained by inversely quantizing (scaling) the transmitted transform coefficient levels and applying the inverse transform. In a particular version, the complete residual block corresponding to the block from which the interlayer interprediction signal arises is transformed using a single transform. (i.e., the entire block is transformed using a single transform of the same size as the prediction block.) In another embodiment, the prediction block is further subdivided into smaller blocks, for example, using hierarchical decomposition. Then, a separate transform is applied to each of the smaller blocks (with different block sizes). In yet another embodiment, the coded unit is divided into smaller prediction blocks. Then, for zero or more of the prediction blocks, the prediction signal arises using one of the methods for interlayer interprediction. Next, the residuals of the entire coding unit are transformed using a single transform. Alternatively, the coding unit is subdivided into different transform units, where the subdivision to form transform units (blocks to which a single transform is applied) is different from the subdivision to decompose the coding unit into prediction blocks.

[0185] The following describes the possibility of performing prediction using base layer residuals and enhancement layer reconstruction. The multiplexing method includes the following: The conventional interpretation signal (obtained by motion-compensated interpolation of an already reconstructed enhancement layer image) is coupled to the base layer residual signal (inverse transform of base layer transformation coefficients, or the difference between base layer reconstruction and base layer prediction) (upsampled / filtered). This method is also called the "BL_resi" mode (as illustrated in Figure 30).

[0186] In short, the predictions for the enhancement layer sample are described below. EL prediction=filter(BL_resi(t0))+MCP_filter(EL_reco(t1)) It is also possible to use two or more hypotheses for the enhancement layer reconstruction signal. For example, EL prediction = filter(BL_resi(t0)) + MCP_filter1(EL_reco(t1)) + MCP_filter2(EL_reco(t2)) The motion-compensated prediction (MCP) filters used on the enhancement layer (EL) reference image are integer or fractional sample precision. The MCP filters used on the EL reference image are either the same as or different from the MCP filters used on the BL reference image during the BL decoding process. A motion vector MV(x, y, t) is defined to indicate a specific location within an EL reference image. The parameters x and y indicate the spatial location within the image. The parameter t is used to describe the temporal index of the reference image and is also called the reference index. Often, the term motion vector is used to refer to only the two spatial components (x, y). The integer part of MV is used to take a set of samples from the reference image. The fractional part of MV is used to select an MCP filter from a set of filters. The taken reference samples are filtered to produce the filtered reference samples. Motion vectors are generally encoded using different predictors. This means that the motion vector predictor is obtained based on an already encoded motion vector (and the syntactic elements potentially indicate one of a set of potential motion vector predictors used), and different vectors are included in the bitstream. The final motion vector is obtained by adding the transmitted motion vector difference to the motion vector predictor. It is usually also possible to obtain the motion parameters for a block completely. Therefore, the list of potential motion parameter candidates is typically constructed based on already encoded data. This list can include motion parameters of spatially adjacent blocks, as well as motion parameters obtained based on motion parameters in co-located blocks within a reference frame. The base layer (BL) residual signal can be defined as one of the following: • Inverse transformation of BL transformation coefficients, or The difference between BL reconstruction and BL prediction, or Regarding BL blocks where the inverse transform of the BL transform coefficient is zero, it can be replaced with another signal obtained from the BL, for example, a high-pass filtered version of the reconstructed BL block, or • A combination of the above methods. To calculate EL prediction elements from the current BL residuals, the regions to be considered in the EL image and the co-located regions in the BL image are identified, and the residual signal is extracted from the identified BL regions. The definition of the co-located regions is constructed so that it describes an integer scaling factor of the BL resolution (e.g., 2 × scalability) or a decimal scaling factor of the BL resolution (e.g., 1.5 × scalability). Alternatively, it can even produce the same EL resolution as the BL resolution (e.g., quality scalability). In the case of quality scalability, the co-located blocks in the BL image have the same coordinates as the EL blocks to be predicted. Co-located BL residuals can be upsampled / filtered to generate filtered BL residual samples. The final EL prediction is obtained by adding the filtered EL reconstruction sample and the filtered BL residual sample.

[0187] A multiplexing method for prediction using base layer reconstruction and enhancement layer difference signals (see Example J) includes the following: The (upsampled / filtered) reconstructed base layer signal is coupled to a motion-compensated prediction signal, which is obtained by a motion-compensated difference image. The difference image represents the difference between the reconstructed enhancement layer signal and the (upsampled / filtered) reconstructed base layer signal with respect to a reference image. This method is also called BL_reco mode.

[0188] This concept is explained in Figure 31. In short, the prediction for the EL sample is described below: EL prediction = filter(BL_reco(t0)) + MCP_filter(EL_diff(t1))

[0189] It is also possible to use two or more hypotheses for the EL difference signal. For example, EL prediction = filter(BL_resi(t0)) + MCP_filter1(EL_diff(t1)) + MCP_filter2(EL_diff(t2))

[0190] Regarding the EL difference signal, the following versions are used. The difference between EL reconstruction and upsampled / filtered BL reconstruction, or The difference between the EL reconstruction before or during a loop filtering stage (such as non-blocking, SAO, ALF) and the upsampled / filtered BL reconstruction.

[0191] The usage of a particular version is fixed within the decoder, or it is indicated at the sequence level, image level, slice level, maximum encoding unit level, encoding unit level, or another partition level. Alternatively, it may depend on different encoding parameters.

[0192] When the EL difference signal is defined using the difference between the EL reconstruction and the upsampled / filtered BL reconstruction, we simply store the EL and BL reconstructions and, using prediction mode, become compliant in quickly calculating the EL difference signal for the block. As a result, we store the memory necessary to store the EL difference signal. However, this incurs a slight computational complexity overhead.

[0193] The MCP filter used on the EL difference image can have integer or fragmentary sample precision. Regarding the MCP of the difference image, a different interpolation filter can be used than that of the MCP of the reconstructed image. Regarding the MCP of the difference image, the interpolation filter is selected based on the characteristics of the corresponding region in the difference image (or based on the encoding parameters in the bitstream, or based on the transmitted information).

[0194] The motion vector MV(x,y,t) is defined to pinpoint a specific location within an EL difference image. The parameters x and y pinpoint the spatial location within the image, and the parameter t is used to describe the temporal index of the difference image.

[0195] The integer part of MV is used to take one set of samples from the difference image, and the fractional part of MV is used to select an MCP filter from one set of filters. The taken difference samples are then filtered to generate filtered difference samples.

[0196] The dynamic range of the difference image can theoretically exceed the dynamic range of the original image. Assuming an 8-bit representation of the image within the range

[0255] , the difference image can have a range [-255 255]. However, in practice, the majority of the amplitude is distributed around ±0. In a preferred embodiment for storing the difference image, a constant offset of 128 is added, and the result is cropped to the range

[0255] and stored as a regular 8-bit image. Subsequently, during the encoding and decoding process, the offset of 128 is subtracted from the difference amplitude read from the difference image and returned.

[0197] Regarding the method of using the reconstructed BL signal, the following versions can be used. This is fixed, or it is shown at the sequence level, image level, slice level, maximum coding unit level, coding unit level, or another partition level. Alternatively, it can depend on a different coding parameter. • Reconstructed base layer samples before deblocking and further in-loop processing (such as adaptive offset filters or adaptive loop filters). • After deblocking, but before further in-loop processing, the reconstructed base layer samples (such as adaptive offset filters or adaptive loop filters) are processed. - Reconfigured base layer samples after deblocking and further in-loop processing (such as adaptive offset filters or adaptive loop filters), or reconfigured base layer samples between multiple i in-loop processing steps.

[0198] To calculate the EL prediction component from the current BL reconstruction, regions in the BL image that are co-located with regions considered in the EL image are identified. The reconstructed signal is then extracted from the identified BL regions. The definition of co-located regions is constructed to describe whether they produce the same EL resolution as an integer scaling factor of the BL resolution (e.g., 2 × scalability), or a fractional scaling factor of the BL resolution (e.g., 1.5 × scalability), or even the BL resolution (e.g., SNR scalability). In the case of SNR scalability, the co-located blocks in the BL image have the same coordinates as the EL blocks to be predicted.

[0199] The final EL prediction is obtained by adding the filtered EL difference sample and the filtered BL reconstruction sample.

[0200] Several possible variations of the mode for coupling the (upsampled / filtered) base layer reconstructed signal with the motion-compensated enhancement difference layer signal are described below. Multiple versions of the method for using the (upsampled / filtered) BL signal are used. The upsampled / filtered BL signal employed for these versions may differ from the interpolation filter used (including interpolation filters that also filter integer sample positions), or the upsampled / filtered BL signal for the second version may be obtained by filtering the upsampled / filtered BL signal for the first version. A selection of one of the different versions is shown in sequence, image, slice, at the maximum encoding unit level, at the encoding unit level, or at another level of image segmentation. Alternatively, it may be inferred from the characteristics of the corresponding reconstructed BL signal or the transmitted encoding parameters. Different filters are used for the upsampled / filtered BL reconstructed signal in BL_reco mode and for the BL residual signal in BL_resi mode. The upsampled / filtered BL signal can also be combined with two or more hypothetical motion-compensated difference signals. This is illustrated in Figure 32.

[0201] Considering the above, prediction is performed using a combination of base layer reconstruction and enhancement layer reconstruction (see Example C). One major difference from the above description with respect to Figures 11, 12, and 13 is the encoding mode for obtaining the intra-layer prediction 34, which is performed temporally rather than spatially. That is, instead of the spatial prediction 30, the temporal prediction 32 is used to form the intra-layer prediction signal 34. Thus, several embodiments described below are readily transferable to the embodiments above Figures 6-10 and Figures 11-13, respectively. The multiplexing method includes the following: The (upsampled / filtered) reconstructed base layer signal is coupled to an inter-prediction signal, where the inter-prediction is obtained by motion-compensated prediction using the reconstructed enhancement layer image. The final prediction signal is obtained by weighting the inter-prediction signal and the base layer prediction signal in a way that different frequency components use different weightings. For example, this can be achieved by one of the following: This involves filtering the base layer prediction signal with a low-pass filter, filtering the inter-layer prediction signal with a high-pass filter, and then summing the resulting filtered signals. The base layer prediction signal and the interlayer prediction signal are transformed and the resulting transformed block is superimposed, where different weighting coefficients are used for different frequency positions. The resulting transformed block is then inversely transformed and used as the enhancement layer prediction signal. Alternatively, the resulting transformed coefficients are added to a scaled transmitted transformed coefficient level and inversely transformed to obtain a reconstructed block before deblocking and in-loop processing.

[0202] This mode may also be called the "BL_comb" mode, as shown in Figure 33.

[0203] In short, the EL forecast is as follows: EL prediction=BL_weighting(BL_reco(t0))+EL_weighting(MCP_filter(EL_reco(t1)))

[0204] In a preferred embodiment, the weighting is done depending on the ratio of EL resolution to BL resolution. For example, when BL should be scaled up by a coefficient within the range [1 1.25), a predetermined set of weightings for EL and BL reconstruction is used. When BL should be scaled up by a coefficient within the range [1.25 1.75], a different set of weightings is used. When BL should be scaled up by a coefficient greater than or equal to 1.75, another different set of weightings is used, and so on.

[0205] In other embodiments relating to spatial intra-layer prediction, it is also possible to perform specific weightings that depend on the scaling coefficient separation base layer and enhancement layer.

[0206] In another preferred embodiment, the weighting is made dependent on the EL block size for prediction. For example, a weighting matrix is ​​defined that specifies the weights for the EL reconstruction transformation coefficients for a 4x4 block in the EL. Then another weighting matrix is ​​defined that specifies the weights for the BL reconstruction transformation coefficients. The weighting matrix for the BL reconstruction transformation coefficients is, for example, as follows: 64, 63, 61, 49, 63, 62, 57, 40, 61,56,44,28, 49, 46, 32, 15, The weighting matrix for the EL reconstruction transformation coefficients is, for example, as follows: 0,2,8,24, 3,7,16,32, 9,18,20,26, 22, 31, 30, 23,

[0207] Similarly, separation weighting matrices are defined for block sizes such as 8x8, 16x16, and 32x32.

[0208] The actual transformations used for frequency-domain weighting are either the same as or different from the transformations used to encode the predicted residuals. For example, integer approximations for DCT can be used for both frequency-domain weighting and for calculating the transformation coefficients of the predicted residuals to be encoded in the frequency domain.

[0209] In another preferred embodiment, the maximum transformation size is defined for frequency-domain weighting to limit the computational complexity. If the EL block size under consideration is larger than the maximum transformation size, the EL and BL reconstructions are spatially separated into a series of adjacent subblocks. Frequency-domain weighting is performed on the subblocks. The final predicted signal is formed by assembling the weighted results.

[0210] Furthermore, weighting can be performed using luminance and chrominance components, or a selected subset of chrominance components.

[0211] The following describes different possibilities for obtaining enhancement layer coding parameters. The coding (or prediction) parameters to be used to reconstruct the enhancement layer block are obtained by a multiplexing method from co-located coding parameters in the base layer. The base layer and the enhancement layer may have different spatial resolutions, or they may have the same spatial resolution.

[0212] In the H.264 / AVC scalable video extension, interlayer motion prediction is performed for macroblock types indicated by the syntactic element base mode flag. If the base mode flag is equal to 1 and the corresponding reference macroblock in the base layer is intercoded, then the enhancement layer macroblock is also intercoded, and all motion parameters are inferred from the co-located base layer block. Otherwise (the base mode flag is equal to 0), for each motion vector, the so-called "motion prediction flag" is transmitted, specifying whether the base layer motion vector is used as the motion vector predictor. If the "motion prediction flag" is equal to 1, the motion vector predictor of the co-located reference block in the base layer is scaled according to the resolution ratio and used as the motion vector predictor. If the "motion prediction flag" is equal to 0, the motion vector predictor is calculated as specified in H.264 / AVC.

[0213] The following describes a method for obtaining enhancement layer coding parameters. The sample array related to the base layer image is decomposed into blocks, each block relating to coding (or prediction) parameters. In other words, all sample locations within a particular block have specific relational coding (or prediction) parameters. The coding parameters include parameters for motion-compensated prediction, including the motion hypothesis, reference index list, motion vector, motion vector predictor identifier, and number of integrated identifiers. The coding parameters also include intra-prediction parameters, such as intra-prediction direction.

[0214] The fact that blocks within the enhancement layer are encoded using co-located information from the base layer can be shown in the bitstream.

[0215] For example, the derivation of enhancement layer coding parameters (see Example T) is constructed as follows: For N×M blocks of the enhancement layer shown using co-located base layer information, the coding parameters related to the sample positions within the blocks are derived based on the coding parameters related to the co-located sample positions in the base layer sample array.

[0216] In a particular embodiment, this process is carried out by the following steps. 1. Deriving coding parameters for each sample position in an N×M enhancement layer block based on the base layer coding parameters. 2. Derivation of the partitions for N×M enhancement layer blocks within a subblock such that all sample positions within a specific subblock have the same related encoding parameters.

[0217] Also, the second step can be omitted.

[0218] Step 1 involves setting the enhancement layer sample position p el function f c This can be done by using and providing the encoding parameter c. That is, c=f c (p el )

[0219] For example, to ensure the minimum block size m × n in the enhancement layer, we use the function f c The function f is given by the following relationship: p,m × n p given by bl It can return the coding parameter c related to this. TIFF2026090372000002.tif77153

[0220] The distance between two adjacent horizontal or vertical base layer sample positions is consequently equal to 1. Both the top-left base layer sample and the top-left enhancement layer sample have the position p=(0,0).

[0221] As another example, function f c (p el ) can return the coding parameter c related to the base layer sample position p el that is closest to the base layer sample position p bl . Also, function f c (p el ) can interpolate the coding parameter when a particular enhancement layer sample position has a fractional component within a unit of the distance between base layer sample positions.

[0222] Before returning the motion parameter, function f c cycles through the spatial replacement components of the motion parameter for the closest available value within the enhancement layer sample upsampling lattice.

[0223] Since each sample position is related to the prediction parameter after step 1, samples of each enhancement layer are predicted after step 1. Nevertheless, in step 2, block partitioning is derived to perform a prediction operation on samples of a larger block or to transform and encode prediction residuals within the derived partitioned blocks.

[0224] Step 2 is performed by classifying enhancement layer sample positions into square or rectangular blocks. Each is decomposed into one of a set of allowed decompositions within the sub-blocks. The square or rectangular blocks correspond to the leaves in a quadtree structure where they can exist at different levels represented in FIG. 34.

[0225] The level and decomposition of each square or rectangular block are determined by performing the following ordered steps. a) Set the highest level to the level corresponding to a block of size N×M. Set the current level to the lowest level, i.e., the level where the square or rectangular block contains a single block of the smallest block size. Go to step b). b) For each square or rectangular block at the current level, if there are possible decompositions of the square or rectangular block, then all sample positions within each subblock are related to the same coding parameter, or related to the coding parameter by a small difference (depending on the magnitude of some difference). That decomposition is a candidate decomposition. From all candidate decompositions, select the one that decomposes the square or rectangular block into the fewest number of subblocks. If the current level is the highest level, proceed to step c). Otherwise, set the current level to the next higher level and proceed to step b). c) End

[0226] function f c This can be selected in such a way that at some level in step b), there is always at least one decomposition candidate.

[0227] While the decomposition of blocks with the same encoding parameters is not limited to square blocks, blocks can be grouped into rectangular blocks. Furthermore, the decomposition is not limited to a quadtree structure. It is also possible to use decomposition structures in which a block is broken down into two rectangular blocks of the same size, or into two rectangular blocks of different sizes. It is also possible to use decomposition structures that use quadtree decomposition up to a certain level, followed by decomposition into two rectangular blocks. Any other block decomposition is also possible.

[0228] In contrast to the SVC interlayer motion parameter prediction mode, the described mode is supported not only at the macroblock level (or the most favored block size) but also at any block size. This means that the mode is not only shown for the most favored block size, but that blocks of the most favored block size (macroblocks in MPEG4, H.264, and coded tree blocks / maximum coded units in HEVC) are hierarchically subdivided into smaller blocks / coded units, and the use of the interlayer motion mode is shown for any favored block size (with respect to the corresponding block). In certain embodiments, this mode supports only a selected block size. Then, the syntactic element indicating the use of this mode is sent only for the corresponding block size. Or, the value of the syntactic element indicating the use of this mode (in another coding parameter) is restricted to correspond to a different block size. Another difference from the interlayer motion parameter prediction mode in the SVC extension of H.264 / AVC is that blocks coded in this mode are not fully intercoded. A block may include intra-encoded subblocks, depending on the co-located base layer signals.

[0229] One of several methods for reconstructing samples of an M×M enhancement layer block using the coding parameters obtained by the method described above is shown in the bitstream. Such methods for predicting an enhancement layer block using the obtained coding parameters include: • For motion compensation, obtain a predictive signal for the enhancement layer block using the obtained motion parameters and the reconstructed enhancement layer reference image. - A combination of (a) an (upsampled / filtered) base layer reconstruction for the current image and (b) a motion compensation signal, using the resulting motion parameters and enhancement layer reference image, which are obtained by subtracting the (upsampled / filtered) base layer reconstruction from the reconstructed enhancement layer image. (a) The base layer residual (upsampled / filtered) for the current image (the difference between the reconstructed signal and the prediction, or the inverse transform of the encoded transformation coefficient value), and (b) a combination of the obtained motion parameters and a motion compensation signal using the reconstructed enhancement layer reference image.

[0230] Assuming another subblock is intercoded, while it is being classified, the process of obtaining partitions within smaller blocks with respect to the current block and obtaining coding parameters for the subblock can be intracoded and classified into several subblocks. For the intercoded subblocks, motion parameters are obtained from the co-located base layer blocks. However, if the co-located base layer blocks are intracoded, then the corresponding subblocks in the enhancement layer are classified as intracoded. For samples of such intracoded subblocks, the enhancement layer signal is predicted by using information from the base layer. For example, • The corresponding (upsampled / filtered) version of the base layer reconstruction is used as the intra-prediction signal. The obtained intra-prediction parameters are used for spatial intra-prediction within the enhancement layer.

[0231] To predict an enhancement layer block using a weighted combination of prediction signals, the following embodiments include a method for generating a prediction signal for an enhancement layer block by combining (a) an enhancement layer intra-prediction signal obtained by spatial or temporal (i.e., difference-compensated) prediction using a reconstructed enhancement layer sample, and (b) a base layer prediction signal which is a (upsampled / filtered) base layer reconstruction for the current image. The final prediction signal is obtained by weighting the enhancement layer intra-prediction signal and the base layer prediction signal in such a way that weighting according to a weighting function is used for each sample.

[0232] For example, the weighting function is realized in the following way: Compare the low-pass filtered version of the original enhancement layer intra-prediction signal v with the low-pass filtered version of the base layer reconstruction u. From this comparison, derive the weights for each sample position that should be used to combine the original intra-prediction signal and the (upsampled / filtered) base layer reconstruction. For example, the weights are obtained by mapping the difference uv with respect to the weights w using the transfer function t. That is, t(uv)=w

[0233] Different weighting functions are used for different block sizes of the current block to be predicted. Furthermore, the weighting functions are modified according to the temporal distance of the reference image from which the interpretation hypothesis is obtained.

[0234] For enhancement layer intraprediction signals, which are intraprediction signals, the weighting function is implemented using different weights that depend, for example, on the position in the current block to be predicted.

[0235] In a preferred embodiment, a method is used to obtain enhancement layer coding parameters. Step 2 of the method uses a set of possible decompositions of square blocks, as shown in Figure 35.

[0236] In a preferred embodiment, the function f c (p el ) is the function f described above with m=4 and n=4 p,m × n (p el It returns coding parameters related to the base layer sample positions given by ).

[0237] In this embodiment, the function f c (p el ) returns the following encoding parameter c. • First, the base layer sample position is p bl =f p,4 ×4(p el ) obtained as. • For example, p bl However, if it has related interprediction parameters (or the same motion parameters) obtained by integrating with a previously encoded base layer block, then c is equal to the motion parameters of the enhancement layer block corresponding to the base layer block used for integration within the base layer (i.e., the motion parameters are copied from the corresponding enhancement layer block). Otherwise, c is p bl It is equal to the coding parameters related to this.

[0238] Furthermore, combinations of the above embodiments are also possible.

[0239] In another embodiment, using co-located base layer information, for an enhancement layer block to be shown, since the intra prediction parameters obtained from an initial set of motion parameters are related to their enhancement layer sample positions, the block is integrated with a block containing these samples (i.e., a copy of the initial set of motion parameters). The initial set of motion parameters consists of an indicator for using one or two hypotheses, a reference index list referring to the first image in a list of reference images, and motion vectors with zero padding.

[0240] In another embodiment, using co-located base layer information, for an enhancement layer block having the obtained motion parameters, enhancement layer samples are first predicted and reconstructed in a certain order. Thereafter, samples having the obtained intra prediction parameters are predicted in the intra reconstruction order. As a result, intra prediction can use the already reconstructed sample values from (a) adjacent inter prediction blocks and (b) the previous adjacent intra prediction blocks in the intra reconstruction order.

[0241] In another embodiment, for an enhancement layer block that is integrated (i.e., takes motion parameters obtained from another inter prediction block), the list of integration candidates additionally includes candidates from the corresponding base layer block. And if the enhancement layer has a higher spatial upsampling ratio than the base layer, additionally, by improving only the available adjacent values in the enhancement layer for the spatial replacement component, it includes up to four candidates obtained from the base layer candidates.

[0242] In another embodiment, the magnitude of the difference used in step 2b) asserts that there is a small difference in the sub-block only if the difference disappears completely. That is, the sub-block can be formed only when all the included sample positions have the same obtained coding parameters.

[0243] In another embodiment, if (a) all included sample positions have obtained motion parameters and no set of sample positions in a block have different obtained motion parameters that are greater than a certain value according to the vector norm applied to the corresponding motion vector, or (b) all included sample positions have obtained intra-prediction parameters and no set of sample positions in a block have different obtained intra-prediction parameters that are greater than a certain angle of intra-prediction in direction, then the magnitude of the difference used in step 2b) is to assert that there is a small difference in the subblock. The parameters obtained for the subblock are calculated by meaningful or median operation. In another embodiment, the partition obtained by inferring encoding parameters from the base layer is further refined based on side information shown in the bitstream. In another embodiment, residual coding for a block whose coding parameters are inferred from the base layer is independent of the partitions within the block inferred from the base layer. For example, this means that a single transform is applied to the block, although the inference of coding parameters from the base layer divides the block into several subblocks, each having a separate set of coding parameters. Alternatively, the block from which the partitions for the subblocks and coding parameters are inferred from the base layer is divided into smaller blocks for the purpose of transform coding the residuals. There, the division into transform blocks is independent of the inferred partitions within the block having different coding parameters.

[0244] In another embodiment, residual coding for blocks whose coding parameters are inferred from the base layer depends on the dividers within the blocks inferred from the base layer. For example, this means that, with respect to transform coding, the division of blocks within a transform block depends on the dividers inferred from the base layer. In one version, a single transform is applied to each subblock having different coding parameters. In another version, the dividers are refined based on side information included in the bitstream. In yet another version, several subblocks are combined into a larger block, as shown in the bitstream, for the purpose of transform coding the residual signals.

[0245] Furthermore, embodiments obtained by combining the aforementioned embodiments are also possible.

[0246] In relation to enhancement layer motion vector coding, the following section describes a method for reducing motion information in scalable video coding applications by providing multiple enhancement layer predictors to efficiently encode motion information in the enhancement layer and using motion information encoded in the base layer. This idea is suitable for scalable video coding, including spatial, temporal, and quality scalability.

[0247] In the H.264 / AVC Interlayer Scalable Video Extension, motion prediction is performed for macroblock types indicated by the syntactic element "base mode flag". If the "base mode flag" is equal to 1 and the corresponding reference macroblock in the base layer is intercoded, then the enhancement layer macroblock is also intercoded. All motion parameters are then inferred from the co-located base layer block. Otherwise (if the "base mode flag" is equal to 0), each motion vector (the syntactic element of the so-called "motion prediction flag") is transmitted and specified regardless of whether the base layer motion vector is used as the motion vector predictor. If the "motion prediction flag" is equal to 1, the motion vector predictor of the co-located reference block in the base layer is scaled according to the resolution ratio and used as the motion vector predictor. If the "motion prediction flag" is equal to 0, the motion vector predictor is calculated as specified in H.264 / AVC. In HEVC, motion parameters are predicted by applying Advanced Motion Vector Competition (AMVP). AMVP features two spatial motion vector predictors and one temporal motion vector predictor competing with each other. Spatial candidates are selected from the positions of adjacent prediction blocks located to the left or above the current prediction block. Temporal candidates are selected from co-located positions in previously encoded images. All spatial and temporal candidate positions are shown in Figure 36.

[0248] After spatial and temporal candidates have been inferred, a redundancy check is performed to introduce zero motion vectors into the list of candidates. The candidate list describing the index is sent to identify the motion vector predictor to be used with the motion vector difference for motion compensation prediction. HEVC further employs a block merging algorithm that aims to reduce coded redundant motion parameters arising from quadtrees based on coding configurations. This is achieved by creating regions consisting of multiple prediction blocks that share specific motion parameters. These motion parameters only need to be coded once for the first prediction block in each region that is sowing the seeds of new motion information. Similar to AMVP, the block merging algorithm constructs a list containing possible merge candidates for each prediction block. The number of candidates is defined by "NumMergeCands," which ranges from 1 to 5 and is shown in the slice header. Candidates are inferred from prediction blocks in the time image co-located with spatially adjacent prediction blocks. Possible sample positions for prediction blocks considered to be candidates are equal to the positions shown in Figure 36. An example of the block merging algorithm with possible prediction block dividers in HEVC is illustrated in Figure 37. The thick lines in Figure 37(a) define all prediction blocks that are merged into a single region to hold specific motion data. This motion data is sent only to block S. The current prediction block to be coded is indicated by "X". The predicted block in the removed region does not yet have any related predicted data because it is the successor to predicted block X in the block scan order. The dots indicate the sample positions of adjacent blocks that are possible spatial integration candidates. Before the possible candidates are inserted into the predictor list, a redundancy check for spatial candidates is performed as shown in Figure 37(b).

[0249] If the spatial and temporal number of candidates is less than "NumMergeCands", additional candidates are provided by merging with existing candidates or by inserting zero-motion vector candidates. If a candidate is added to the list, it has an index used to identify the candidate. As new candidates are added to the list, the merged index (starting from 0) is incremented, with the list completed by the last candidate identified by index "NumMergeCands" - 1. Fixed-length codewords are used to encode the merged candidate index to ensure independent operation of candidate list derivation and bitstream parsing.

[0250] The following section describes a method for using a multiple enhancement layer predictor, which includes a predictor obtained from the base layer, to encode the motion parameters of the enhancement layer. Motion information already encoded for the base layer can be used to significantly reduce the motion data rate while encoding the enhancement layer. This method includes the possibility of directly obtaining all the motion data for the prediction block from the base layer. In this case, the additional motion data does not need to be encoded. In the following description, a term prediction block refers to a prediction unit in HEVC (an M×N block in H.264 / AVC) and is understood as a general set of samples in an image.

[0251] The first part of the current section concerns extending the list of motion vector prediction candidates with a base layer motion vector predictor (see Example K). Base layer motion vectors are added to the motion vector predictor list during enhancement layer coding. This is achieved by inferring one or more motion vector predictors from the co-located prediction blocks of the base layer and using them as candidates in the list of predictors for motion compensation prediction. The co-located prediction blocks of the base layer are located in the center, left, top, right, or bottom of the current block. If the base layer prediction block at a selected location does not contain motion relation data or lies outside the current range and is therefore not currently accessible, then a binary location can be used to infer a motion vector predictor. These binary locations are represented in Figure 38.

[0252] The motion vectors inferred in the base layer are scaled according to the resolution ratio before they are used as predictor candidates. Similar to motion vector differences, an index describing a list of motion vector predictor candidates is sent to a predictor block specifying the final motion vector to be used for motion-compensated prediction. In contrast to the scalable extensions of the H.264 / AVC standard, the embodiments presented herein do not constitute the use of motion vector predictors in co-located blocks within a reference image—rather, it is available in a list within another predictor and described by the index sent. In one embodiment, the motion vector is obtained from the center position C1 of the co-located prediction block in the base layer and is added to the beginning of the candidate list as the first entry. The candidate list of motion vector predictors is expanded by one item. If there is no motion data available in the base layer for sample position C1, the list structure is left untouched. In another embodiment, any sequence of sample positions in the base layer is checked for motion data. If motion data is found, the motion vector predictor for the corresponding position is inserted into the candidate list and is available for motion compensation prediction in the enhancement layer. Furthermore, the motion vector predictor obtained from the base layer is inserted into the candidate list for any other position in the list. In yet another embodiment, if certain constraints are met, the base layer motion predictor is simply inserted into the candidate list. These constraints include the value of the integration flag of the co-located reference block, which must be equal to zero. Another constraint is the width of the prediction block in the enhancement layer being equal to the width of the co-located prediction block in the base layer with respect to the resolution ratio. For example, in the application of K x spatial scalability, suppose the width of the co-located blocks in the base layer is equal to N, and the width of the prediction blocks to be encoded in the enhancement layer is K. * If it is equal to N, only the motion vector predictor is inferred. In another embodiment, one or more motion vector predictors from several sample locations in the base layer are added to the candidate list in the enhancement layer. In yet another embodiment, candidates having motion vector predictors inferred from co-located blocks replace spatial or temporal candidates in the list rather than expanding the list. It is also possible to include multiple motion vector predictors derived from base layer data in the motion vector predictor candidate list.

[0253] The second part concerns expanding the list of integrated candidates with base layer candidates (see Example K). Motion data from one or more co-located blocks in the base layer is added to the integrated candidate list. This method allows for the possibility of creating integrated regions that share specific motion parameters across the base and enhancement layers. As in the previous section, as represented in Figure 38, the base layer block covering the co-located sample at the central position is not limited to this central position but can be obtained from any position in its immediate vicinity. If any motion data is not available or accessible with respect to a given position, a binary position can be selected to infer possible integrated candidates. Before the obtained motion data is inserted into the integrated candidate list, it is scaled according to the resolution ratio. An index describing the integrated candidate list is sent and defines the motion vectors, which are used for motion-compensated prediction. However, the method also suppresses possible motion predictor candidates that depend on the motion data of the prediction blocks in the base layer. In this embodiment, the motion vector predictor of a co-located block in the base layer covering sample position C1 in Figure 38 is considered a possible merge candidate for encoding the current prediction block in the enhancement layer. However, if the “merge_flag” of the reference block is equal to 1, or if the co-located reference block contains no motion data at all, the motion vector predictor is not inserted into the list. In any other case, the obtained motion vector predictor is added to the merge candidate list as the second entry. Note that in this embodiment, the length of the merge candidate list is preserved and not expanded. In another embodiment, as shown in Figure 38, one or more motion vector predictors are obtained from prediction blocks covering any of the sample positions so that they are added to merge the candidate list. In another embodiment, one or more motion vector predictors from the base layer are added to the merge candidate list at any position. In another embodiment, if certain restrictions are permitted, only one or more motion vector predictors are added to the merge candidate list. Such restrictions include the width of the prediction blocks in the enhancement layer matching the width of the co-located blocks in the base layer (with respect to the resolution ratio described in the section of the previous embodiment for motion vector prediction). Another restriction in another embodiment is a value of "merge_flag" equal to 1. In another embodiment, the length of the merged candidate list is extended by the number of motion vector predictors inferred from the co-located reference blocks in the base layer.

[0254] The third part of this specification relates to rearranging the motion parameter (or integration) candidate list using base layer data (see Example L), and describes the process of rearranging the integration candidate list according to information already encoded in the base layer. If the co-located base layer block covering the sample of the current block is a motion-compensated prediction with candidates derived from a particular original, then the corresponding enhancement layer candidate from the equivalent original (if any) is placed at the top of the integration candidate list as the first entry. This step is equivalent to describing this candidate having the lowest index. The lowest index assigns the simplest code word to this candidate. In this embodiment, the co-located base layer block is motion-compensated and predicted together with candidates arising from the prediction block covering sample position A1, as shown in Figure 38. If the merged candidate list of prediction blocks in the enhancement layer includes a candidate whose motion vector predictor arises from the corresponding sample position A1 in the enhancement layer, then this candidate is placed in the list as the first entry. As a result, this candidate is indexed by index 0 and thus assigned the shortest fixed-length codeword. In this embodiment, this step is performed with respect to the merged candidate list in the enhancement layer after the derivation of the motion vector predictor of the co-located base layer block. Therefore, the reordering process assigns the lowest index to the candidate arising from the corresponding block as the motion vector predictor of the co-located base layer block. The second lowest index is assigned to the candidate derived from the co-located block in the base layer, as described in the second part of this section. Furthermore, the reordering process is performed only if the "merge_flag" of the co-located block in the base layer is equal to 1. In another embodiment, the reordering process is performed regardless of the value of the "merge_flag" of the co-located prediction block in the base layer. In another embodiment, a candidate with a corresponding original motion vector predictor is placed at any position in the merged candidate list. In another embodiment, the reordering process removes all other candidates from the merged candidate list. Here, only candidates whose motion vector predictor has the same original motion vector predictor used for motion compensation prediction of the co-located block in the base layer remain in the list. In this case, a single candidate is utilized, and no index is sent at all.

[0255] The fourth part of this specification relates to reordering a list of motion vector predictor candidates using base layer data (see Example L), and embodies the process of reordering a list of motion vector predictor candidates using the motion parameters of a base layer block. If a co-located base layer block covering a sample of the current predictor block uses motion vectors from a particular original, then the corresponding motion vector predictor from the original in the enhancement layer is used as the first entry in the motion vector predictor list of the current predictor block. This results in assigning the cheapest codeword to this candidate. In this embodiment, the co-located base layer blocks are motion-compensated and predicted together with candidates arising from the prediction block covering sample position A1, as shown in Figure 38. If the list of motion vector predictor candidates for a block in the enhancement layer includes a candidate whose motion vector predictor arises from the corresponding sample position A1 in the enhancement layer, this candidate is placed in the list as the first entry. As a result, this candidate is indexed by index 0 and therefore assigned the shortest fixed-length codeword. In this embodiment, this step is performed with respect to the list of motion vector predictors in the enhancement layer after the derivation of the motion vector predictors for the co-located base layer blocks. Therefore, the reordering process assigns the lowest index to the candidate arising from the corresponding block as the motion vector predictor for the co-located base layer block. The second lowest index is assigned to the candidate derived from the co-located block in the base layer, as described in the first part of this section. Furthermore, the reordering process is performed only if the "merge_flag" of the co-located blocks in the base layer is equal to 0. In another embodiment, the reordering process is performed regardless of the value of the "merge_flag" of the co-located prediction blocks in the base layer. In another embodiment, a candidate with a corresponding original motion vector predictor is placed at any position in the motion vector predictor candidate list.

[0256] The following concerns enhancement layer coding of transformation coefficients.

[0257] In state-of-the-art video and image coding, the residuals of the predicted signal are previously transformed, and the resulting quantized transformation coefficients are shown in the bitstream. This coefficient coding is then followed by a fixed scheme.

[0258] Depending on the transformation size (with respect to luma residuals: 4×4, 8×8, 16×16, 32×32), different scanning directions are defined. Given the first and last positions in the scanning order, these scans must determine which coefficient positions may be important and, as a result, need to be encoded. In all scans, the last position must be shown in the bitstream, but the first coefficient is set to be the DC coefficient at position (0,0). The bitstream is encoded by encoding the (horizontal) x and (vertical) y positions within the transformation block. Starting from the last position, signaling of important coefficients is performed in reverse scanning order until the DC position is reached.

[0259] For transformation sizes 16x16 and 32x32, only one scan, namely a "diagonal scan," is defined. However, for transformation blocks of sizes 2x2, 4x4, and 8x8, "vertical" and "horizontal" scans are also available. However, the use of vertical and horizontal scans is limited to the residuals of the intra-predictive coding unit. The scan actually used is derived from the instruction mode of that intra-prediction. Instruction modes with indices in the range of 6 and 14 result in vertical scans. Instruction modes with indices in the range of 22 and 30 result in horizontal scans. All residual instruction modes result in diagonal scans.

[0260] Figure 39 shows the diagonal, vertical, and horizontal scans defined for a 4x4 transform block. The coefficients of larger transforms are subdivided into 16 subgroups of coefficients. These subgroups allow for hierarchical coding of important coefficient locations. Subgroups indicated as unimportant do not contain important coefficients. Transforms for 8x8 and 16x16 are shown, along with their scans and the subgroup divisions to which they relate in Figures 40 and 41, respectively. Large arrows indicate the scan order of the coefficient subgroups.

[0261] In zigzag scanning, for blocks larger than 4x4, subgroups consist of 4x4 pixel blocks scanned in the zigzag scan. Subgroups are scanned in a zigzag manner. Figure 42 shows vertical scanning for 16x16 transformation, as proposed in JCTVC-G703.

[0262] The following paragraphs describe extensions to transform coefficient coding. These include the introduction of new scanning modes (methods for scanning into transform blocks and assigning modified coding to key coefficient positions). These extensions enable better fitting of different coefficient distributions within the transform block, and as a result, achieve coding gain within the ratio distortion function.

[0263] A new realization for vertical and horizontal scanning patterns is introduced for 16x16 and 32x32 conversion blocks. In contrast to previously proposed scanning patterns, the sizes of the scanning subgroups are 16x1 for horizontal scanning and 1x16 for vertical scanning, respectively. Subgroups with sizes of 8x2 and 2x8 are also selected, respectively. The subgroups themselves are scanned in the same manner.

[0264] Vertical scanning is efficient due to the transformation coefficients located within something like an expanded row. This is observed in images that include horizontal edges.

[0265] Horizontal scanning is efficient due to the transformation coefficients located within something resembling an expanded column. This is observed in images containing vertical edges.

[0266] Figure 43 illustrates the implementation of vertical and horizontal scanning for a 16x16 transformation block. Each coefficient subgroup is defined as either a row or a column. Vertical and horizontal scanning is the introduced scanning pattern, which allows for the encoding of coefficients within rows by scanning column by column. For a 4x4 block, the first row is scanned, followed by the remainder of the first column, then the remainder of the second row, then the remainder of the coefficients in the second column, then the remainder of the third row, and finally the remainder of the fourth column and row.

[0267] For larger blocks, the block is divided into 4x4 subgroups. These 4x4 blocks are scanned using vertical and horizontal scanning, and the subgroups are scanned using vertical and horizontal scanning itself.

[0268] Vertical and horizontal scanning is used when the coefficients are located in the first row and column of a block. In this way, the coefficients are scanned faster than when using another scanning method (e.g., diagonal scanning). This is observed for images that contain both horizontal and vertical edges.

[0269] Figure 44 shows vertical and horizontal scanning with respect to a 16x16 transformation block.

[0270] Other scanning methods are also possible. For example, all combinations between scanning and subgroups can be used. For instance, using a horizontal scan for a 4x4 block and a diagonal scan for the subgroups, the appropriate choice of scanning is applied by selecting a different scan for each subgroup.

[0271] It should be noted that the transformation coefficients are rearranged after quantization on the encoder side, allowing for different scans in the same way that conventional encoding is used. On the decoder side, the transformation coefficients are rearranged before (or after and before) scaling and inverse transformations, after conventional decoding.

[0272] The different parts of the base layer signal are used to obtain encoding parameters from the base layer signal. The base layer signal includes the following: • Co-configured reconfigured base layer signal • Co-located residual base layer signal • The estimated residual signal of the enhancement layer, obtained by subtracting the enhancement layer prediction signal from the reconstructed base layer signal. • Image segmentation of the base layer frame

[0273] [Gradient parameters] The gradient parameters are derived as follows: For each pixel in the examined block, a gradient is calculated. From these gradients, the magnitude and angle are calculated. The angle that occurs within the block is the block angle. The angle is rounded to use only three directions: horizontal (0°), vertical (90°), and diagonal (45°).

[0274] [Edge detection] The edge detector is applied to the investigated blocks as follows: First, the block is smoothed by an n×n smooth filter (e.g., a Gaussian filter). A gradient matrix of size m x m is used to calculate the gradient of each pixel. The size and angle of every pixel are calculated. The angles are rounded to use only three directions: horizontal (0°), vertical (90°), and diagonal (45°). For every pixel with a size greater than a predetermined threshold of 1, adjacent pixels are checked. If an adjacent pixel has a size greater than a threshold of 2 and has the same angle as the current pixel, the counter for this angle is incremented. For the entire block, the counter with the highest value is selected as the block's angle.

[0275] [Obtain the base layer coefficients through the previous transformation] For a particular TU, the investigated and co-arranged signals (reconstructed base layer signal / residual base layer signal / estimated enhancement layer signal) are transformed in the frequency domain to obtain coding parameters from the frequency domain of the base layer signal. Preferably, this is done using the same transformation used by that particular enhancement layer TU. The resulting base layer transformation coefficients may or may not be quantized. Ratio-distortion quantization with a modified lambda is used to obtain a coefficient distribution comparable to that of the enhancement layer block.

[0276] [Scan effectiveness score for a specific distribution and scan] The scan effectiveness score for a specific important coefficient distribution is defined as follows: Let's represent each position in the investigated blocks by an index, in the order of the scanned scan. Then, the sum of the index values ​​of the important coefficient positions is defined as the effective score of this scan. As a result, the smaller the score of a scan, the better the efficiency of a particular distribution.

[0277] [Selecting the appropriate scan pattern for conversion coefficient coding] If several scans are available for a particular TU, a rule must be defined to select only one of the scans.

[0278] [Method for selecting a scanning pattern] The selected scan can be obtained directly from the already decoded signal (without transmitting any additional data). This can be achieved either by basing it on the characteristics of the co-arranged base layer signal or by utilizing only the enhancement layer signal. The scanning pattern can be obtained from the EL signal by the following method. • The aforementioned state-of-the-art derivation rules. • Use a scanning pattern for chromatic residuals selected for co-located luminance residuals. - Define a fixed mapping between the encoding mode and the scan pattern used. • Obtain the scan pattern from the last important coefficient position (proportional to the estimated fixed scan pattern). In a preferred embodiment, the scanning pattern is selected depending on the last position already decoded, as follows:

[0279] The last position is represented as x and y coordinates within the transformation block and is already decoded (a fixed scan pattern is estimated for the decoding process of the last position, with respect to the scan that depends on the last one to be encoded; this is the scan pattern of the leading edge of that TU). Let T be a defined threshold that depends on a particular transformation size. If neither the x nor y coordinates of the last significant position exceed T, then a diagonal scan is selected.

[0280] Otherwise, x is compared to y. If x exceeds y, horizontal scanning is selected and vertical scanning is not selected. The preferred value of T for a 4x4 TU is 1. The preferred value of T for a TU larger than 4x4 is 4.

[0281] In another preferred embodiment, the derivation of the scanning pattern described in the previous embodiment is limited to TUs of sizes 16×16 and 32×32. It is further limited to luminance signals only.

[0282] Furthermore, the scan pattern is obtained from the BL signal. Any of the aforementioned encoding parameters can be used to obtain a selected scan pattern from the base layer signal. In particular, the gradients of the co-arranged base layer signals are calculated and compared to a predefined threshold, and / or potentially found edges can be utilized.

[0283] In a preferred embodiment, the scanning direction is obtained depending on the block gradient angle, as follows: For gradients quantized in the horizontal direction, vertical scanning is used. For gradients quantized in the vertical direction, horizontal scanning is used. Otherwise, diagonal scanning is selected.

[0284] In another preferred embodiment, the scan pattern is obtained as described in the previous embodiment, but only with respect to those transformation blocks where the number of block angle occurrences exceeds a threshold. The remaining transformation units are decoded using the scan pattern of the leading edge of the TU.

[0285] Assuming that the base layer coefficients for co-located blocks are valid and clearly shown in the base layer data stream or calculated by the previous transformation, the base layer coefficients can be used in the following way: For each available scan, the cost of encoding the base layer coefficients is evaluated. The scan with the lowest cost is used to decode the enhancement layer coefficients. The effective score for each available scan is calculated for the base layer coefficient distribution. The scan with the lowest score is used to decode the enhancement layer coefficients. The distribution of base layer coefficients within the transformation block is classified into one of a predefined set of distributions associated with a particular scanning pattern. The scanning pattern is selected based on the last important base layer coefficients.

[0286] If the co-located base layer blocks are predicted using intra-prediction, then the intra-direction of that prediction can be used to obtain the enhancement layer scanning pattern.

[0287] Furthermore, the transformation size of the jointly arranged base layer blocks is used to obtain the scan pattern.

[0288] In a preferred embodiment, the scan pattern is derived from the BL signal only with respect to TUs representing the residuals of the INTRA_COPY mode predicted blocks. These co-located base layer blocks are then intra-predicted. A modified state-of-the-art scan selection is used for these blocks. In contrast to the state-of-the-art scan selection, the intra-prediction direction of the co-located base layer blocks is used to select the scan pattern.

[0289] The scan pattern index is shown within the bitstream (see Example R). The scan pattern of the conversion block is selected by the encoder in terms of rate distortion and is then shown in the bitstream.

[0290] A specific scan pattern can be encoded by indicating its index within a list of available candidate scan patterns. This list can be a fixed list of scan patterns defined for a particular transformation size, or it can be dynamically filled during the decoding process. Dynamically filling the list allows for the adaptive selection of those scan patterns. The scan pattern likely encodes a particular coefficient distribution most efficiently. In doing so, the number of available scan patterns for a particular TU can be reduced, and consequently, signaling the index within that list is not very expensive. If the number of scan patterns in a particular list can be reduced to one, then signaling is not necessary. For a particular TU, the process of selecting a candidate scan pattern follows predetermined rules that utilize any of the aforementioned encoding parameters and / or specific characteristics of that particular TU. These include: TU represents the residual of the luminance / chrominance signal. TU has a specific size. • TU represents the residual for a specific prediction mode. The last important position within TU is known by the decoder and belongs to a specific subdivision of TU. • TU is a portion of one I / B / P slice. The coefficients of TU are quantized using specific quantization parameters.

[0291] In a preferred embodiment, the list of scan pattern candidates includes three scans for all TUs: “diagonal scan”, “vertical scan”, and “horizontal scan”.

[0292] Another embodiment can be obtained by including any combination of scanning patterns in the candidate list.

[0293] In certain preferred embodiments, the list of candidate scanning patterns includes one of the following scans: diagonal scanning, vertical scanning, and horizontal scanning.

[0294] However, the scan pattern selected by the (previously) cutting-edge scan derivation is initially set to be in the list. Another candidate is added to the list only if a particular TU has a size of 16×16 or 32×32. The order of the remaining scan patterns depends on the last important coefficient position.

[0295] (Note: Diagonal scanning is always the first pattern in the list that estimates 16x16 and 32x32 transformations.)

[0296] If the x-coordinate exceeds the y-coordinate, then the horizontal scan is selected next. The vertical scan is then placed in the last position. Otherwise, the vertical scan is placed in the second position, following the horizontal scan.

[0297] Another preferred embodiment can be obtained by further restricting the conditions, as there is one or more candidates in the list.

[0298] In another embodiment, if the coefficients of the conversion block represent the residuals of the luminance signal, then vertical and horizontal scanning are simply added to the candidate lists of 16x16 and 32x32 conversion blocks.

[0299] In another embodiment, if both the x and y coordinates of the last important position are greater than a certain threshold, the vertical and horizontal scan is added to the candidate list of transformation blocks. This threshold depends on the mode and / or TU size. A preferred threshold is 3 for all sizes greater than 4x4 and 1 for 4x4 TUs.

[0300] In another embodiment, if either the x or y coordinate of the last significant position is greater than a certain threshold, the vertical and horizontal scan is simply added to the candidate list of transformation blocks. This threshold depends on the mode and / or TU size. A preferred threshold is 3 for all sizes greater than 4x4 and 1 for 4x4 TUs.

[0301] In another embodiment, if both the x and y coordinates of the last important position are greater than a certain threshold, the vertical and horizontal scan is simply added to the candidate list of 16x16 and 32x32 transformation blocks. This threshold can be a size-dependent mode and / or TU. A preferred threshold is 3 for all sizes greater than 4x4 and 1 for 4x4 TUs.

[0302] In another embodiment, if either the x or y coordinates of the last important position are greater than a certain threshold, the vertical and horizontal scan is simply added to the candidate list of 16x16 and 32x32 transformation blocks. This threshold is a size-dependent mode and / or TU. A preferred threshold is 3 for all sizes greater than 4x4 and 1 for 4x4 TUs.

[0303] In any of the embodiments described, a specific scan pattern is shown in the bitstream. The signaling itself is performed at different signaling levels. In particular, signaling may be performed at any node (all sub-TUs of that node, which use the shown scan and the same candidate list index) of the residual quadtree, at the CU / LCU level, or at the slice level, with respect to each TU that descends within a subgroup of TUs having the shown scan pattern.

[0304] The index in the candidate list can be transmitted using fixed-length coding, variable-length coding, arithmetic coding (including context-adaptive binary arithmetic coding), or PIPE coding. If context-adaptive coding is used, the context is obtained based on the parameters of adjacent blocks, the coding mode described above, and / or the characteristics of the particular TU itself.

[0305] In a preferred embodiment, context-adaptive coding is used to indicate the index in the list of candidate scan patterns of the TU. However, the context model is obtained based on the transformed size and / or position of the last significant position in the TU. Any of the methods described above for obtaining a scan pattern can be used to obtain a contextual model for a particular TU, which is shown in the obvious scan pattern.

[0306] The following modifications are used in the enhancement layer to encode the last critical scan location. • Separate context models are used for all or a subset of encoding modes, using base layer information. It is also possible to use different context models for different modes, each with its own base layer information. The context model can depend on data within co-located base layer blocks (e.g., transformation coefficient distribution within the base layer, gradient information within the base layer, last scan position within the co-located base layer block). • The last scan position can be encoded as the difference between that position and the last base layer scan position. If the last scan position is encoded within the TU by signaling the x and y positions, then the contextual model of the second shown coordinates can depend on the values ​​of the first signaling. To obtain a scan pattern independent of the last significant location, one of the aforementioned methods is used to obtain a contextual model that indicates the last significant location.

[0307] In certain versions, the derivation of the scan pattern depends on the last important position: If the last scan position is encoded within the TU by indicating its x and y positions, then the contextual model of the second coordinate can rely on those scan patterns, which are still possible candidates, given that the first coordinate is already known. If the last scan position is encoded within the TU by indicating its x and y positions, then the contextual model of the second coordinate can depend on whether a unique scan pattern is already selected when the first coordinate is already known.

[0308] In another version, the scan pattern derivation is independent of the last important location. The context model can depend on the scan patterns used within a particular TU. • Any of the methods described above to obtain the scan pattern can be used to obtain a contextual model to indicate the last important location.

[0309] To encode important locations and important flags within the TU (subgroup flags and / or important flags for a single transformation coefficient), the following modifications are used within the enhancement layer, respectively: • Separate context models are used for all or a subset of encoding modes that use base layer information. It is also possible to use different context models for different modes that have base layer information. The contextual model can rely on data within co-located base layer blocks (e.g., the number of important transformation coefficients for a given frequency position). • One of the methods described above is used to obtain a scan pattern, and then to obtain a contextual model to indicate important locations and / or their levels. A generalized template is used to evaluate both the number of significant already encoded transformation coefficient levels in the spatial neighborhood of the coefficients to be encoded, and the number of significant transformation coefficients in the co-located base layer signals at similar frequency positions. A generalized template is used to evaluate both the significant number of already encoded transformation coefficient levels in the spatial neighborhood of the coefficients to be encoded, and the significant transformation coefficient levels in the co-located base layer signals at similar frequency locations. The context modeled for subgroup flags depends on the scan pattern used and / or the specific transformation size.

[0310] Different context initialization tables are used for the base layer and the enhancement layer. Context model initialization for the enhancement layer is modified in the following way: The enhancement layer uses a separate set of initialization values. The enhancement layer uses separate sets of initialization values ​​for different operating modes (spatial / temporal, or quality scalability). The enhancement layer context model, which has corresponding parts within the base layer, uses the state of those corresponding parts as its initial state. The algorithm for obtaining the initial state of the context model is dependent on the base layer QP and / or delta QP.

[0311] Next, the possibility of encoding a suitable enhancement layer later on is explained using the base layer data. This section describes how to generate an enhancement layer prediction signal in a scalable video coding system. The method uses a base layer decoded from image sample information to infer the values ​​of prediction parameters. The values ​​of the prediction parameters are not transmitted in the encoded video bitstream, but are used to form the prediction signal for the enhancement layer. Thus, the overall bit ratio required to encode the enhancement layer signal is reduced.

[0312] State-of-the-art hybrid video encoders decompose the source image into blocks of varying sizes, typically following a hierarchical structure. For each block, the video signal is predicted from spatially adjacent blocks (intra-prediction) or from previously temporally encoded images (inter-prediction). The difference between the prediction and the actual image is the transformation and quantization. The resulting prediction parameters and transformation coefficients are entropy-encoded to form the encoded video bitstream. Matching decoders follow the steps in the reverse order… Scalable video, which encodes a bitstream, consists of different layers: a base layer that provides the complete decodeable video, and an enhancement layer that is added and used for decoding. The enhancement layer can provide higher spatial resolution (spatial scalability), temporal resolution (temporal scalability), or quality (SNR scalability). In older standards like H.264 / AVC SVC, syntactic elements such as motion vectors, reference image indices, or intra-prediction modes are predicted directly from their corresponding syntactic elements in the encoded base layer. Within the enhancement layer, the mechanism exists to switch between them at the block level, using predicted signals obtained from base layer syntactic elements or predicted from other enhancement layer syntactic elements or decoded enhancement layer samples.

[0313] In the following section, the base layer data is used by the decoder to obtain the enhancement layer parameters.

[0314] [Method 1: Derivation of candidate motion parameters] For each block (a) of the image in the spatial or qualitative enhancement layer, the corresponding block (b) of the image in the base layer is determined. It covers the same image region. The interprediction signal for block (a) of the enhancement layer is formed using the following method: 1. Candidate motion compensation parameter sets are determined, for example, from temporally or spatially adjacent enhancement layer blocks or their derivations. 2. Motion compensation is performed to form interpredictive signals within the enhancement layer with respect to each candidate's motion compensation parameter set. 3. The best motion compensation parameter set is selected by minimizing the magnitude of the error between the predicted signal for the enhancement layer block (a) and the reconstructed signal for the base layer block (b). For spatial scalability, the base layer block (b) is spatially upsampled using an interpolation filter.

[0315] A motion compensation parameter set includes a specific combination of motion compensation parameters.

[0316] Motion compensation parameters can be a motion vector, a reference image index, and a choice between one or two predictions and another parameter.

[0317] In the binary embodiment, candidate motion compensation parameter sets are used from the base layer block. Interpretation is performed within the base layer (using the reference image of the base layer). To apply the magnitude of the error, the reconstructed signal of base layer block (b) is used directly without upsampling. The selected optimal motion compensation parameter set is applied to the reference image of the enhancement layer to form the prediction signal of block (a). When the motion vector is applied in the spatial enhancement layer, the motion vector is scaled according to the resolution change. Both the encoder and decoder can perform the same prediction steps to create the same predicted signal by selecting the optimal set of motion compensation parameters from the available candidates. These parameters are not shown in the encoded video bitstream.

[0318] The selection of the prediction method is indicated in the bitstream and encoded using entropy coding. Within a hierarchical block subdivision structure, this coding method can be selected at any sublevel, or alternatively, only a subset of the coding hierarchy. In an alternative embodiment, the encoder can transmit an improved motion parameter set prediction signal to the decoder. The improved signal differentially contains the encoded values ​​of the motion parameters. The improved signal is entropy coded.

[0319] In an alternative embodiment, the decoder generates a list of best candidates. The indices of the motion parameter set used are shown in the encoded video bitstream. The indices are entropy encoded. In the embodiment, the list can be ordered by increasing the magnitude of the error.

[0320] The example uses a HEVC Adaptive Motion Vector Prediction (AMVP) candidate list to generate candidate motion compensation parameter sets. Another embodiment uses a list of HEVC integrated mode candidates to generate candidate motion compensation parameter sets.

[0321] [Method 2: Derivation of motion vectors] For a block (a) of the spatial or quality enhancement layer image, a corresponding block (b) of the base layer image covering the same image region is determined.

[0322] The interprediction signal for block (a) of the enhancement layer is formed using the following method: 1. A motion vector predictor is selected. 2. The motion estimation of a defined set of search locations is performed on the reference image of the enhancement layer. 3. For each search position, the magnitude of the error is determined, and the motion vector with the smallest error is selected. 4. The prediction signal for block (a) is formed using the selected motion vector.

[0323] In an alternative embodiment, the search is performed on the reconstructed base layer signal. With respect to spatial scalability, the selected motion vectors are scaled according to the change in spatial resolution before generating the prediction signal in step 4.

[0324] The search location can be either full resolution or sub-pel resolution. The search can also perform multiple steps, for example, determining the best full pel location first, followed by another set of candidates based on the selected full pel location. For instance, the search can terminate quickly if the error magnitude is below a defined threshold.

[0325] Both the encoder and decoder can perform the same prediction step to select the optimal motion vector from the candidates and generate the same predicted signal. These vectors are not shown in the encoded video bitstream.

[0326] The selection of the prediction method is indicated in the bitstream and encoded using entropy coding. Within a hierarchical block subdivision structure, this coding method is selected at any sublevel, or at only a subset of alternative coding hierarchies. In an alternative embodiment, the encoder can transmit an improved motion vector prediction signal to the decoder. The improved signal is entropy coded.

[0327] The example uses the algorithm described in Method 1 to select a motion vector predictor.

[0328] Another embodiment uses an HEVC-appropriate motion vector prediction (AMVP) method to select motion vector predictors from temporally or spatially adjacent blocks in the enhancement layer.

[0329] [Method 3: Derivation of Intra Prediction Mode] For each block (a) in the enhancement layer (n) image, a corresponding block (b) is determined that covers the same region in the reconstructed base layer (n-1) image.

[0330] In a scalable video decoder, for each base layer block (b), an intra-prediction signal is formed using an intra-prediction mode (p) inferred by the following algorithm. 1) The intra-prediction signal is generated for each available intra-prediction mode, following the rules for intra-prediction in the enhancement layer, but using sample values ​​from the base layer. 2) Best prediction mode (p best The value is determined by minimizing the magnitude of the error (e.g., the sum of absolute differences) between the intra-predicted signal and the decoded base layer block (b). 3) The prediction selected in step 2) (p best The ) mode is used to generate a prediction signal for enhancement layer block (a) according to an intra-prediction rule for the enhancement layer.

[0331] Both the encoder and decoder are in the best prediction mode (p best ) can be selected and the same steps can be performed to form a matched prediction signal. Actual intra-prediction mode (p best ) is not thus shown in the encoded video bitstream.

[0332] The selection of the prediction method is indicated in the bitstream and encoded using entropy coding. Within a hierarchical block subdivision structure, this coding mode is selected at any sublevel, or alternatively, at only a subset of the coding hierarchy. An alternative embodiment uses samples from the enhancement layer in step 2) to generate an intra-prediction signal. With respect to the spatially scalable enhancement layer, the base layer is upsampled using an interpolation filter to apply the magnitude of the error.

[0333] An alternative embodiment involves using an enhancement layer block with a smaller block size (a i ) can be divided into multiple blocks (for example, a 16x16 block (a) can be divided into 16 4x4 blocks (a i It can be divided into (a). The algorithm described above is divided into each subblock (a i ) and the corresponding base layer block (b i ) applies to block (a i After the prediction of ), residual coding is applied, and the result is block (a i+1 It is used to predict ).

[0334] An alternative embodiment is the predicted intra-prediction mode (p best To determine ) (b) or (b i Use sample values ​​around the ) . For example, a 4x4 block (a) of the spatial enhancement layer (n) i ) corresponds to the 2x2 base layer block (b i When (b) has i Samples around ) predict the predicted intra-prediction mode (pbest ) is used for the determination of 4x4 blocks (c i Used to form ).

[0335] In an alternative embodiment, the encoder can transmit an improved intra-predictive direction signal to the decoder. For example, in a video codec such as HEVC, most intra-predictive modes correspond to the angles used by boundary pixels to form the predictive signal. The offset to the optimal mode is the predicted intra-predictive mode (p best It is transmitted as the difference to ). The improved mode is entropy coded.

[0336] Intra-predicted modes are typically encoded based on their probabilities. In H.264 / AVC, the most likely modes are determined based on the modes used in the (spatial) neighborhood of a block. In the HEVC list, the most likely modes are created. These most likely modes are selected using fewer symbols in the bitstream than the total number of modes required. An alternative embodiment is to have the predicted intra-predicted modes (p) for block (a) (determined as described in the algorithm above) as members of the list of most likely modes, or as the most likely modes. best Use ).

[0337] [Method 4: Intra-prediction using boundary regions] In a scalable video decoder for forming an intra-predictive signal for a block (a) of a scalable or quality enhancement layer (see Figure 45), lines of samples (b) from the surrounding region of the same layer are used to fill the block region. These samples are taken from an already encoded region (usually, but not required on the upper and left boundaries).

[0338] The following alternative transformations are used to select these pixels. a) If pixels in the surrounding region have not yet been encoded, their pixel values ​​are not used to predict the current block. b) If a pixel in the surrounding region has not yet been encoded, its pixel value can be obtained from an already encoded adjacent pixel (for example, by iteration). c) If the pixels in the surrounding region have not yet been encoded, the pixel values ​​are obtained from the pixels in the corresponding region of the decoded base layer image.

[0339] To form the intra-prediction of block (a), the adjacent lines of pixel (b) (obtained as described above) are each line (a) of block (a) j It is used as a template to fill in the blanks.

[0340] Block (a) line (a j ) are filled one by one along the x-axis. To achieve the best possible prediction signal, the rows of the template sample (b) are related to the line (a j Prediction signal (b') for ) j ) Shifts along the y-axis to form.

[0341] To find the optimal prediction within each line, shift offset (o j ) is the resulting predicted signal (a j This is determined by minimizing the magnitude of the error between the sample value of the corresponding line in the base layer and the sample value of the corresponding line in the base layer.

[0342] For example, (o j If ) is a non-integer value, the interpolation filter is as shown in (b'7), (a j This can be used to map the value of (b) to the integer sample position of ).

[0343] If spatial scalability is used, an interpolation filter is used to create a matching number of sample values ​​for the corresponding lines in the base layer.

[0344] The filling direction (x-axis) can be horizontal (left and right), vertical (up and down), diagonal, or any other angle. The sample used for template line (b) is a sample directly adjacent to the block along the x-axis. Template line (b) is shifted along the y-axis, forming a 90° angle with respect to the x-axis.

[0345] To find the optimal orientation of the x-axis, a filler intra-prediction signal is generated for block (a). The angle with the smallest error magnitude between the prediction signal and the corresponding base layer block is selected. The number of possible angles can be limited.

[0346] Both the encoder and decoder run the same algorithm to determine the best predicted angle and offset. No explicit angle or offset information needs to be shown in the bitstream. In the alternative embodiment, only the sample of the base layer image is offset (o i It is used to determine ).

[0347] In an alternative embodiment, the predicted offset (o i The improvement (e.g., the difference value) is shown in the bitstream. Entropy coding can be used to encode the improved offset value.

[0348] In an alternative embodiment, the improvement in the predicted direction (e.g., the difference value) is shown in the bitstream. Entropy coding is used to encode the improvement direction value.

[0349] For example, line (b' j If ) is used for prediction, an alternative embodiment uses a threshold to select. For example, if the optimal offset (o j If the error magnitude for ) is less than the threshold, then line (c i ) is a block line (a j This is used to determine the value of the optimal offset (oj If the error magnitude for ) is greater than or equal to the threshold, then the (upsampled) base layer signal is a block line (a j It is used to determine the value of ).

[0350] [Method 5: Alternative Prediction Parameters] Other predictive information is inferred, for example, for block partitioning into subblocks, in the same manner as in methods 1-3.

[0351] With respect to a block (a) of the image in the spatial or quality enhancement layer, the corresponding block (b) of the image in the base layer is determined. It covers the same image region.

[0352] The prediction signal for block (a) of the enhancement layer is formed using the following method. 1) Predictive signals are generated for each possible value of the tested parameter. 2) Best prediction mode (p best This is determined by minimizing the magnitude of the error (e.g., the sum of absolute differences) between the predicted signal and the decoded base layer block (b). 3) The prediction selected in step 2) (p best The ) mode is used to generate the prediction signal for enhancement layer block (a).

[0353] Both the encoder and decoder can perform the same prediction step to select the optimal prediction mode from the possible candidates and generate the same prediction signal. The actual prediction mode is not shown in the encoded video bitstream.

[0354] The choice of prediction method is indicated within the bitstream and can be encoded using entropy coding. Within a hierarchical block subdivision structure, this coding method can be either selected at any sublevel or only with respect to a subset of the coding hierarchy.

[0355] The following description briefly summarizes some of the embodiments described above.

[0356] [Enhancement layer coding with a multiplexing method for generating an intra-prediction signal using reconstructed base layer samples] Main embodiment: With regard to encoding blocks in the enhancement layer, a multiplexing method for generating an intra-prediction signal using samples from the reconstructed base layer is provided, in addition to a method for generating a prediction signal based solely on samples from the reconstructed enhancement layer.

[0357] Sub-examples: The multiplexing method includes the following: The reconstructed (upsampled / filtered) base layer signal is used directly as the enhancement layer prediction signal. The multiplexing method includes the following: The (upsampled / filtered) reconstructed base layer signal is coupled to the spatial intra-prediction signal, where the spatial intra-prediction is obtained based on difference samples with respect to adjacent blocks. The difference sample represents the difference between the reconstructed enhancement layer signal and the (upsampled / filtered) reconstructed base layer signal (see Example A). The multiplexing method includes the following: A conventional spatial intra-prediction signal (obtained using samples from adjacent reconstructed enhancement layers) is coupled to a base layer residual signal (inverse transform of base layer transformation coefficients, or the difference between base layer reconstruction and base layer prediction) (see Example B). The multiplexing method includes the following: The (upsampled / filtered) reconstructed base layer signal is coupled to a spatial intra-prediction signal, where the spatial intra-prediction is obtained based on samples from the reconstructed enhancement layer of adjacent blocks. The final prediction signal is obtained by weighting the spatial prediction signal and the base layer prediction signal in a way that different frequency components use different weightings (see Example C1). This can be achieved, for example, by one of the following: ○The base layer prediction signal is filtered by a low-pass filter, the spatial intra prediction signal is filtered by a high-pass filter, and the resulting filtered signals are added together (see Example C2). ○ The base layer prediction signal and the enhancement layer prediction signal are transformed, and the resulting transformed block is superimposed. Different weighting coefficients are used for different frequency positions (see Example C3). The resulting transformed block is inversely transformed and used as the enhancement layer prediction signal. Alternatively, the resulting transformed coefficients are added to a scaled transmitted transformed coefficient level and then inversely transformed to obtain a block reconstructed before deblocking and i-in-loop processing (see Example C4). Regarding the method of using the reconstructed base layer signal, the following versions are used: This is fixed, or it is indicated at the sequence level, image level, slice level, maximum encoded unit level, or encoded unit level. Or it is created depending on different encoding parameters. ○ Samples of the reconstructed base layer before deblocking and in-loop processing (as samples, adaptive offset filters or adaptive loop filters). ○ Samples of the reconstructed base layer after deblocking and before in-loop processing (as samples, an adaptive offset filter or an adaptive loop filter). ○ A sample of the reconfigured base layer after deblocking and in-loop processing (as an adaptive offset filter or adaptive loop filter), or a sample of the reconfigured base layer during multiple in-loop processing steps (see Example D). • Multiple versions of the method using the (upsampled / filtered) base layer signal are used. The upsampled / filtered base layer signal employed for these versions differs among the interpolation filters used (including interpolation filters that filter integer sample positions). Alternatively, the upsampled / filtered base layer signal for the second version is obtained by filtering the upsampled / filtered base layer signal for the first version. A selection of one of the different versions is indicated at the sequence level, image level, slice level, maximum encoded unit level, and encoded unit level. It is inferred from the characteristics of the corresponding reconstructed base layer signal or the transmitted encoding parameters (see Example E). Different filters are used to upsample / filter the reconstructed base layer signal (see Example E) and the base layer residual signal (see Example F). For base layer blocks where the residual signal is zero, it is replaced by another signal obtained from the base layer (e.g., a high-pass filtered version of the reconstructed base layer block) (see Example G). Regarding modes using spatial intra-prediction, unavailable adjacent samples in the enhancement layer (depending on a specific coding order) are replaced with corresponding samples from the upsampled / filtered base layer signal (see Example H). Regarding modes that use spatial intra-prediction, the coding of the intra-prediction mode is changed. The most likely list of modes includes intra-prediction modes for co-located base layer signals. In certain versions, the enhancement layer image is decoded in a two-stage process. In the first stage, only the blocks that use the base layer signal (without adjacent blocks) or the inter-prediction signal are decoded and reconstructed for prediction. In the second stage, the residual blocks that use adjacent samples for prediction are reconstructed. With respect to the blocks reconstructed in the second stage, the spatial intra-prediction concept is extended (see Example I). Based on the usefulness of the already reconstructed blocks, not only the samples adjacent to the upper and left sides of the current block, but also the samples adjacent to the lower and right sides are used for spatial intra-prediction.

[0358] [Enhancement layer coding with a multiplexing method for generating interprediction signals using reconstructed base layer samples] Main embodiment: With regard to encoding blocks in the enhancement layer, a multiplexing method for generating an interprediction signal using reconstructed base layer samples is provided in addition to a method for generating a prediction signal based solely on reconstructed enhancement layer samples.

[0359] Sub-examples: The multiplexing method includes the following: The conventional interprediction signal (obtained by motion-compensated interpolation of an already reconstructed enhancement layer image) is coupled to the base layer residual signal (inverse transform of the base layer transformation coefficients, or the difference between the base layer reconstruction and the base layer prediction) (upsampled / filtered). The multiplexing method includes the following: The (upsampled / filtered) reconstructed base layer signal is coupled to a motion-compensated prediction signal, which is obtained by a motion-compensated difference image. The difference image represents the difference between the reconstructed enhancement layer signal and the (upsampled / filtered) reconstructed base layer signal with respect to a reference image (see Example J). The multiplexing method includes the following: The (upsampled / filtered) reconstructed base layer signal is coupled to the interprediction signal, where the interprediction is obtained by motion-compensated prediction using the reconstructed enhancement layer image. The final prediction signal is obtained by weighting the interprediction signal and the base layer prediction signal in a way that different frequency components use different weightings (see Example C). This can be achieved, for example, by one of the following: ○The base layer prediction signal is filtered using a low-pass filter, the inter-layer prediction signal is filtered using a high-pass filter, and the resulting filtered signals are added together. ○ The base layer prediction signal and the interlayer prediction signal are transformed, and the resulting transformed block is superimposed. Different weighting coefficients are used for different frequency positions. The resulting transformed block is inversely transformed to obtain a reconstructed block before deblocking and in-loop processing, and is used as the enhancement layer prediction signal, or the resulting transformed coefficients are added to a scaled transmitted transformed coefficient level and then inversely transformed. Regarding the method of using the reconstructed base layer signal, the following versions are used: it is fixed, or it is indicated at the sequence level, image level, slice level, maximum encoded unit level, or encoded unit level; or it is created depending on other encoding parameters. ○ Samples of the reconstructed base layer before deblocking and in-loop processing (as samples, adaptive offset filters or adaptive loop filters). ○ Samples of the reconstructed base layer after deblocking and before in-loop processing (as samples, an adaptive offset filter or an adaptive loop filter). ○ A sample of the reconfigured base layer after deblocking and in-loop processing (as an adaptive offset filter or adaptive loop filter), or a sample of the reconfigured base layer during multiple in-loop processing steps (see Example D). For base layer blocks where the residual signal is zero, it is replaced by another signal obtained from the base layer (e.g., a high-pass filtered version of the reconstructed base layer block) (see Example G). • Multiple versions of the method using the (upsampled / filtered) base layer signal are used. The upsampled / filtered base layer signal employed for these versions differs among the interpolation filters used (including interpolation filters that filter integer sample positions). Alternatively, the upsampled / filtered base layer signal for the second version is obtained by filtering the upsampled / filtered base layer signal for the first version. A selection of one of the different versions is indicated at the sequence level, image level, slice level, maximum encoded unit level, and encoded unit level. It is inferred from the characteristics of the corresponding reconstructed base layer signal or the transmitted encoding parameters (see Example E). Different filters are used to upsample / filter the reconstructed base layer signal (see Example E) and the base layer residual signal (see Example F). Regarding motion compensation prediction of the difference image (the difference between the enhancement layer reconstruction and the upsampled / filtered base layer signal) (see Example J), ​​different interpolation filters are more often used for motion compensation prediction of the reconstructed image. Regarding motion compensation prediction of the difference image (the difference between the enhancement layer reconstruction and the upsampled / filtered base layer signal) (see Example J), ​​the interpolation filter is selected based on the characteristics of the corresponding region in the difference image (or based on the encoding parameters or based on the information transmitted in the bitstream).

[0360] [Enhancement layer motion parameter coding] Main example: Use of multiple enhancement layer predictors and at least one predictor obtained from the base layer for enhancement layer motion parameter coding.

[0361] Sub-examples: • Add the (scaled) base layer motion vector to the motion vector predictor list (see Example K). ○ Use of a base layer block covering the co-located sample at the central position of the current block (another possible derivation). ○Scale motion vector according to resolution ratio. • Add the motion data of the jointly placed base layer blocks to the list of candidates for integration (see Example K). ○ Use of a base layer block covering the co-located sample at the central position of the current block (another possible derivation). ○Scale motion vector according to resolution ratio. ○If the "integration_flag" in the base layer is equal to 1, do not add it. • Reordering of the integration candidate list based on base layer integration information (see Example L) ○If a jointly placed base layer block is merged into a specific candidate, the corresponding enhancement layer candidate will be used as the first entry in the enhancement layer merger candidate list. • Reordering of the motion predictor candidate list based on base layer motion predictor information (see Example L). ○If a co-configured base layer block uses a specific motion vector predictor, the corresponding enhancement layer motion vector predictor will be used as the first entry in the enhancement layer motion vector predictor candidate list. The derivation of the integration index (i.e., the candidates for the current block to be integrated) is based on base layer information within the co-located blocks (see Example M). For example, if a base layer block is integrated into a particular adjacent block, and that is shown in a bitstream where an enhancement layer block is also integrated, then the integration index is not transmitted at all. Instead, the enhancement layer block is integrated into the same adjacent block (but within the enhancement layer) as a co-located base layer block.

[0362] [Enhancement stratification and motion parameter inference] Main example: Inference of enhancement layer decomposition and motion parameters based on base layer decomposition and motion parameters (this example may require combining with one of the sub-examples).

[0363] Sub-examples: Obtain motion parameters for N×M subblocks of the enhancement layer based on co-located base layer motion data. Group blocks with the same obtained parameters (or parameters with small differences) into larger blocks. Determine the prediction and coding units. (See Example T) The motion parameters include the motion hypothesis, reference index list, motion vector, motion vector predictor identifier, and number of integrated identifiers. • To provide one of the multiplexing methods for generating enhancement layer prediction signals. Such methods include: ○ Motion compensation using the obtained motion parameters and the reference image of the reconstructed enhancement layer. ○(a) a (upsampled / filtered) base layer reconstruction for the current image, and (b) a motion compensation signal using the obtained motion parameters, combined with a reference image of the enhancement layer, which is generated by subtracting the (upsampled / filtered) base layer reconstruction from the reconstructed enhancement layer image. ○(a) combining the (upsampled / filtered) base layer residuals for the current image (the difference between the reconstructed signal and the prediction, or the inverse transform of the encoded transformation coefficient values), and (b) a motion compensation signal using the obtained motion parameters, with the reconstructed enhancement layer reference image. If the co-configured blocks in the base layer are intra-encoded, then the corresponding M×N blocks (or CUs) in the enhancement layer are also intra-encoded. There, the intra-prediction signal is obtained using the base layer information (see Example U). For example, ○The corresponding base layer reconstruction (upsampled / filtered) is used as the intra-prediction signal (see Example U). The intra-prediction mode is obtained based on the intra-prediction mode used in the base layer. This intra-prediction mode is then used for spatial intra-prediction in the enhancement layer. If a co-located base layer block for an M×N enhancement layer block (sub-block) is integrated into a previously encoded base layer block (or has the same motion parameters), then the M×N enhancement layer (sub) block is also integrated into the enhancement layer block corresponding to the base layer block used for integration within the base layer (i.e., the motion parameters are copied from the corresponding enhancement layer block) (see Example M).

[0364] [Coding / contextual modeling at the transformation coefficient level] Main examples: Encoding transformation coefficients using different scanning patterns; modeling the context based on the encoding mode and / or base layer data for the enhancement layer, and performing different initializations for the context model.

[0365] Sub-examples: • Introduce one or more additional scanning patterns, e.g., horizontal and vertical scanning patterns. Redefine subblocks for additional scanning patterns. Instead of 4x4 subblocks, e.g., 16x1 or 1x16 subblocks are used. Or, 8x2 or 8x2 subblocks are used. Additional scanning patterns are introduced only for blocks of a specific size, e.g., 8x8 or 16x16, that are greater than or equal to that size (see Example V). • (If the encoded block flag is equal to 1,) the selected scan pattern is shown in the bitstream (see Example N). A fixed context is used to indicate the corresponding syntactic element. Alternatively, context derivation for the corresponding syntactic element can rely on one of the following: ○ The gradient of the co-located reconstructed base layer signal or reconstructed base layer residual, or the edge detected in the base layer signal. ○ Distribution of transformation coefficients within jointly arranged base layer blocks. The selected scan is obtained directly from the base layer signal (without transmitting any additional data) based on the characteristics of the co-arranged base layer signal (see Example N). ○ The gradient of the co-located reconstructed base layer signal or reconstructed base layer residual, or the edge detected in the base layer signal. ○ Distribution of transformation coefficients within jointly arranged base layer blocks. • Different scans are implemented in a way that the conversion coefficients are reordered after quantization on the encoder side, and conventional encoding is used. On the decoder side, the conversion coefficients are decoded conventionally and reordered before (or after scaling and before inverse transformation). The following modifications are used within the enhancement layer to encode important flags (subgroup flags and / or important flags for a single transformation coefficient): ○A separate context model is used for all or a subset of encoding modes that use base layer information. It is also possible to use different context models for different modes that have base layer information. Contextual modeling can rely on data from co-located base layer blocks (e.g., the number of important transformation coefficients for a particular frequency position) (see Example O). ○A generalized template is used that evaluates both the number of already encoded significant transformation coefficient levels in the spatial neighborhood of the coefficient to be encoded, and the number of significant transformation coefficients in the co-located base layer signal at the same frequency location (see Example O). • The following modifications are used in the enhancement layer to encode the last critical scan location. ○The separated context model is used for all or a subset of encoding modes that use base layer information. It is also possible to use different context models for different modes that have base layer information (see Example P). Contextual modeling can rely on data within co-located base layer blocks (e.g., transformation coefficient distribution within the base layer, gradient information within the base layer, last scan position within co-located base layer blocks). ○The last scan position is encoded as the difference from the last base layer scan position (see Example S). • How to use different contextual initialization tables for the base layer and the enhancement layer.

[0366] [Coding of the backward adaptive enhancement layer using base layer data] Main example: Use of base layer data to obtain enhancement layer coding parameters.

[0367] Sub-examples: • Obtain a merge candidate based on the (potentially upsampled) base layer reconstruction. In the enhancement layer, only the use of merge is shown. However, in practice, the candidate used to merge the current block is obtained based on the reconstructed base layer signal. Thus, for all merge candidates, the magnitude of the error between the (potentially upsampled) base layer signal for the current enhancement layer block and the corresponding predicted signal (obtained using motion parameters on the merge candidate) is evaluated for all merge candidates (or a subset thereof). Then, the merge candidate relating to the minimum error magnitude is selected. The error magnitude is also calculated in the base layer using the reconstructed base layer signal and the base layer reference image (see Example Q). • Obtain a combined candidate based on the (potentially upsampled) base layer reconstruction. The motion vector difference is inferred based on the reconstructed base layer, although it is not encoded. Determine a motion vector predictor for the current block and evaluate a defined set of searches located around the motion vector predictor. For each search location, determine the magnitude of the error between the (potentially upsampled) base layer signal for the current enhancement layer block and the replaced reference frame (the replacement is given by the search location). Select the search location / motion vector that yields the minimum error magnitude. The search is divided into several stages. For example, a full Pell search is performed first. Then, a half Pell search is performed around the full Pell vector. Then, a quarter Pell search is performed around the best full / half Pell vector. The search is also performed within the base layer using the reconstructed base layer signal and the base layer reference image. The found motion vectors are then scaled according to the resolution change between the base layer and the enhancement layer (see Example Q). Obtain an intra-predictive mode based on the (potentially upsampled) base layer reconstruction. The intra-predictive mode is inferred based on the reconstructed base layer, although it is not encoded. For each possible intra-predictive mode (or a subset thereof), determine the magnitude of the error between the (potentially upsampled) base layer signal and the intra-predictive signal for the current enhancement layer block (using the tested prediction nodes). Select the prediction mode that yields the minimum error magnitude. The error magnitude calculation is performed in the base layer using the reconstructed base layer signal and the intra-predictive signal within the base layer. Furthermore, the intra-block can be decomposed into 4x4 blocks (or other block sizes). For each 4x4 block, a separate intra-predictive mode is determined (see Example Q). The intra-predicted signal is determined by the row or column direction of the boundary sample having the reconstructed base layer signal. To obtain the transition between the adjacent sample and the current line / column, the magnitude of the error is calculated between the transitioned line / column of the adjacent sample and the reconstructed base layer signal. The shift that yields the minimum magnitude of the error is then selected. As the adjacent sample, a sample from the (upsampled) base layer or a sample from the enhancement layer is used. Alternatively, the magnitude of the error can be calculated directly within the base layer (see Example W). • Use backward adaptive techniques to derive other coding parameters, such as block partitioning.

[0368] A more concise overview of the above-described embodiment is presented below. In particular, the above-described embodiment will be explained.

[0369] A1) Scalable video decoders are The base layer signals (200a, 200b, 200c) are reconstructed (80) from the encoded data stream (6). The enhancement layer signal (360) is reconstructed (60), Reconstruction (60) is, To obtain the interlayer prediction signal (380), the reconstructed base layer signals (200a, 200b, 200c) are subjected to resolution or quality improvement (220). The difference signal (260) between the already reconstructed portion of the enhancement layer signal (400a or 400b) and the interlayer prediction signal (380) is calculated. In order to obtain a spatial intra-predicted signal, a difference signal is spatially predicted (260) from a second portion (460) of the difference signal that is spatially adjacent to the first portion and belongs to an already reconstructed portion of the enhancement layer signal (360), in a first portion (440, as shown in Figure 46) that is jointly arranged with the portion of the enhancement layer signal (360) to be reconstructed. To obtain the enhancement layer prediction signal (420), the interlayer prediction signal (380) and the spatial intra-prediction signal are coupled (260). The configuration includes predictively reconstructing the enhancement layer signal (360) (320, 580, 340, 300, 280) using the enhancement layer prediction signal (420). According to Embodiment A1, the base layer signals are reconstructed by the base layer decoding stage 80 from the encoded data stream 6 or substream 6a, respectively, using the aforementioned block-based prediction method that includes, for example, the base layer residual signals 640 / 480, insofar as the base layer residual signals 640 / 480 are involved. However, other alternative reconstructions are also possible. With respect to the reconstruction of the enhancement layer signal 360 by the enhancement layer decoding stage 60, any improvement in resolution or quality that the reconstructed base layer signals 200a, 200b, or 200c undergo means, for example, upsampling in the case of resolution improvement, copying in the case of quality improvement, or tone mapping from n bits to m bits (m>n) in the case of bit depth improvement. The difference signal is calculated on a pixel-by-pixel basis. That is, pixels in which the enhancement layer signal and the prediction signal 380 are jointly arranged are subtracted from each other. This is done for each pixel position. Spatial prediction of the difference signal is performed in some way, such as by transmitting intra-prediction parameters, such as the intra-prediction direction, in the encoded data stream 6 or in the substream 6b, and then copying / interpolating already reconstructed pixels adjacent to the portion of the enhancement layer signal 360 that is currently to be reconstructed, along this intra-prediction direction in the current portion of the enhancement layer signal. Combinations mean addition, weighted sums, or more elaborate combinations, such as combinations that weight the contributions in the frequency domain differently. Predictive reconstruction of the enhancement layer signal 360 using the enhancement layer prediction signal 420 means the entropy decoding and inverse transform of the enhancement layer residual signal 540 and the combination of the enhancement layer prediction signal 420 and the latter 540, as shown in the figure.

[0370] B1) Scalable video decoder, The base layer residual signal (480) is decoded (100) from the encoded data stream (6). The enhancement layer signal (360) is reconstructed (60), Reconstruction (60) is, To obtain the interlayer residual prediction signal (380), the reconstructed baselayer residual signal (480) is subjected to resolution or quality improvement (220). To obtain an enhancement layer intra-predicted signal, the portion of the enhancement layer signal (360) that should be reconstructed is spatially predicted (260) from the already reconstructed portion of the enhancement layer signal (360). To obtain the enhancement layer prediction signal (420), the interlayer residual prediction signal and the enhancement layer intra prediction signal are coupled (260), The configuration includes predictively reconstructing (340) the enhancement layer signal (360) using the enhancement layer prediction signal (420). Decoding of the base layer residual signal from the encoded data stream is performed using entropy decoding and inverse transform, as shown in the figure. Furthermore, the scalable video decoder optionally performs reconstruction of the base layer signal itself by obtaining a base layer prediction signal 660 and predictively decoding it by concatenating this signal with the base layer residual signal 480. As just mentioned, this is simply optional. With respect to the reconstruction of the enhancement layer signal, improvements in resolution or quality are performed as indicated above with respect to Example A). Furthermore, with respect to the spatial prediction of the enhancement layer signal portion, this spatial prediction is performed as illustrated in A) for different signals. Similar considerations are valid with respect to combinations and predictive reconstructions. However, it is noted that the base layer residual signal 480 in Example B) is not limited to being equal to the obviously shown version of the base layer residual signal 480. Rather, it is possible for the scalable video decoder to subtract any reconstructed base layer signal version 200 having the base layer prediction signal 660. As a result, a base layer residual signal 480 is obtained that deviates from the obviously shown version due to the deviation arising from a filter function such as filter 120 or 140. Furthermore, the latter state is valid for another embodiment in which the base layer residual signal is involved in interlayer prediction.

[0371] C1) Scalable video decoder, The base layer signals (200a, 200b, 200c) are reconstructed (80) from the encoded data stream (6). The enhancement layer signal (360) is reconstructed (60), Reconstruction (60) is, To obtain the interlayer prediction signal (380), the reconstructed base layer signal (200) is subjected to resolution or quality improvement (220), To obtain an enhancement layer intra-predicted signal, the portion of the enhancement layer signal (360) that should be reconstructed is spatially or temporally predicted (260) from the already reconstructed portion of the enhancement layer signal (360) (400a,b in the case of "spatial" reconstruction; 400a,b,c in the case of "temporal" reconstruction). To obtain the enhancement layer prediction signal (420) such that the weighting of the interlayer prediction signal and the enhancement layer intra-prediction signal (380) contributing to the enhancement layer prediction signal (420) is var...

Claims

1. The base layer residual signal (480) of the base layer signal (200) is decoded (100) from the encoded data stream (6). The enhancement layer signal (360) is reconstructed (60), The reconstruction (60) of the enhancement layer signal is The following configuration is used to decode the conversion coefficient block (402) of the conversion coefficients representing the enhancement layer signal from the encoded data stream, Based on the base layer residual signal or the base layer signal, a sub-block subdivision is selected from a set of possible sub-block subdivisions. The conversion coefficient block traverses the positions of the conversion coefficients within the units of the subblock (412), which are regularly subdivided according to the subdivision of the selected subblock, such that all positions within one subblock are traversed in a continuous manner, immediately following the next subblock in the subblock order defined within the subblock. For the subblock currently visited, From the data stream, decode the syntactic element (416) indicating whether the currently visited subblock has an important transformation coefficient. If the syntactic element (416) indicates that the currently visited subblock does not have any important conversion coefficients, then set the conversion coefficients in the currently visited subblock to zero. If a syntactic element indicates that the currently visited subblock has an important transformation coefficient, the system is configured to include decoding a syntactic element (418) from the data stream that indicates the level of the transformation coefficient in the currently visited subblock. A scalable video decoder featuring the following characteristics.

2. A scalable video decoder according to claim 1, characterized in that the base layer signal (200) is decoded from the encoded data stream (6) by prediction, and the base layer residual signal is configured to represent the predicted residual of the prediction signal for the base layer signal.

3. A scalable video decoder according to claim 1 or 2, characterized in that it is configured to spatially, temporally, and / or interlayer predict the enhancement layer signal, and to reconstruct the enhancement layer signal (60) by applying an inverse transform block of transformation coefficients as a prediction residual to the prediction of the enhancement layer signal.

4. A scalable video decoder according to claim 1, characterized in that it is configured to select a sub-block subdivision from a set of possible sub-block subdivisions based on the base layer signal.

5. A scalable video decoder according to any one of claims 1 to 4, characterized in that it is configured to detect an edge in the base layer residual signal or a portion of the base layer signal corresponding to the conversion coefficient block, and to select a subblock subdivision from the set of possible subblock subdivisions by setting the subblock extension of the selected subblock subdivision to be elongated along a spatial frequency axis perpendicular to the edge.

6. A scalable video decoder according to any one of claims 1 to 4, characterized in that it uses the base layer residual signal or the spectral decomposition of a portion of the base layer signal corresponding to the conversion coefficient block, and is configured to select the sub-block subdivision from a set of possible subdivisions by setting the sub-block extension of the selected sub-block subdivision to be elongated along a spatial frequency axis perpendicular to the spatial frequency axis in which the spectral energy distribution of the spectral decomposition becomes narrower.

7. The scalable video decoder according to claim 6, characterized in that it is configured to form the spectral decomposition of the base layer residual signal or a portion of the base layer signal by actually applying a spatial domain to frequency domain conversion to the base layer residual signal or the base layer signal.

8. A scalable video decoder according to claim 6, characterized in that it is configured to combine and scale a conversion coefficient block of the base layer residual signal corresponding to the conversion coefficient block representing the enhancement layer signal, and to associate it with the portion that overlaps with the portion of the base layer residual signal, thereby forming the spectral decomposition of the base layer residual signal or the portion of the base layer signal.

9. The base layer residual signal (480) of the base layer signal (200) is decoded (100) from the encoded data stream (6). The enhancement layer signal (360) is reconstructed (60), The reconstruction (60) of the enhancement layer signal is The following configuration is used to decode the conversion coefficient block representing the enhancement layer signal from the encoded data stream, Based on the base layer residual signal or the base layer signal, a sub-block subdivision is selected from a set of possible sub-block subdivisions. The conversion coefficient block is traversed in such a continuous manner that all positions within one subblock immediately lead to the next subblock in the order of subblocks defined within the subblock, the conversion coefficient block is subdivided regularly according to the subdivision of the selected subblock, and the positions of the conversion coefficients within the subblock units are traversed. For the subblock currently visited, From the data stream, decode the syntactic element indicating whether the currently visited subblock has important transformation coefficients. If the syntactic element indicates that the currently visited subblock does not have any important conversion coefficients, then set the conversion coefficients in the currently visited subblock to zero. If a syntactic element indicates that the currently visited subblock has important transformation coefficients, then the process includes decoding a syntactic element from the data stream that indicates the level of transformation coefficients in the currently visited subblock. A scalable video decoding method characterized by the following.

10. The base layer residual signal (480) of the base layer signal (200) in the encoded data stream (6) is encoded, Encode the enhancement layer signal (360), The encoding of the enhancement layer signal is The following configuration encodes a block of conversion coefficients representing the enhancement layer signal from the encoded data stream: Based on the base layer residual signal or the base layer signal, a sub-block subdivision is selected from a set of possible sub-block subdivisions. The conversion coefficient block is traversed in such a continuous manner that all positions within one subblock immediately lead to the next subblock in the order of subblocks defined within the subblock, the conversion coefficient block is subdivided regularly according to the subdivision of the selected subblock, and the positions of the conversion coefficients within the subblock units are traversed. For the subblock currently visited, From the data stream, encode a syntactic element indicating whether the currently visited subblock has important transformation coefficients. If a syntactic element indicates that the currently visited subblock has a significant transformation coefficient, the system is configured to include encoding a syntactic element from the data stream that indicates the level of the transformation coefficient in the currently visited subblock. A scalable video encoder featuring the following characteristics.

11. The base layer residual signal (480) of the base layer signal (200) in the encoded data stream (6) is encoded, Encode the enhancement layer signal (360), The encoding of the enhancement layer signal is The following configuration encodes a block of conversion coefficients representing the enhancement layer signal from the encoded data stream: Based on the base layer residual signal or the base layer signal, a sub-block subdivision is selected from a set of possible sub-block subdivisions. The conversion coefficient block is traversed in such a continuous manner that all positions within one subblock immediately lead to the next subblock in the order of subblocks defined within the subblock, the conversion coefficient block is subdivided regularly according to the subdivision of the selected subblock, and the positions of the conversion coefficients within the subblock units are traversed. For the subblock currently visited, From the data stream, encode a syntactic element indicating whether the currently visited subblock has important transformation coefficients. If a syntactic element indicates that the currently visited subblock has important transformation coefficients, then the method includes encoding a syntactic element from the data stream that indicates the level of transformation coefficients in the currently visited subblock. A scalable video encoding method characterized by the following.

12. A computer program having the program code, wherein when the program code is executed on a computer, the computer executes the scalable video decoding method described in claim 9 or the scalable video encoding method described in claim 11.