Scalable video coding using the derivation of sub - partitioning of sub - blocks for prediction from a base layer

By analyzing the spatial variation of base layer coding parameters and weighting prediction signals based on spatial frequency components, the scalable video coding technique achieves improved coding efficiency and compression ratios.

JP7682959B2Active Publication Date: 2025-05-26DOLBY VIDEO COMPRESSION LLC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2023122307
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2012-10-01
Filing Date
2023-07-27
Publication Date
2025-05-26
Estimated Expiration
2033-10-01

AI Technical Summary

Technical Problem

Current scalable video coding techniques face challenges in achieving higher coding efficiency.

Method used

The proposed solution involves evaluating the spatial variation of base layer coding parameters to select the appropriate subdivision of sub-blocks for enhancement layer prediction, and weighting inter-layer and intra-enhancement layer prediction signals differently for various spatial frequency components to optimize prediction accuracy.

Benefits of technology

This approach enhances coding efficiency by improving the accuracy of enhancement layer prediction signals, leading to a higher compression ratio.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007682959000019
    Figure 0007682959000019
  • Figure 0007682959000020
    Figure 0007682959000020
  • Figure 0007682959000021
    Figure 0007682959000021
Patent Text Reader

Abstract

To provide a subblock division system in a scalable video coding.SOLUTION: In a set of sub divisions of a subblock, capable of blocking an enhancement layer by evaluating a spatial change of a base layer coding on a base layer signal, a derivation / selection of the sub division of the subblock to be used for the prediction of the enhancement layer is more effectively performed. Therefore, if so, less signal overhead has to be used for signaling to each sub division of this subblock in a stream of data of the enhancement layer. Each sub division of the subblock selected in this manner can be used when coding / decoding the signal of the enhancement layer predictively.SELECTED DRAWING: Figure 14
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to scalable video coding.

Background Art

[0002] In non-scalable coding, intra coding refers to a coding technique that uses only data of already-coded portions of the current image (e.g., reconstructed samples, coding modes, or symbol statistics), rather than reference data of already-coded images. For example, an intra-coded image (or intra image) is used within a broadcast bitstream at so-called random access points to synchronize a decoder to the bitstream. Also, intra images are used to limit error propagation in an error-prone environment. Generally, since images that can be used as reference images are not available here, the first image of an encoded video sequence must be encoded as an intra image. Usually, intra images are used at scene cuts where a prediction signal suitable for temporal prediction cannot normally be provided.

[0003] Furthermore, the intra coding mode is also used for specific regions / blocks within so-called inter images. There, they may perform better than the inter coding mode with respect to rate-distortion efficiency. This is often the case within flat regions, as well as in regions where temporal prediction is performed quite poorly (occlusions, objects that are partially dissolved or faded).

[0004] In scalable coding, the concept of intra coding (coding of intra pictures and intra blocks within inter pictures) is extended to all pictures belonging to the same access unit or time instance. Thus, the intra coding mode for spatial or quality enhancement layers can instantaneously increase coding efficiency while making use of inter-layer prediction from lower layer pictures. This means that not only can the already coded parts within the picture of the current enhancement layer be used for intra prediction, but also the lower layer pictures already coded at the same time instance. The latter concept is also referred to as inter-layer intra prediction.

[0005] In state-of-the-art hybrid video coding standards (such as H.264 / AVC or HEVC), the pictures of a video sequence are partitioned into sample blocks. The block size can be fixed, or the coding method can provide a hierarchical structure that allows the blocks to be further sub-divided into smaller block sizes. Usually, the reconstruction of a block is obtained by generating a prediction signal for the block and adding the transmitted residual signal. Usually, the residual signal is transmitted using transform coding, which means that a quantization index list for the transform coefficients (also referred to as transform coefficient levels) is transmitted using entropy coding techniques. And on the decoder side, these transmitted transform coefficient levels are scaled and inverse-transformed to obtain the residual signal to be added to the prediction signal. The residual signal is generated either by intra prediction (using only the data already transmitted for the current time instance) or by inter prediction (using data already transmitted for different time instances).

[0006] If inter prediction is used, the prediction block is obtained by motion compensated prediction using samples of a frame that has already been reconstructed. This can be done by uni-directional prediction (using one reference image and a set of motion parameters). Alternatively, the prediction signal can be generated by multi-hypothesis prediction. In the latter case, two or more prediction signals are superimposed. That is, for each sample, a weighted average is constructed to form the final prediction signal. The multiple prediction signals (that are superimposed) are generated using different motion parameters for different hypotheses (e.g., different reference images or motion vectors). Also, for uni-directional prediction, it is possible to multiply the samples of the motion compensated prediction signal by a constant factor and add a constant offset to form the final prediction signal. Also, such scaling and offset correction is used for all hypotheses or for selected hypotheses in multi-hypothesis prediction.

[0007] In current state-of-the-art video coding techniques, the intra prediction signal for a block is obtained by predicting samples from the spatial neighborhood of the current block, which in turn is the block reconstructed prior to the current block according to the blocks being processed in order. In the latest standards, various prediction techniques that perform prediction in the spatial domain are utilized. A refined granular directional prediction mode, with or without filtering the samples of adjacent blocks, is extended for specific angles to generate the prediction signal. Further, there are plane-based and DC-based prediction modes that use samples of adjacent blocks to generate a flat prediction plane or a DC prediction block.

[0008] In older video coding standards (e.g., H.263, MPEG-4), intra prediction was performed within the transform domain. In this case, the transmitted coefficients were inverse quantized. Then, for a subset of the transform coefficients, the transform coefficient values were predicted using the corresponding reconstructed transform coefficients of adjacent blocks. The inverse quantized transform coefficients were added to the predicted transform coefficient values, and the reconstructed transform coefficients were used as input to the inverse transform. The output of the inverse transform formed the final reconstructed signal for the block.

[0009] In scalable video coding as well, base layer information is utilized to support the prediction process for the enhancement layer. In the state-of-the-art video coding standard for scalable coding (SVC extension of H.264 / AVC), there is one additional mode for improving the coding efficiency of the intra prediction process within the enhancement layer. This mode is signaled at the macroblock level (a block of 16×16 luma samples). This mode is supported only when the collocated samples in the lower layer are coded using the intra prediction mode. If this mode is selected for a macroblock within the quality enhancement layer, the prediction signal is assembled by the collocated samples of the reconstructed lower layer signal before the non-blocking filter operation. If the inter-layer intra prediction mode is selected within the spatial enhancement layer, the prediction signal is generated by upsampling the collocated reconstructed base layer signal (after the non-blocking filter operation). An FIR filter is used for upsampling. Generally, for the inter-layer intra prediction mode, an additional residual signal is transmitted by transform coding. Also, if it is correspondingly signaled within the bitstream, the transmission of the residual signal can be omitted (inferred to be equal to zero). The final reconstructed signal is obtained by adding the reconstructed residual signal (obtained by scaling the transmitted transform coefficient levels and applying the inverse spatial transform) to the prediction signal. SUMMARY OF THE INVENTION

Problems to be Solved by the Invention

[0010] However, in scalable video coding, it is preferable to be able to achieve higher coding efficiency.

[0011] Therefore, an object of the present invention is to provide a concept for scalable video coding that achieves higher coding efficiency.

Means for Solving the Problems

[0012] This object is achieved by the subject matter of the independent claims filed simultaneously.

[0013] One embodiment of the present invention is that scalable video coding is more efficient by evaluating the spatial Variation of the base layer coding parameters on the base layer signal and deriving / selecting the sub-division of the sub-blocks to be used for enhancement layer prediction within a set of sub-divisions of possible sub-blocks of the enhancement layer block. For this purpose, if so, less signaling overhead has to be spent in the enhancement layer data stream to signal this sub-division of the sub-blocks. The sub-division of the sub-blocks thus selected can be used when predictively coding / decoding the enhancement layer signal.

[0014] One embodiment of the present invention is that within scalable video coding, a better predictor for predictively coding an enhancement layer signal forms an enhancement layer prediction signal from an inter-layer prediction signal and an intra-enhancement layer prediction signal by weighting them differently for different spatial frequency components to obtain the enhancement layer prediction signal, i.e., by forming a weighted average of the inter-layer prediction signal and the intra-enhancement layer prediction signal in the portion to be currently reconstructed. Thus, the weights by which the inter-layer prediction signal and the intra-enhancement layer prediction signal contribute to the enhancement layer prediction signal vary for different spatial frequency components. For this reason, it is possible to interpret the enhancement layer prediction signal from the inter-layer prediction signal and the intra-enhancement layer prediction signal in a way that is optimized with respect to the spectral characteristics of the individual contributing components, i.e., the inter-layer prediction signal on the one hand and the intra-enhancement layer prediction signal on the other hand. For example, due to the improved resolution or quality obtained from the reconstructed base layer signal, the inter-layer prediction signal may be more accurate at low frequencies compared to high frequencies. As for the intra-enhancement layer prediction signal, the characteristics are the opposite. That is, its accuracy can increase for high frequencies compared to low frequencies. In this example, at low frequencies, the contribution of the inter-layer prediction signal to the enhancement layer prediction signal, with their respective weights, exceeds the contribution of the intra-enhancement layer prediction signal to the enhancement layer prediction signal. And as for high frequencies, it does not exceed the contribution of the intra-enhancement layer prediction signal to the enhancement layer prediction signal. For this reason, a more accurate enhancement layer prediction signal can be achieved. As a result, the coding efficiency increases, resulting in a higher compression ratio.

[0015] Various embodiments are described for incorporating various possibilities for the concepts just outlined into any scalable video encoding based on the concepts. For example, the formation of the weighted average can be performed either in the spatial domain or in the transform domain. Performing spectral weighted averaging requires individual contributions, i.e., the transform to be performed on the inter-layer prediction signal and the intra-enhancement layer prediction signal. However, for example, avoid spectrally filtering either the inter-layer prediction signal or the intra-enhancement layer prediction signal in the spatial domain, including Processing . However, performing the formation of the spectral weighted average in the spatial domain avoids the detour of the individual contributions to the weighted average via the transform domain. The decision as to which domain is actually selected to perform the formation of the spectral weighted average may depend on whether the scalable video data stream contains the residual signal in the form of transform coefficients for the part that should currently be configured within the enhancement layer signal. If not, the detour via the transform domain is stopped. On the other hand, if the residual signal is present, the detour via the transform domain is more advantageous as it allows adding directly to the spectral weighted average in the transform domain for the transmitted residual signal in the transform domain. Processing Avoid spectrally filtering either the inter-layer prediction signal or the intra-enhancement layer prediction signal in the spatial domain, including Processing . However, performing the formation of the spectral weighted average in the spatial domain avoids the detour of the individual contributions to the weighted average via the transform domain. The decision as to which domain is actually selected to perform the formation of the spectral weighted average may depend on whether the scalable video data stream contains the residual signal in the form of transform coefficients for the part that should currently be configured within the enhancement layer signal. If not, the detour via the transform domain is stopped. On the other hand, if the residual signal is present, the detour via the transform domain is more advantageous as it allows adding directly to the spectral weighted average in the transform domain for the transmitted residual signal in the transform domain.

[0016] One embodiment of the present invention is that information available from base layer encoding / decoding, i.e., base layer hints, can be utilized to make motion compensation prediction in the enhancement layer more efficient by more efficiently encoding enhancement layer motion parameters. In particular, a set of motion parameter candidates collected from adjacent already reconstructed blocks of a frame of the enhancement layer signal is likely to be augmented by one or more sets of base layer motion parameters of blocks of the base layer signal (the base layer signal collocated with the blocks of the frame of the enhancement layer signal), and as a result, the available quality of the set of motion parameter candidates is improved based on the fact that motion compensation prediction of a block of the enhancement layer signal can be performed by selecting one of the motion parameter candidates of the augmented set of motion parameter candidates and using the selected motion parameter candidate for prediction. Additionally, or alternatively, the list of motion parameter candidates of the enhancement layer signal can be ordered depending on the base layer motion parameters involved in base layer encoding / decoding. For this reason, the probability distribution for selecting an enhancement layer motion parameter from the ordered list of motion parameter candidates can be compressed such that, for example, a clearly signaled index syntax element can be encoded using fewer bits (e.g., using entropy coding, etc.). Further, additionally, or alternatively, the index used within base layer encoding / decoding can be used as a basis for determining an index within the list of motion parameter candidates for the enhancement layer. For this reason, any signaling of an index for the enhancement layer can be completely avoided. Or, simply the prediction deviation thus determined for the index can be transmitted within the enhancement layer substream, and as a result, the encoding efficiency is improved.

[0017] One embodiment of the present invention is that, if the sub - division of the sub - blocks of each transformation coefficient block is controlled based on the base layer residual signal or the base layer signal, the encoding based on the sub - blocks of the enhancement layer transformation coefficient block can be made more efficient. In particular, by utilizing each base layer hint, the sub - blocks become longer along the horizontal spatial frequency axis with respect to the edge extension observable from the base layer residual signal or the base layer signal. For this reason, the shape of the sub - blocks is, with an increasing probability, such that each sub - block is filled with either almost entirely significant transformation coefficients (i.e., transformation coefficients not quantized to zero) or insignificant transformation coefficients (i.e., only transformation coefficients quantized to zero), while, with a decreasing probability, any sub - block can be adapted to the estimated distribution of the energy of the transformation coefficients of the enhancement layer transformation coefficient block such that the number of significant transformation coefficients on one side is the same as the number of insignificant transformation coefficients on the other side. However, due to the fact that sub - blocks having no significant transformation coefficients are efficiently signaled within the data stream, for example, by using only one flag, and due to the fact that sub - blocks almost entirely filled with significant transformation coefficients do not require a waste of signaling amount for encoding the insignificant transformation coefficients scattered therein, the encoding efficiency for encoding the enhancement layer transformation coefficient block increases.

[0018] One embodiment of the present invention is that the coding efficiency of scalable video coding can be increased by substituting the missing intra prediction parameter candidates in the spatial neighborhood of the current block in the enhancement layer with the intra prediction parameters of the collocated blocks of the base layer signal. Therefore, the coding efficiency for coding the intra prediction parameters is expected to increase due to the improved prediction quality of the set of intra prediction parameters in the enhancement layer, or more precisely, is likely to increase. A suitable predictor for the intra prediction parameters for the intra predicted blocks in the enhancement layer is useful, and as a result, increases the likelihood that the signaling of the intra prediction parameters for each enhancement layer block can be performed with fewer bits on average.

[0019] Further advantageous implementations are described in the dependent claims.

[0020] Preferred embodiments are described in detail below with reference to the drawings.

Brief Description of the Drawings

[0021]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 15c

Figure 15d

Figure 16

Figure 17

Figure 18

Figure 19a

Figure 19b

Figure 19c

Figure 20

Figure 21

Figure 22

Figure 23

Figure 24a

Figure 24b

Figure 25

Figure 26

Figure 27

Figure 28

Figure 29

Figure 30

Figure 31

Figure 32

Figure 33

Figure 34

Figure 35

Figure 36

Figure 37

Figure 38

Figure 39

Figure 40

Figure 41

Figure 42

Figure 43

Figure 44

Figure 45

Figure 46

DETAILED DESCRIPTION OF THE INVENTION

[0022] FIG. 1 shows, in a general way, an embodiment of a scalable video encoder incorporating the embodiments outlined below. The scalable video encoder of FIG. 1 is generally denoted using reference numeral 2 and receives and encodes video 4. The scalable video encoder 2 is configured to encode video 4 into a data stream 6 in a scalable manner. That is, the data stream 6 has a first portion 6a having video 4 encoded therein with a first information content amount, and another portion 6b having video 4 encoded therein with an information content amount greater than that of the first portion 6a. For example, the information content amounts of portions 6a and 6b may differ in terms of quality or fidelity, i.e., the amount of deviation per pixel from the original video 4 and / or in terms of spatial resolution. However, other forms of different information content amounts can also be applied, for example, in terms of color fidelity. Portion 6a is called a base layer data stream or a base layer sub-stream, while portion 6b may be called an enhancement layer data stream or an enhancement layer sub-stream.

[0023] The scalable video encoder 2 is configured to utilize redundancy between versions 8a and 8b of the reconstructible video 4 from the base layer substream 6a without the enhancement layer substream 6b on the one hand and from both substreams 6a and 6b on the other hand. To do so, the scalable video encoder 2 may use inter-layer prediction.

[0024] As shown in FIG. 1, the scalable video encoder 2 may selectively receive two versions 4a and 4b of the video 4. Both versions 4a and 4b have different information capacities from each other, just as the base layer substream 6a and the enhancement layer substream 6b do. Thus, for example, the scalable video encoder 2 is configured to generate substreams 6a and 6b such that the base layer substream 6a has the version 4a encoded therein. On the other hand, the enhancement layer data stream (substream) 6b has the version 4b encoded therein using inter-layer prediction based on the base layer substream 6b. The encoding of both substreams 6a and 6b may be irreversible.

[0025] Even if the scalable video encoder 2 only receives the original version of the video 4, the scalable video encoder 2 may be configured to internally derive two versions 4a and 4b therefrom by obtaining the base layer version 4a, for example, by spatial downscaling and / or tone mapping from a higher bit depth to a lower bit depth.

[0026] Figure 2 shows a scalable video decoder that matches the scalable video encoder 2 of FIG. 1 and is shown in a similar manner suitable for incorporating the embodiments outlined below. The scalable video decoder of FIG. 2 is generally designated by reference numeral 10. The scalable video decoder is generally configured to decode (decipher) the encoded data stream 6 so as to reconstruct the enhancement layer version 8b of the video therefrom, or, if, for example, part 6b is not available due to transmission loss or the like, to reconstruct the base layer version 8a of the video therefrom, provided that both parts 6a and 6b of the data stream 6 reach the scalable video decoder 10 in a complete manner. That is, the scalable video decoder 10 is configured to be able to reconstruct version 8a from only the base layer substream 6a and to be able to reconstruct version 8b using inter-layer prediction from both parts 6a and 6b.

[0027] Before explaining the embodiments of the present invention in more detail below, that is, before explaining the embodiments showing how the embodiments of FIGS. 1 and 2 are specifically implemented, a more detailed implementation of the scalable video encoder and decoder of FIGS. 1 and 2 will be described with respect to FIGS. 3 and 4. FIG. 3 shows a scalable video encoder 2 comprising a base layer encoder 12, an enhancement layer encoder 14 and a multiplexer 16. The base layer encoder 12 is configured to encode the base layer version 4a of the input video, while the enhancement layer encoder 14 is configured to encode the enhancement layer version 4b of the video. Accordingly, the multiplexer 16 receives the base layer substream 6a from the base layer encoder 12 and the enhancement layer substream 6b from the enhancement layer encoder 14 and multiplexes both of them into the encoded data stream 6 when outputting.

[0028] As shown in FIG. 3, both encoders 12 and 14 may be predictive encoders that use, for example, spatial prediction and / or temporal prediction to encode their respective input versions 4a and 4b into their respective sub-streams 6a and 6b. In particular, encoders 12 and 14 may each be hybrid video block encoders. That is, each of encoders 12 and 14 may be configured to encode each input version of the video on a block-by-block basis, for example, while different prediction modes are selected for each block into which the images or frames of video versions 4a and 4b are each sub-divided. The various prediction modes of base layer encoder 12 may include spatial and / or temporal prediction modes, while enhancement layer encoder 14 may additionally support an inter-layer prediction mode. The sub-division into blocks may be different between the base layer and the enhancement layer. The prediction mode, prediction parameters for the selected prediction mode for the various blocks, prediction residuals, and optionally, the block sub-division of each video version may be described by each encoder 12, 14 using a respective syntax that includes syntax elements that can be encoded in turn into their respective sub-streams 6a, 6b using entropy coding. Inter-layer prediction may be utilized, for example, one or more times to predict samples of the enhancement layer video, prediction mode, prediction parameters, and / or block sub-division as mentioned in the examples of 2, 3. Accordingly, both the base layer encoder 12 and the enhancement layer encoder 14 each include prediction encoders 18a, 18b followed by entropy encoders 19a, 19b. On the other hand, prediction encoders 18a, 18b form a syntax element stream from their respective inbound versions 4a and 4b using predictive coding. The entropy encoders entropy-encode the syntax elements output by their respective prediction encoders 18a, 18b. As just mentioned, the inter-layer prediction of encoder 2 may be relevant in the case of differences within the encoding procedure of the enhancement layer.Accordingly, the prediction encoder 18b is shown to be connected to one or more of the prediction encoder 18a, its output, and the entropy encoder 19a. Similarly, the entropy encoder 19b may optionally utilize inter-layer prediction, for example, by predicting the context used for entropy encoding from the base layer. Accordingly, the entropy encoder 19b is shown to be optionally connected to any of the elements of the base layer encoder 12.

[0029] In the same way as FIG. 2 with respect to FIG. 1, FIG. 4 shows a possible implementation of the scalable video decoder 10 that conforms to the scalable video encoder of FIG. 3. Accordingly, the scalable video decoder 10 of FIG. 4 includes a demultiplexer 40 that receives the data stream 6 to obtain the sub-streams 6a and 6b, a base layer decoder 80 configured to decode the base layer sub-stream 6a, and an enhancement layer decoder 60 configured to decode the enhancement layer sub-stream 6b. As shown, the decoder 60 is connected to the base layer decoder 80 to receive information therefrom for utilizing inter-layer prediction. Thereby, the base layer decoder 80 can reconstruct the base layer version 8a from the base layer sub-stream 6a. And the enhancement layer decoder 60 is configured to reconstruct the enhancement layer version 8b of the video using the enhancement layer sub-stream 6b. Similar to the scalable video encoder of FIG. 3, the base layer decoder 60 and the enhancement layer decoder 80 each internally include entropy decoders 100, 320, followed by prediction decoders 102, 322.

[0030] To simplify the understanding of the following embodiments, FIG. 5 illustratively shows various versions of video 4, namely base layer versions 4a and 8a that deviate from each other merely by encoding losses. Similarly, enhancement layer versions 4b and 8b merely deviate from each other by encoding losses respectively. The base layer signal and the enhancement layer signal may each be composed of a series of images 22a and 22b. They are shown in FIG. 5 so as to be registered with each other (i.e., in addition to the temporally corresponding image 22b of the enhancement layer signal, also the image 22a of the base layer version) along the time axis 24. As described above, image 22b may represent video 4 with a higher spatial resolution and / or with a higher fidelity, etc. (e.g., with a higher bit depth of the sample values of the image). Solid lines and dotted lines are used to show the encoding / decoding order as defined between images 22a, 22b. According to the example shown in FIG. 5, the encoding / decoding order is such that the base layer image 22a at a given time / moment crosses in front of the enhancement layer image 22b at the same time moment of the enhancement layer signal, crossing images 22a and 22b. With respect to the time axis 24, images 22a, 22b may be crossed by the encoding / decoding order 26 in the order of the provision time. However, an order deviating from the order of the provision time of images 22a, 22b is also possible. Neither encoder 2 nor decoder 10 needs to encode / decrypt continuously along the encoding / decoding order 26. Rather, encoding / decryption may be used in parallel. The encoding / decoding order 26 can define the availability between portions of the base layer signal and the enhancement layer signal adjacent to each other in a spatial, temporal and / or interlayer sense. As a result, when encoding / decoding the current portion of the enhancement layer, the available portion of the current enhancement layer portion is defined via the encoding / decoding order. Thus, since only adjacent portions available according to this encoding / decoding order 26 are used by the encoder for prediction, the decoder accesses the same information source to rectify the prediction.

[0031] For the following figures, how the scalable video encoder or decoder described above with respect to FIGS. 1-4 forms an embodiment of the present invention according to one embodiment of the present application will be described. Possible examples of the embodiments described below are discussed using the designation "Embodiment C".

[0032] In particular, FIG. 6 illustrates the image 22b of the enhancement layer signal shown here using reference numeral 360 and the image 22a of the base layer signal shown here using reference numeral 200. Temporally corresponding images of different layers are shown in the manner shown relative to each other with respect to the time axis 24. Using hatching, the portions within the base layer signal 200 and the enhancement layer signal 36 that have already been encoded / decoded in accordance with the encoding / decoding order are distinguished from the portions that have not yet been encoded or decoded in accordance with the encoding / decoding order shown in FIG. 5. Also, FIG. 6 shows a portion 28 of the enhancement layer signal 360 that is currently being encoded / decoded.

[0033] According to the presently described embodiment, the prediction of portion 28 uses both intra-layer prediction within the enhancement layer itself and inter-layer prediction from the base layer to predict portion 28. However, the predictions are combined such that they contribute to the final prediction of portion 28 in a spectrally varying manner. As a result, in particular, the ratio between both contributions varies spectrally.

[0034] In particular, the portion 28 is spatially or temporally predicted from the already reconstructed portion of the enhancement layer signal 400, i.e., the portion indicated by the hatching in the enhancement layer signal 400 in FIG. 6. The spatial prediction is explained using the arrow 30. On the other hand, the temporal prediction is explained using the arrow 32. The temporal prediction can include, for example, motion compensation prediction according to which motion vector information is transmitted within the enhancement layer substream for the current portion 28. The motion vector indicates the replacement of a portion of the reference picture of the enhancement layer signal 400 to be copied in order to obtain the temporal prediction of the current portion 28. The spatial prediction 30 can include the spatially adjacent portion to be estimated within the current portion 28, the already encoded / decoded portion of the picture 22b, and the spatially adjacent current portion 28. For this purpose, intra prediction information such as the estimation (or angular) direction can be signaled within the enhancement layer substream for the current portion 28. Also, a combination of the spatial prediction 30 and the temporal prediction 32 can be used as well. In any case, as a result, the intra prediction signal 34 of the enhancement layer is obtained as explained in FIG. 7.

[0035] To obtain another prediction of the current portion 28, inter-layer prediction is used. For this purpose, the base layer signal 200 is spatially and temporally corresponding to the current portion 28 of the enhancement layer signal 400 at the portion 36, and the inter-layer prediction signal for the current portion 28 undergoes a resolution or quality improvement in order to obtain an increasing potential resolution. The improvement procedure is explained using the arrow 38 in FIG. 6, resulting in the inter-layer prediction signal 39 as shown in FIG. 7.

[0036] Accordingly, two prediction contributions 34 and 39 exist for the current portion 28. And a weighted average of both contributions is formed for the current portion 28 in such a way that the weights by which the inter-layer prediction signal and the in-enhancement-layer prediction signal contribute to the enhancement-layer prediction signal 42 vary differently for spatial frequency components, as schematically shown at 44 in FIG. 7. FIG. 7 exemplarily shows a case where, for every spatial frequency component, the weights by which the prediction signals 34 and 38 contribute to the final prediction signal, for all spectral components, however, add the same value 46 in a state where the ratio between the weight applied to the prediction signal 34 and the weight applied to the prediction signal 39 varies spectrally.

[0037] On the other hand, the prediction signal 42 can be directly used for the current portion 28 by the enhancement-layer signal 400. Alternatively, the residual signal can be provided in the enhancement-layer sub-stream 6b of the current portion 28, which is brought about by a combination 50 with the prediction signal 42, for example, as shown by the addition in FIG. 7, within the reconstructed version 54 of the current portion 28. As an intermediate note, it should be noted that both the scalable video encoder and the decoder can be hybrid video decoders / encoders that use transform coding to encode / decode the prediction residuals and use predictive coding.

[0038] Summarizing the descriptions of FIGS. 6 and 7, the enhancement layer substream 6b can include an intra prediction parameter 56 for controlling spatial and / or temporal prediction 30, 32 for the current portion 28, and optionally a weighting parameter 58 for controlling the formation of a spectrum weighted average 41, and residual information 59 for signaling a residual signal 48. On the other hand, a scalable video encoder accordingly determines all of these parameters 56, 58, 59 and inserts the parameters 56, 58, 59 into the enhancement layer substream 6b. A scalable video decoder uses the parameters 56, 58, 59 to reconstruct the current portion 28 as outlined above. All of these elements 56, 58, 59 can undergo some quantization. And accordingly, the scalable video encoder can determine the use of these parameters / elements, i.e., quantization as a rate / distortion cost function. Interestingly, the encoder 2 uses the parameters / elements 56, 58, 59 thus determined to serve as a basis for any prediction for the portion of the enhancement layer signal 400 that follows, for example, in the order of encoding / decoding, so as to obtain a reconstructed version 54 for the current portion 28.

[0039] There are different possibilities for the weighting parameters 58 and the way they control the formation 41 of the spectral weighted average. For example, the weighting parameter 58 can signal only one of two states for the current portion 28, namely, one state that activates the formation of the spectral weighted average as previously described, and the other state that deactivates the contribution of the interlayer prediction signal 38. As a result, the final enhancement layer prediction signal 42 is then created only by the enhancement layer internal prediction signal 34. Instead, the weighting parameter 58 for the current portion 28 can switch between the activation of one spectral weighted average formation and the interlayer prediction signal 39 that forms the enhancement layer prediction signal 42 alone on the other hand. Also, the weighting parameter 58 can be designed to signal one of the three mentioned states / options. Alternatively, additionally the weighting parameter 58 can further signal, with respect to the current portion 28, the spectral variation of the ratio between the weights by which the prediction signals 34 and 39 contribute to the final prediction signal 42 (variation) with respect to which the spectral weighted average formation 41 can be controlled. It will be explained later that the spectral weighted average formation 41 can involve filtering one or both of the prediction signals 34 and 39 before adding them, for example, using a high-pass filter and / or a low-pass filter. In that case, the weighting parameter 58 can signal the filter characteristics of the filter to be used for the prediction of the current portion 28. As an alternative, the spectral weighting in step 41 can be achieved by the individual weighting of the spectral components in the transform domain, and thus, in this case, the weighting parameter 58 can signal / set the values of the individual weightings of these spectral components, as will be explained below.

[0040] Additionally, or alternatively, the weighting (weighting) parameter for the current portion 28 can signal whether the spectral weighting within step 41 is performed in the transform domain or the spatial domain.

[0041] FIG. 9 illustrates an embodiment for performing spectral weighted average composition within a spatial region. The prediction signals 39 and 34 are illustrated as being obtained in the form of respective pixel arrays that coincide with the pixel raster of the current portion 28. To perform the spectral weighted average composition, the pixel arrays of both prediction signals 34 and 39 are shown to receive a filter Processing . FIG. 9 exemplarily shows the filter Processing by showing filter kernels 62 and 64 that shift the pixel arrays of prediction signals 34 and 39 so as to perform, for example, a FIR filter Processing . However, an IIR filter Processing is also possible. Furthermore, only one of the prediction signals 34 and 39 may receive the filter Processing . Since the transfer functions of both filters 62 and 64 are different, the addition 66 of the results of the filtering Processing of the pixel arrays of prediction signals 39 and 34 results in the spectral weighted average composition result, i.e., the enhancement layer prediction signal 42. In other words, the addition 66 simply adds the juxtaposed samples in the filtered prediction signals 39 and 34 using filters 62 and 64, respectively. As a result, 62 to 66 result in the spectral weighted average composition 41. FIG. 9 illustrates that in the case of the residual information 59 existing in the form of transform coefficients, signaling to the residual signal 48 within the transform domain and the inverse transform 68 can be used to yield the spatial region in the form of a pixel array 70, and as a result, a combination 52 that yields the reconstructed version 55 can be realized by a simple pixel-wise addition of the residual signal array 70 and the enhancement layer prediction signal 42.

[0042] Again, it is to be recalled that the prediction is performed by a scalable video encoder and decoder using prediction for reconstruction within the decoder and encoder, respectively.

[0043] FIG. 10 illustratively shows how spectral weighted averaging is performed within the transform domain. Here, the pixel arrays of the prediction signals 39 and 34 each undergo transforms 72 and 74, respectively, resulting in spectral decompositions 76 and 78, respectively. Each spectral decomposition 76 and 78 has one transform coefficient per spectral component, and a transform coefficient array is created. Each transform coefficient block 76 and 78 is multiplied by a corresponding block of weights, namely, blocks 82 and 84. As a result, for each spectral component, the transform coefficients of blocks 76 and 78 are individually weighted. For each spectral component, the weighted values of blocks 82 and 84 add a value common to all spectral components. However, this is not obligatory. In fact, the multiplier 86 between blocks 76 and 82 and the multiplier 88 between blocks 78 and 84 each represent a spectral filter within the transform domain Processing The transform coefficient / spectral component unit addition 90 then concludes the spectral weighted averaging formation 41 to yield a transform domain version of the enhancement layer prediction signal 42 in the form of a block of transform coefficients. In the case of the residual signal 59 signaled to the residual signal 48 in the form of a transform coefficient block, as shown in FIG. 10, the residual signal 59 can be combined with the transform coefficient block representing the enhancement layer prediction signal 42, simply by a transform coefficient unit addition or another combination 52, to yield a reconstructed version of the current portion 28 within the transform domain. Thus, the inverse transform 84 applied to the additional result of the combination 52 yields the pixel array reconstructing the current portion 28, i.e., the reconstructed version 54.

[0044] As described above, the parameters present within the enhancement layer sub-stream 6b for the current portion 28, such as the residual information 59 or the weighting parameter 58, can signal whether the mean formation 41 is to be performed within the transform domain shown in FIG. 10 or within the spatial domain according to FIG. 9. For example, if the residual information 59 indicates the absence of any transform coefficient block for the current portion 28, the spatial domain is used. Alternatively, the weighting parameter 58 switches between both domains regardless of whether the residual information 59 includes transform coefficients or not.

[0045] Thereafter, it is explained that a difference signal can be calculated and managed between the already reconstructed portion of the enhancement layer signal and the inter-layer prediction signal in order to obtain an intra-layer enhancement layer prediction signal. The spatial prediction of the difference signal in the first portion located in the currently reconstructed portion of the enhancement layer signal can be used to spatially predict the difference signal from a second portion that is spatially adjacent to the first portion of the difference signal and belongs to the already reconstructed portion of the enhancement layer signal. Alternatively, the temporal prediction of the difference signal in the first portion located in the portion of the enhancement layer signal to be currently reconstructed can be used to obtain a temporally predicted difference signal from a second portion of the difference signal belonging to a previously reconstructed frame of the enhancement layer signal. The combination of the inter-layer prediction signal and the predicted difference signal may then be used to obtain an intra-layer enhancement layer prediction signal that is combined with the inter-layer prediction signal. Alternatively, the temporal prediction of the difference signal in the first portion juxtaposed to the portion of the enhancement layer signal to be currently reconstructed can be used to obtain a temporally predicted difference signal from a second portion of the difference signal belonging to a previously reconstructed frame of the enhancement layer signal. The combination of the inter-layer prediction signal and the predicted difference signal may then be used to obtain an intra-layer enhancement layer prediction signal that is combined with the inter-layer prediction signal.

[0046] For the following figures, it is described how a scalable video encoder or decoder as described above with respect to FIGS. 1-4 can be executed to form embodiments of the present application according to another form of the application.

[0047] To explain this embodiment, reference is made to FIG. 11. FIG. 11 shows the possibility of performing spatial prediction 30 of the current portion 28. As a result, the following description of FIG. 11 can be combined with the description regarding FIGS. 6-10. In particular, the embodiments described below are later described with respect to the illustrated example implementations by referring to Examples X and Y.

[0048] The situation shown in FIG. 11 corresponds to that shown in FIG. 6. That is, the base layer signal 200 and the enhancement layer signal 400 are shown. The already encoded / decoded portions are shown using hatching. Within the enhancement layer signal 400, the portions that are currently to be encoded / decoded have adjacent blocks 92 and 94. Here, by way of example, for both blocks 92 and 94 having the same size as the current block 28, block 92 is depicted above the current portion 28 and block 94 is depicted to the left. However, the size match is not obligatory. Rather, the portions of the blocks into which the image 22b of the enhancement layer signal 400 is sub-divided may have different sizes. They are not even restricted to being quadrilaterals. They may be rectangles or other shapes. Further, the current block 28 has adjacent blocks that are not explicitly represented in FIG. 11. However, the adjacent blocks have not yet been decoded / encoded. That is, the adjacent blocks follow in the order of encoding / decoding and as a result are not available for prediction. Beyond this, there may be blocks adjacent to the current block 28 that are different from the already encoded / decoded blocks 92 and 94 in accordance with the encoding / decoding order, such as block 96 that is diagonally adjacent, for example, at the upper left corner of the current block 28. However, blocks 92 and 94 are the predetermined adjacent blocks that serve to predict the intra prediction parameters for the current block 28 that is the subject of the intra prediction 30 in the example considered here. The number of such predetermined adjacent blocks is not limited to two. It may be more or even one.

[0049] A scalable video encoder and a scalable video decoder can determine a set of predetermined adjacent blocks, here blocks 92, 94, from a set of already encoded adjacent blocks. Here, blocks 92 to 96 depend on a predetermined sample position 98 within the current portion 28, such as its upper left sample. For example, only those already encoded adjacent blocks of the current portion 28 can form a set of "predetermined adjacent blocks" that includes sample positions immediately adjacent to the predetermined sample position 98. In any case, the adjacent already encoded / decoded blocks include samples 102 adjacent to the current block 28 based on the sample values for which the area of the current block 28 is to be spatially predicted. For this purpose, spatial prediction parameters such as 56 are signaled within the enhancement layer substream 6b. For example, the spatial prediction parameter for the current block 28 indicates the spatial direction in which the sample value of the sample 102 is to be copied into the area of the current block 28.

[0050] In any case, at least as far as the related spatially corresponding area of the temporally corresponding image 22a is concerned, as described above, block-by-block prediction is used, for example, using block-by-block selection between a spatial prediction mode and a temporal prediction mode, when spatially predicting the current block 28, the scalable video decoder / encoder uses the base layer substream 6a to already reconstruct (and in the case of the encoder, encode) the base layer 200.

[0051] In FIG. 11, some blocks 104 into which the images 22a arranged in time of the base layer signal 200 are sub-divided are in and around an area that locally corresponds to the currently illustrated portion 28. It is exactly the case of spatially predicted blocks within the enhancement layer signal 400, and the spatial prediction parameters are included or signaled within the base layer substream for the blocks 104 within the base layer signal 200 for which the selection of the spatial prediction mode is signaled.

[0052] Here, by way of example, for the coded data stream for block 28 where spatial intra-layer prediction 30 is selected, in order to enable reconstruction of the enhancement layer signal, the intra prediction parameters are used and coded within the following bitstream.

[0053] Intra prediction parameters are often coded using the concept of the most likely intra prediction parameters, which is a fairly small subset of all possible intra prediction parameters. The set of most likely intra prediction parameters includes, for example, one, two, or three intra prediction parameters. On the other hand, for example, the set of all possible intra prediction parameters can include 35 intra prediction parameters. If an intra prediction parameter is included in the set of most likely intra prediction parameters, it can be signaled in the bitstream with a small number of bits. If an intra prediction parameter is not included in the set of most likely intra prediction parameters, its signaling in the bitstream requires more bits. Thus, the amount of bits to be spent on syntax elements to signal the intra prediction parameters for the currently intra predicted block depends on the quality of the set of most likely, or perhaps advantageous, intra prediction parameters. Assuming that this concept can be used to appropriately derive the set of most likely intra prediction parameters, on average fewer bits are required to code the intra prediction parameters.

[0054] Typically, the set of most likely intra prediction parameters is selected in such a way that it includes the intra prediction parameters of directly adjacent blocks and / or additionally often uses intra prediction parameters, for example in the form of initial setting parameters. For example, since the main gradient directions of adjacent blocks are the same, it is generally advantageous to include the intra prediction parameters of adjacent blocks within the set of most likely intra prediction parameters.

[0055] However, if adjacent blocks are not coded in the spatial intra prediction mode, their parameters are not available at the decoder side.

[0056] In scalable coding, however, it is possible to use the intra prediction parameters of collocated base layer blocks. Thus, according to the embodiments outlined below, this situation is exploited for non-coded adjacent blocks within the spatial intra prediction mode by using the intra prediction parameters of collocated base layer blocks.

[0057] As a result, according to FIG. 11, a possibly advantageous set of intra prediction parameters for the current enhancement layer block is configured by examining the intra prediction parameters of the predefined adjacent blocks and, for example, by exceptionally re-partitioning the blocks collocated in the base layer in case the respective predefined adjacent blocks do not have the appropriate intra prediction parameters associated therewith since they are not coded in the intra prediction mode.

[0058] First, it is checked whether a predefined adjacent block, such as block 92 or 94 of the current block 28, is predicted using the spatial intra prediction mode. That is, it is checked whether the spatial intra prediction mode is selected for that adjacent block. Thereby, the intra prediction parameters of that adjacent block are included in a possibly advantageous set of intra prediction parameters for the current block 28 or, if any, alternatively, in the intra prediction parameters of the collocated base layer block 108. This process can be carried out for each of the predefined adjacent blocks 92 and 94.

[0059] For example, if each of the pre-determined adjacent blocks is not an in-space prediction block, instead of using initial setting prediction, etc., the intra prediction parameter of block 108 of the base layer signal 200 is included in a set of perhaps advantageous inter prediction parameters for the current block 28 juxtaposed to the current block 28. For example, the juxtaposed block 108 is determined using the pre-determined sample position 98 of the current block 28. That is, the block 108 covers the position 106 that locally corresponds to the pre-determined sample position 98 in the temporally arranged image 22a of the base layer signal 200. Naturally, a further check can be performed as to whether this juxtaposed block 108 in the base layer signal 200 is actually an in-space prediction block. In the case of FIG. 11, it is illustratively explained that this is the case. However, if the juxtaposed block is also not encoded in the intra prediction mode, a set of perhaps advantageous intra prediction parameters can be left with no contribution for its pre-determined adjacent blocks. Or, the initial setting intra prediction parameters can be used as an alternative instead. That is, the initial setting intra prediction parameters are inserted into a set of perhaps advantageous intra prediction parameters.

[0060] Therefore, if the block 108 juxtaposed to the current block 28 is an in-space prediction, the intra prediction parameter signaled in the base layer substream 6a is used for the pre-determined adjacent blocks 92 or 94 of the current block 28 that have no intra prediction parameters because the intra prediction parameters are encoded using another prediction mode such as the temporal prediction mode as an alternative.

[0061] According to another embodiment, in certain cases, if each of the predetermined adjacent blocks is in an intra prediction mode, the intra prediction parameters of the predetermined adjacent blocks are replaced by the intra prediction parameters of the collocated base layer blocks. For example, further checks such as whether the intra prediction parameters meet a predetermined criterion can be performed for any of the predetermined adjacent blocks in the intra prediction mode. If the predetermined criterion is not met by the intra prediction parameters of the adjacent blocks, but is met by the intra prediction parameters of the collocated base layer blocks, then the replacement is performed regardless of the very adjacent blocks that are intra-coded. For example, if the intra prediction parameters of the adjacent blocks do not represent an angular intra prediction mode (but, for example, a DC or planar intra prediction mode), but the intra prediction parameters of the collocated base layer blocks represent an angular intra prediction mode, then the intra prediction parameters of the adjacent blocks can be replaced by the intra prediction parameters of the base layer blocks.

[0062] The inter prediction parameters for the current block 28 are then determined based on the enhancement layer substream 6b for the current block 28 and syntax elements present in the coded data stream such as perhaps a set of advantageous intra prediction parameters. That is, the syntax elements can be coded using fewer bits in the case of the inter prediction parameters for the current block 28 which are perhaps members of a set of advantageous intra prediction parameters than in the case of the remaining members of the set of possible intra prediction parameters which perhaps do not lead to a set of advantageous intra prediction parameters.

[0063] The set of possible intra prediction parameters may include several angular mode directions that follow the fact that the current block is filled by copying from already encoded / decoded adjacent samples along the angular direction of each mode / parameter, a DC mode that follows the fact that the samples of the current block are set to a constant value determined based on, for example, several averages, such as already encoded / decoded adjacent samples, and a planar mode that follows the fact that the samples of the current block are set to a value distribution that follows a linear function of the slopes and intercepts of x and y, determined based on, for example, already encoded / decoded adjacent samples.

[0064] FIG. 12 shows the possibility of how an alternative of the spatial prediction parameters obtained from the collocated blocks 108 of the base layer can be used together with the syntax elements signaled within the enhancement layer substream. FIG. 12 shows an enlarged view of the current block 28 together with the adjacent already encoded / decoded samples 102 and the predetermined adjacent blocks 92 and 94. Also, FIG. 12 exemplarily shows the angular direction 112 indicated by the spatial prediction parameters of the collocated blocks 108.

[0065] The syntax element 114 signaled within the enhancement layer substream 6b for the current block 28 can signal, for example, as shown in FIG. 13, a conditionally encoded index 118 into a list 122, here illustratively shown as the angular direction 124, which is the result of possible advantageous intra prediction parameters. Or, if, hypothetically, the actual intra prediction parameter 116 is not within the most likely set 122 and is an index 123 within a list 125 of possible intra prediction modes that are potentially excluded as shown at 127, then the candidates in list 122, as a result, identify the actual intra prediction parameter 116. Encoding of the syntax element can consume fewer bits in the case of the actual intra prediction parameter belonging within list 122. For example, the syntax element can include a flag and an index field. The flag indicates whether to include or exclude members of list 122 and whether the index refers to either list 122 or list 125, i.e., whether to include or exclude in members of list 122. Or, the syntax element includes a field that identifies either a member 124 of list 122 or one of the escape codes. And, in the case of an escape code, the syntax element includes a second field that identifies a member from list 125 that includes or excludes members of list 122. The order within member 124 in list 122 can be determined, for example, based on default rules.

[0066] Accordingly, the scalable video decoder obtains, or recovers, the syntax element 114 from the enhancement layer sub-stream 6b. And the scalable video encoder can insert the syntax element 114 into the enhancement layer sub-stream 6b. And then, for example, the syntax element 114 is used to index one spatial prediction parameter from the list 122. When forming the list 122, the above-mentioned alternative can be executed by checking whether the pre-determined adjacent blocks 92 and 94 are of the spatial prediction coding mode type. Otherwise, as described above, it is checked whether the collocated block 108 is a spatially predicted block in turn, and if so, the spatial prediction parameter of the enhancement layer sub-stream, such as the angular direction 112, used to spatially predict this collocated block 108 is included in the list 122. Also, if the base layer block 108 does not contain a suitable intra prediction parameter, the list 122 can be left without contribution from the respective pre-determined adjacent blocks 92 or 94. Because, in order to avoid the list 122 being empty, for example, since both of the pre-determined adjacent blocks 92, 98 are, for example, inter predicted, at least one of the members 124 is unconditionally determined using the initial intra prediction parameter, similar to the collocated block 108 lacking a suitable intra prediction parameter. Alternatively, it may also be allowed that the list 122 is empty.

[0067] Of course, the embodiments described with respect to FIGS. 11 to 13 can be connected to the embodiments outlined above with respect to FIGS. 6 to 10. In particular, the intra prediction obtained using the spatial intra prediction parameter drawn bypassing the base layer according to FIGS. 11 to 13 can represent the enhancement layer internal prediction signal 34 of the embodiments of FIGS. 6 to 10 because it is combined with the inter-layer prediction signal 38 in a spectrally weighted manner as described above.

[0068] For the following drawings, as described with respect to FIGS. 1 - 4, how a scalable video encoder or decoder can form embodiments of the present application according to another embodiment of the application can be described. Later, for the embodiments described below, additional implementation examples are presented with reference to Examples T and U.

[0069] Referring to FIG. 14, images 22b and 22a of enhancement layer signal 400 and base layer signal 200 are shown in a time registration method respectively. Currently, the portion to be encoded / decoded is indicated by 28. According to the current embodiment, the base layer signal 200 is predictively encoded by a scalable video encoder using base layer encoding parameters that spatially vary () the base layer signal and is predictively reconstructed by a scalable video decoder. Spatial variation (variation) is shown in FIG. 14 using the hatched portion 132 where the base layer encoding parameters used to predictively encode / reconstruct the base layer signal 200 are constant, and is surrounded by a non-hatched region where the base layer encoding parameters change when transitioning from the hatched portion 132 to the non-hatched region. According to the embodiment outlined above, the enhancement layer signal 400 is encoded / reconstructed within a unit of blocks. The current portion 28 is such a block. According to the embodiment outlined above, the sub-block sub-division for the current portion 28 is within the juxtaposed portion 134 of the base layer signal 200, i.e., within the spatially juxtaposed portion of the temporally corresponding image 22a of the base layer signal 200, the spatial None variation (variation) is selected from one set of possible sub-block sub-divisions based on.

[0070] Specifically, instead of signaling within the sub - division information of the enhancement layer sub - stream 6b for the current portion 28, the above description presents selecting the sub - division of the sub - blocks within the set of possible sub - divisions of the current portion 28 such that the sub - division of the selected sub - blocks is the coarsest within the set of possible sub - divisions of the sub - blocks. Therein, when shifted onto the juxtaposed portion 134 of the base layer signal, the base layer coding parameters are sufficiently the same as each other within each sub - block of the sub - division of each sub - block. For ease of understanding, refer to FIG. 15a. FIG. 15a shows the portion 28 depicting the spatial variation of the base layer coding parameters within the juxtaposed portion 134 using hatching. Specifically, the portion 28 shows three instances of the sub - division of the various sub - blocks applied to block 28. In particular, the quadtree sub - division is illustratively used in the case of FIG. 15a. That is, the set of possible sub - divisions of the sub - blocks is the quadtree sub - division or is defined thereby. And the three specific examples of the sub - division of the sub - blocks of the portion 28 represented in FIG. 15a belong to different hierarchical levels of the quadtree sub - division of block 28. From bottom to top, the level or coarseness of the sub - division of block 28 within the sub - blocks increases. At the highest level, the portion 28 remains as it is. At the next lower level, block 28 is sub - divided into four sub - blocks. And at least one of the latter is further sub - divided into four sub - blocks at the next lower level and so on. In FIG. 15a, at each level, the quadtree sub - division is selected at the location where the number of sub - blocks is the smallest and which is not a sub - block overlapping with the base layer coding parameter change boundary. That is, in the case of FIG. 15a, it can be seen that the quadtree sub - division of block 28 to be selected for sub - dividing block 28 is the lowest one shown in FIG. 15a. Here, the base layer coding parameters of the base layer are constant within each portion juxtaposed to each sub - block of the sub - division of the sub - blocks. (variation)

[0071] ​ Therefore, the sub-division information for block 28 need not be signaled within the enhancement layer sub-stream 6b. As a result, the coding efficiency is increased. Moreover, as outlined, the method of obtaining the sub-division is appropriate regardless of the current position of portion 28 with respect to any grid (lattice), or any alignment of the sample array of the base layer signal 200. Also, in particular, the sub-division derivation works in the case of a fragmented spatial resolution ratio between the base layer and the enhancement layer.

[0072] Based on the sub-division of the sub-blocks of portion 28 thus determined, portion 28 is pre-predictively reconstructed / encoded. Regarding the above description, it should be noted that different possibilities exist to "measure" the coarseness of the sub-division of the different available sub-blocks of the current block 28. For example, the coarseness magnitude can be determined based on the number of sub-blocks. The more sub-blocks each sub-division of a sub-block has, the lower its level. This definition is clearly not applicable in the case of FIG. 15a where the "coarseness magnitude" is determined by the combination of the number of sub-blocks of each sub-division of a sub-block and the smallest size of all the sub-blocks of each sub-division of a sub-block.

[0073] To be complete, FIG. 15b illustratively shows the case of selecting one possible sub-division of a sub-block from one set of sub-divisions of the sub-blocks available for the current block 28 when illustratively using the sub-division of FIG. 35 as the available set. Different hatches (and non-hatches) indicate regions in the base layer signal that are juxtaposed to each other and have the same base layer coding parameters associated therewith.

[0074] As described above, the selection outlined traverses the sub - division of the possible sub - blocks according to a certain continuous order, such as in the order of increasing or decreasing levels of coarseness, and within each sub - block of the sub - division of each sub - block, in a situation where the base - layer coding parameters are sufficiently similar to each other, it can be carried out by selecting the sub - division of the possible sub - blocks from the sub - division of the possible sub - blocks. (When using traversal according to increasing levels of coarseness) It is no longer applicable. Or, (when using traversal according to decreasing levels of coarseness) it is applied incidentally at first. Optionally, all possible sub - divisions can be tested.

[0075] In the descriptions above FIGS. 14 and 15a, 15b, the broad term "base - layer coding parameters" is used in the preferred embodiments, but these base - layer coding parameters represent base - layer prediction parameters, that is, parameters related to the formation of the prediction of the base - layer signal and not related to the formation of the prediction residuals. Thus, for example, the base - layer coding parameters can include a prediction mode that distinguishes between spatial prediction and temporal prediction, prediction parameters for blocks / portions of the base - layer signal assigned to spatial prediction such as angular direction, prediction parameters for blocks / portions of the base - layer signal assigned to temporal prediction such as motion parameters, etc., and can be composed of such.

[0076] However, interestingly, within a given sub - block, the definition of "sufficient" similarity of the base - layer coding parameters only determines / defines a subset of the base - layer coding parameters. For example, the similarity can be determined based only on the prediction mode. Or also, prediction parameters that further adjust spatial prediction and / or temporal prediction can form parameters on which the similarity of the base - layer coding parameters within a given sub - block depends.

[0077] Furthermore, as already outlined above, in order to be sufficiently similar to each other, within a given sub-block, the base layer coding parameters may need to be exactly equal to each other within each sub-block. Alternatively, the degree of similarity used may need to be within a given range of intervals in order to meet the "similarity" criterion.

[0078] As outlined above, the sub-division of the selected sub-block is not only predicted from the base layer signal or the amount transferred. Rather, the base layer coding parameters themselves are transferred to the enhancement layer signal so as to obtain, based thereon, enhancement layer coding parameters for the sub-blocks of the sub-division of the sub-block obtained by transferring the sub-division of the selected sub-block from the base layer signal to the enhancement layer signal. As far as motion parameters are concerned, for example, scaling is used to account for the transfer from the base layer to the enhancement layer. Preferably, only those parts or syntax elements of the prediction parameters of the base layer are used to set the sub-blocks of the sub-division of the current part of the sub-block obtained from the base layer that affect the degree of similarity. By this degree, the fact that these syntax elements of the prediction parameters within each sub-block of the sub-division of the selected sub-block are somehow similar to each other ensures that the syntax elements of the base layer prediction parameters used to predict the corresponding prediction parameters of the sub-block of the current part 308 are similar or even equal to each other. As a result, in the first case allowing for some variation, some important "meanings" of the syntax elements of the base layer prediction parameters corresponding to the parts of the base layer signal covered by each sub-block can be used as predictors for the corresponding sub-blocks. However, also, only the parts of the syntax elements contributing to the degree of similarity are used to predict the prediction parameters of the sub-blocks of the sub-division of the enhancement layer by simply adding the transfer of the sub-division itself so as to merely infer or pre-set the mode of the sub-block of the current part 28.

[0079] One such possibility of not using only sub-partitioned inter-layer prediction from the base layer to the enhancement layer is now explained with respect to the following figure (Figure 16). Figure 16 shows an image 22b of an enhancement layer signal 400 and an image 22a of a base layer signal 200 in a registered manner along a presentation time axis 24.

[0080] According to the embodiment of FIG. 16, the base layer signal 200 is predictively reconstructed by a scalable video decoder by sub-dividing the frame 22a of the base layer signal 200 within intra-blocks and inter-blocks, and is predictively encoded by the use of a scalable video encoder. According to the example of FIG. 16, the latter sub-division is made by a two-stage method. First, the frame 22a is usually sub-divided along its periphery into the largest blocks or the largest coding units, indicated by reference numeral 302 in FIG. 16, using double lines. Then, each of the largest blocks 302 is subordinated to a hierarchical quadtree sub-division within the coding units forming the aforementioned intra-blocks and inter-blocks. As a result, they are the leaves of the quadtree sub-division of the largest block 302. In FIG. 16, reference numeral 304 is used to indicate these leaf blocks or coding units. Usually, solid lines are used to indicate the periphery of these coding units. On the other hand, spatial intra prediction is used for intra-blocks. Temporal inter prediction is used for inter-blocks. However, the prediction parameters related to the spatial intra prediction and the temporal inter prediction are set within the units of the smaller blocks, respectively, into which the intra- and inter-blocks or coding units 304 are sub-divided. Such a sub-division is exemplarily shown in FIG. 16 for one of the coding units 304 using reference numeral 306 to indicate the smaller blocks. The smaller blocks 304 are outlined using dotted lines. That is, in the case of the embodiment of FIG. 16, the spatial video encoder has the opportunity to select between one spatial prediction and the other temporal prediction for each coding unit 304 of the base layer. However, as far as the enhancement layer signal is concerned, the degree of freedom increases. Here, in particular, the frame 22b of the enhancement layer signal 400 is assigned to each one of a set of prediction modes including not only spatial intra prediction and temporal inter prediction but also inter-layer prediction as outlined in more detail below within the coding units into which the frame 22b of the enhancement layer signal 400 is sub-divided.The sub - division within these coding units can be done in the same way as described for the base - layer signal. First, frame 22b can be sub - divided into columns and rows of the largest contoured block using double lines for sub - division within the contoured coding units, using normal solid lines during the hierarchical quadtree sub - division process.

[0081] One coding unit 308 of the current image 22b of the enhancement - layer signal 400 is exemplarily inferred to be assigned to the inter - layer prediction mode and is shown using diagonal lines. In a similar way to FIGS. 14, 15a, and 15b, FIG. 16 shows at 312 how the sub - division of the coding unit 308 is predictively obtained by local transfer from the base - layer signal. In particular, the local region superimposed by the coding unit 308 is shown at 312. Within this region, the dotted lines indicate the boundaries between adjacent blocks of the base - layer signal, or more generally the boundaries where the base - layer coding parameters of the base - layer probably change. As a result, these boundaries are the boundaries of the prediction blocks 306 of the base - layer signal 200 and partially coincide with the boundaries between adjacent coding units 304 of the base - layer signal 200 or, equally, between the largest adjacent coding units 302. The dotted lines in 312 indicate the sub - division of the current coding unit 308 within the prediction block derived / selected by local transfer from the base - layer signal 200. Details regarding local transfer were described above.

[0082] According to the embodiment of FIG. 16, as already explained above, not only the sub - division within the prediction block but also from the base - layer is adopted. Rather, the prediction parameters of the base - layer signal used within the region 312 are used to obtain the prediction parameters to be used for performing prediction for the prediction block of the coding unit 308 of the enhancement - layer signal 400.

[0083] In particular, according to the embodiment of FIG. 16, the sub-division into the prediction block is not only obtained from the base layer signal, but also the prediction mode is used within the base layer signal 200 to encode / reconstruct each region locally covered by each sub-block of the obtained sub-division. One example is as follows. To obtain the sub-division of the encoding unit 308 according to the above, the prediction mode can be used in relation to the relevant base layer signal 200. Mode-specific prediction parameters can be used to determine the "similarity" discussed above. Thus, the different hatches shown in FIG. 16 can correspond to different prediction blocks 306 of the base layer. Each of the different prediction blocks 306 has an intra or inter prediction mode, i.e., a spatial or temporal prediction mode associated therewith. As explained above, in order to be "sufficiently similar", the prediction mode used within the region juxtaposed with each sub-block of the sub-division of the encoding unit 308 and the specific prediction parameters for each prediction mode within the sub-area may have to be exactly equal to each other. Alternatively, some variation (variation) might be tolerable.

[0084] In particular, according to the embodiment of FIG. 16, all the blocks indicated by the hatches extending from the upper left to the lower right are covered by the prediction blocks 306 in which the locally corresponding portions of the base layer signal have a spatial intra prediction mode associated therewith, so that they can be set in the intra prediction block of the encoding unit 308. On the other hand, the other blocks, i.e., the blocks indicated by the hatches extending from the lower left to the upper right, are covered by the prediction blocks 306 in which the locally corresponding portions of the base layer signal have a temporal inter prediction mode associated therewith, so that they can be set in the inter prediction block.

[0085] On the other hand, for an alternative embodiment, the derivation of the prediction can be stopped here within the encoding unit 308 where details for performing the prediction are. That is, the derivation of the sub-division of the encoding unit 308 into prediction blocks and the assignment of these prediction blocks into prediction blocks encoded using non-temporal prediction or spatial prediction and prediction blocks encoded using temporal prediction can be restricted, which does not follow the embodiment of FIG. 16.

[0086] According to the latter embodiment, all prediction blocks of the encoding unit 308 having a non-temporal prediction mode assigned thereto receive non-temporal prediction such as spatial prediction while using prediction parameters derived from the prediction parameters of the locally coincident intra-block of the base layer signal 200 as the enhancement layer prediction parameters of these non-temporal mode blocks. As a result, such a derivation is related to the spatial prediction parameters of the intra-blocks locally juxtaposed of the base layer signal 200. For example, such spatial prediction parameters may be an indication of the angular direction in which the spatial prediction is performed. As outlined above, due to the definition of the similarity itself, some averaging of the spatial base layer prediction parameters superimposed by each non-temporal prediction block of the encoding unit 308 is used to derive the prediction parameters of each non-temporal prediction block.

[0087] Alternatively, all prediction blocks of the encoding unit 308 having an assigned non-temporal prediction mode can receive inter-layer prediction in the following manner. First, the base layer signal undergoes resolution or quality improvement to obtain an inter-layer prediction signal in those regions spatially juxtaposed at least to the non-temporal prediction mode prediction blocks of the encoding unit 308. Then, next, these prediction blocks of the encoding unit 308 are predicted using the inter-layer prediction signal.

[0088] Scalable video decoders and encoders can, by default, cause all of the encoding units 308 to receive spatial prediction or inter-layer prediction. Alternatively, the scalable video encoder / decoder can support both alternatives and signal them within the encoded video data stream signal. Its version is used as long as it pertains to the non-temporal prediction mode prediction blocks of the encoding unit 308. In particular, the decision between both alternatives can be signaled within the data stream, for example, for any size of the encoding unit 308 individually.

[0089] As long as it pertains to another prediction block of the encoding unit 308, the encoding unit 308 can receive temporal inter prediction using prediction parameters that can be derived from the prediction parameters of the inter-block that are locally consistent as if it were the case of a non-temporal prediction mode prediction block. As a result, the derivation is related, in turn, to the motion vectors assigned to the corresponding parts of the base layer signal.

[0090] For all other coding units that have both the assigned spatial intra prediction mode and temporal inter prediction mode, another coding unit receives spatial prediction or temporal prediction in the following way. In particular, another coding unit is further sub-divided into prediction blocks that have the prediction mode assigned to it. The prediction mode is common to all of the prediction blocks within the coding unit and, in particular, is the same prediction mode assigned to each coding unit. That is, different from a coding unit such as the coding unit 308 and having an inter-layer prediction mode related to it, a coding unit having a related spatial intra prediction mode or a temporal inter prediction mode is sub-divided into prediction blocks of the same prediction mode. That is, the prediction mode only inherits from each coding unit derived by the sub-division of each coding unit.

[0091] The sub-division of all coding units including 308 can be a quadtree sub-division into prediction blocks.

[0092] A further difference between an inter-layer prediction mode coding unit such as the symbolization unit 308 and a coding unit in a spatial intra prediction mode or a temporal inter prediction mode is when the prediction blocks of the spatial intra prediction mode coding unit or the temporal inter prediction mode coding unit are respectively made to receive spatial prediction and temporal prediction. The prediction parameters are set, for example, by signaling within the enhancement layer substream 6b without depending on the base layer signal 200 or the like. Even the sub-division of those other coding units having an inter-layer prediction mode related thereto such as the coding unit 308 can be signaled within the enhancement layer signal 6b. That is, an inter-layer prediction mode coding unit such as 308 has the advantage of the necessity for signaling at a low bit transmission rate. According to an embodiment, the mode index of the coding unit 308 itself does not need to be signaled within the enhancement layer substream. Optionally, another parameter can be transmitted for the coding unit 308, such as a prediction parameter residual, for an individual prediction block. Additionally or alternatively, the prediction residual for the coding unit 308 can be transmitted / signaled within the enhancement layer substream 6b. On the other hand, a scalable video decoder searches for this information from the enhancement layer substream, and a scalable video encoder according to the current embodiment determines these parameters and inserts these parameters into the enhancement layer substream 6b.

[0093] In other words, the prediction of the base layer signal 200 is made using base layer coding parameters in such a way that the base layer coding parameters spatially vary the base layer signal 200 above within the unit of the base layer block 304. The prediction modes available for the base layer can include, for example, spatial and temporal prediction. The base layer coding parameters further include prediction mode individual prediction parameters such as the angular direction for the spatially predicted block 304 and motion vectors for the temporally predicted block 304. The individual prediction parameters of the latter prediction mode can vary the base layer signal within a unit smaller than the base layer block 304, i.e., within the aforementioned prediction block 306. In order to satisfy the requirements outlined before sufficient similarity, it may be necessary that the prediction modes of all base layer blocks 304 whose sub-division regions of each possible sub-block overlap are equal to each other. And only the sub-division of each sub-block can be put into the selection candidate list to obtain the sub-division of the selected sub-block. However, the requirements are even stricter. There may also be that the individual prediction parameters of the prediction mode of the prediction block, whose common regions of the sub-division of each sub-block overlap, must also be equal to each other. For each sub-block of this sub-division of each sub-block and the corresponding region within the base layer signal, only the sub-division of the sub-block that satisfies this requirement can be put into the selection candidate list to obtain the sub-division of the finally selected sub-block.

[0094] In particular, as briefly outlined above, there are various possibilities for how to perform the selection within a set of possible sub-block partitions. To outline this in more detail, reference is made to FIGS. 15c and 15d. Assume that set 352 surrounds sub-partitions 354 of all possible sub-blocks of current block 28. Of course, FIG. 15c shows merely an example. The set 352 of possible or available sub-partitions of current block 28 is known to scalable video decoders and scalable video encoders by default or can be signaled within an encoded data stream, such as a sequence of images or the like. According to the example of FIG. 15c, each member of set 352, i.e., each available sub-partition 354 of a sub-block, is subject to a check 356 to check whether the regions within the juxtaposed portions 108 of the base layer signal to be sub-divided, by transferring each sub-partition 354 of the sub-block from the enhancement layer to the base layer, are simply superimposed by prediction block 306 and encoding unit 304. Then, it is checked whether the base layer encoding parameters meet the requirement of sufficient similarity. For example, refer to the exemplary sub-division marked with reference number 354. According to this exemplary available sub-division of the sub-block, current block 28 is sub-divided into four quadrants / sub-blocks 358. And the upper left sub-block corresponds to region 362 within the base layer. Obviously, this region 362 does not correspond to another sub-division into four blocks of the base layer, i.e., into the prediction blocks, and as a result, it overlaps with two prediction blocks 306 and two encoding units 304 that represent the prediction blocks themselves. Thus, if all the base layer encoding parameters of these prediction blocks overlapping region 362 meet the similarity criterion, and this is also the case for all sub-blocks / quadrants of the possible sub-division 354 of the sub-block and their corresponding regions in the base layer with overlapping encoding parameters, then this possible sub-division 354 of the sub-block meets the sufficient requirements for all regions covered by the sub-blocks of the sub-division of the sub-block and belongs to the set 364 of sub-divisions of the sub-block.Then, within this set 364, the coarsest sub-division is selected as indicated by arrow 366, and as a result, from set 352, a sub-division 368 of the selected sub-block is obtained.

[0095] Obviously, it is preferable to avoid performing the check 356 for all members of set 352. Thus, as shown in FIG. 15d and as described above, the possible sub-divisions 354 can be traversed to increase or decrease the coarseness. The traversal is indicated using the double-headed arrow 372. FIG. 15d shows that for at least some of the sub-divisions of the available sub-blocks, the levels or magnitudes of coarseness are equal to each other. In other words, the ordering according to the increasing or decreasing levels of coarseness can be ambiguous. However, since only one of such equally coarse possible sub-divisions of the sub-block belongs to set 364, this does not prevent the search for the sub-division of the largest sub-block belonging to set 364. Thus, when traversing in the direction of the level where the coarseness increases or along the direction where the level of coarseness decreases in the sub-division of the possible sub-block traversed second to last, which is the sub-division 354 of the sub-block to be selected, when the result of the reference check 356 changes from being filled to not being filled or from not being filled to being filled, as soon as the most coarse possible sub-division 368 of the sub-block can be found.

[0096] For the following figures, a scalable video encoder or decoder as described above with respect to FIGS. 1 - 4 is implemented to form an embodiment of the present application according to another embodiment of the present application. Possible embodiments of the examples described below are presented below with reference to Examples K, A, and M.

[0097] To describe the embodiments, reference is made to FIG. 17. FIG. 17 shows the possibilities for the time prediction 32 of the current portion 28. As a result, the following description of FIG. 17 can be combined with the description regarding FIGS. 6 to 10 as long as it is related to the combination with the interlayer prediction signal. Alternatively, it is combined with the description regarding FIGS. 11 to 13 as long as it is related to the combination with the time interlayer prediction mode.

[0098] The situation shown in FIG. 17 corresponds to the situation shown in FIG. 6. That is, the base layer signal 200 and the enhancement layer signal 400 are shown together with the already encoded / decoded portions indicated using hatching. Within the enhancement layer signal 400, the portions that are currently to be encoded / decoded here are, by way of example, the adjacent blocks 92 and 94 described as the block 92 above and the 94 to the left of the current portion 28. Both blocks 92 and 94 have, by way of example, the same size as the current block 28. However, the size match is not essential. Rather, the portions of the blocks into the image 22b of the enhancement layer signal 400 that are sub-divided do not have different sizes. They are not even restricted to rectangles. They may be rectangular, or other shapes. The current block 28 having other adjacent blocks is not clearly described in FIG. 17. However, the other adjacent blocks have not yet been encoded / decoded. That is, they follow in the order of encoding / decoding and as a result are not available for prediction. In addition to this, there are blocks other than the blocks 92 and 94 that have already been encoded / decoded in accordance with the encoding / decoding order, by way of example, the block 96 that is diagonally above and to the left of the current block 28, adjacent to the current block 28. However, in the example considered here, the blocks 92 and 94 serve to predict the inter prediction parameters for the current block 28 that undergoes the inter prediction 30, and the adjacent blocks are predetermined. The number of such predetermined adjacent blocks is not limited to two. It may be greater than 1, or may simply be 1. The discussion of possible embodiments is presented with respect to FIGS. 36 to 38.

[0099] Scalable video encoders and scalable video decoders can determine a set of predetermined adjacent blocks, here the set of blocks 92, 94, from a set of already encoded adjacent blocks, here the set of blocks 92 to 96, that depend, for example, on a predetermined sample position 98 within the current portion 28 such as the sample in the upper left. For example, only those already encoded adjacent blocks of the current portion 28 that include sample positions directly adjacent to the predetermined sample position 98 can form a set of "predetermined adjacent blocks". Further possibilities are described with respect to FIGS. 36-38.

[0100] In any case, the previously encoded / decoded portion 502 of the enhancement layer signal 400, which is replaced from the collocated position of the current block 28 by the motion vector 504 according to the decoding / encoding order, includes sample values reconstructed based on sample values of the portion 28 that can be predicted by mere copying or interpolation. For this purpose, the motion vector 504 is signaled within the enhancement layer substream 6b. For example, the temporal prediction parameter for the current block 28 indicates a displacement vector 506 that shows the replacement of the portion 502 from the collocated position of the portion 28 within the reference image 22b by interpolation, optionally, for being copied onto the samples of the portion 28.

[0101] In any case, when temporally predicting the current block 28, the scalable video decoder / encoder has already reconstructed (and, in the case of the encoder, already encoded) the base layer 200 using the base layer substream 6a. At least as long as the relevant spatially corresponding region of the temporally corresponding image 22a is so related, block-by-block prediction is used as described above, and, for example, block-by-block selection between a spatial prediction mode and a temporal prediction mode is used.

[0102] In FIG. 17, several blocks 104 in which the temporally juxtaposed images 22a of the base layer signal 200 are sub-divided are in a region that locally corresponds to the current portion 28 and are illustratively drawn around it. Exactly that is the case when having spatially predicted blocks within the enhancement layer signal 400. The spatial prediction parameters are included in or signaled in the base layer sub-stream 6a for those blocks 104 within the base layer signal 200. The selection of the spatial prediction mode is signaled for the base layer signal 200.

[0103] Here, illustratively, for a block 28 for which the temporal intra-layer prediction 32 is selected, in order to enable the reconstruction of the enhancement layer signal from the encoded data stream, inter-prediction parameters such as motion parameters are determined using any of the following methods.

[0104] The first possibility is described with respect to FIG. 18. In particular, first, a set 512 of motion parameter candidates 514 is collected from, or generated from, adjacent already reconstructed blocks of the frame such as blocks 92 and 94 that are predetermined. The motion parameters are motion vectors. The motion vectors of blocks 92 and 94 are represented using arrows 516 and 518 in which 1 and 2 are respectively marked (therein). As shown in the illustration, these motion parameters 516 and 518 directly form the candidates 514. Some candidates are formed by combining motion vectors such as 518 and 516 as shown in FIG. 18.

[0105] Furthermore, a set 522 of one or more base layer motion parameters 524 of the blocks 108 of the base layer signal 200 juxtaposed to the portion 28 is collected from, or generated from, the base layer motion parameters. In other words, the motion parameters related to the blocks 108 juxtaposed within the base layer are used to obtain one or more base layer motion parameters 524.

[0106] At that time, one or more base layer motion parameters 524, or a scaled version thereof, are added 526 to the set 512 of motion parameter candidates 514 to obtain an extended motion parameter candidate set 528 of motion parameter candidates. This may be done in a variety of ways such as simply adding the base layer motion parameter 524 at the end of the list of candidates 514, or in a different way outlined with respect to FIG. 19a.

[0107] Next, at least one of the motion parameter candidates 532 of the extended motion parameter candidate set 528 is selected. A temporal prediction 32 is then performed using the selected one of the motion parameter candidates of the extended motion parameter candidate set by motion compensation prediction of portion 28. The selection 534 is signaled within a data stream such as sub-stream 6b for portion 28 by an index 536 within the list / set 528, or can be performed in another way described with respect to FIG. 19a.

[0108] As described above, it is checked whether the base layer motion parameter 523 has been encoded within an encoded data stream such as base layer sub-stream 6a using merge. And if, hypothetically, the base layer motion parameter 523 has been encoded within the encoded data stream using merge, the addition 526 is suppressed.

[0109] The motion parameters described according to FIG. 18 can be related to only the motion vectors (motion vector prediction), or to a complete set of motion parameters including the number of motion hypotheses for each block, reference index list, partitioning information (merge). Thus, the "scaled version" is derived from the scaling of the motion parameters used within the base layer signal according to the spatial resolution ratio between the base layer signal and the enhancement layer signal in the case of spatial scalability. Depending on the method of the encoded data stream, the encoding / decoding of the base layer motion parameters of the base layer signal can be involved in, for example, spatial or temporal motion vector prediction, or merging.

[0110] The incorporation 526 of motion parameters 523 used in the collocated portion 108 of the base layer signal into the set 528 of merge / motion vector candidates 532 enables a very effective indexing among the intra layer candidates 514 and one or more inter layer candidates 524. The selection 534 may involve explicit signaling of an index into an extended set / list of motion parameter candidates in the enhancement layer signal 6b, for each prediction block or similar, for each coding unit. Alternatively, the selected index 536 can also be inferred from other information in the enhancement layer signal 6b or from inter layer information.

[0111] According to the possibility of FIG. 19a, the formation 542 of the final motion parameter candidate list for the enhancement layer signal for part 28 is only optionally performed as outlined with respect to FIG. 18. That is, the formation 542 may also be 528 or 512. However, the list 528 / 512 is ordered 544 depending on base layer motion parameters such as the motion parameters represented by the motion vectors 523 of the collocated base layer blocks 108. For example, the rank of a member, i.e., a motion parameter candidate 532 or 514 of the list 528 / 512, is determined based on the deviation of each member with respect to a potentially scaled version of the motion parameter 523. The greater the deviation, the lower the rank of each member 532 / 512 in the ordered list 528 / 512'. As a result, the ordering 544 involves determining the magnitude of the deviation for each member 532 / 514 of the list 528 / 512. The selection 534 of one candidate 532 / 512 in the ordered list 528 / 512' is performed and controlled via an explicitly signaled index syntax element 536 in the coded data stream to obtain the enhancement layer motion parameters from the ordered motion parameter candidate list 528 / 512' for the enhancement layer signal part 28. Then, in turn, the temporal prediction 32 is performed using the selected motion parameters, where the index 536 points to 534, by motion compensation prediction of the enhancement layer signal part 28.

[0112] Regarding the motion parameters referred to in FIG. 19a, the motion parameters described above with respect to FIG. 18 are applied. Decoding of the base layer motion parameters 520 from the coded data stream can (optionally) involve spatial or temporal motion vector prediction, or merge. Ordering can be made according to the magnitude that measures the difference between each enhancement layer motion parameter candidate and the base layer motion parameter of the base layer signal, with respect to the block of the base layer signal juxtaposed to the current block of the enhancement layer signal. That is, for the current block of the enhancement layer signal, a list of enhancement layer motion parameter candidates can be determined first. Next, it is just described that the ordering is performed. Below, the selection is performed with explicit signaling.

[0113] Alternatively, the ordering 544 may be made according to a magnitude that increases the difference between the base layer motion parameter 523 of the base layer signal associated with the block 108 of the base layer signal juxtaposed to the current block 28 of the enhancement layer signal and the base layer motion parameter 546 of the spatially and / or temporally adjacent blocks 548 within the base layer. Next, the determined ordering within the base layer is transferred to the enhancement layer. As a result, the enhancement layer motion parameter candidates are ordered such that they have the same ordering as the determined ordering with respect to the corresponding base layer candidates. In this regard, when the associated base layer block 548 is spatially / temporally juxtaposed to the adjacent enhancement layer blocks 92 and 94 associated with the considered enhancement layer motion parameter, it can be said that the base layer motion parameter 546 corresponds to the enhancement layer motion parameters of the adjacent enhancement layer blocks 92, 94. Alternatively, when the adjacency relationship (left adjacent, upper adjacent, A1, A2, B1, B2, B0, or refer to FIGS. 36 to 38 for further examples) between the associated base layer block 548 and the block 108 juxtaposed to the current enhancement layer block 28 is the same as the adjacency relationship between the current enhancement layer block 28 and the enhancement layer adjacent blocks 92, 94 respectively, it may be said that the base layer motion parameter 546 corresponds to the enhancement layer motion parameters of the adjacent enhancement layer blocks 92, 94. Based on the base layer ordering, the selection 534 is then performed by explicit signaling.

[0114] To explain this in more detail, refer to FIG. 19b. FIG. 19b shows the first of the outlined alternatives for obtaining an enhancement layer that is ordered for a list of motion parameter candidates using base layer hints. FIG. 19b shows the current block 28 of the alternative and the positions of three different predetermined samples, namely, illustratively, the upper left sample 581, the lower left sample 583, and the upper right sample 585. The examples are to be construed as illustrative only. A set of predetermined adjacent blocks includes, illustratively, four types of adjacencies. The adjacent block 94a covers the sample position 587 that is adjacent directly above the sample position 581. The adjacent block 94b includes or covers the sample position 589 that is located adjacent directly above the sample position 585. Similarly, the adjacent blocks 92a and 92b include the immediately adjacent sample positions 591 and 593 that are located to the left of the sample positions 581 and 583. Also, note that the number of predetermined adjacent blocks can vary despite a predetermined number of decision rules, as will be explained with respect to FIGS. 36 - 38. Nevertheless, the predetermined adjacent blocks 92a, 92b, 94a, 94b are distinguishable by their decision rules.

[0115] According to an alternative to FIG. 19b, juxtaposed blocks within the base layer are determined for each of the predetermined adjacent blocks 92a, 92b, 94a, 94b. For example, for this purpose, the upper left samples 595 of each adjacent block are used. This is the case when having the current block 28 with respect to the upper left sample 581 formally mentioned in FIG. 19a. This is illustrated using the dotted arrows in FIG. 19b. Thereby, for each of the predetermined adjacent blocks, a corresponding block 597 is found in addition to the juxtaposed block 108 and the juxtaposed current block 28. Using the motion parameters m1, m2, m3, m4 of the juxtaposed base layer block 597 and their respective differences with respect to the base layer motion parameter m of the juxtaposed base layer block 108, the enhancement layer M1, M2, M3, M4 of the predetermined adjacent blocks 92a, 92b, 94a, 94b are ordered within list 528 or 512. For example, the larger the distance of any of m1 to m4, the higher the corresponding enhancement layer motion parameters M1 to M4. That is, a higher index list may be required to index from list 528 / 512' in the same state. The absolute difference can be used for the magnitude of the distance. Similarly, the motion parameter candidates 532 or 514 can be rearranged within the list with respect to their ranks which are a combination of the enhancement layer motion parameters M1 to M4.

[0116] FIG. 19c shows an alternative in which the corresponding blocks in the base layer are determined in a different way. In particular, FIG. 19c shows the predefined adjacent blocks 92a, 92b, 94a, 94b of the current block 28, and the juxtaposed block 108 of the current block 28. According to the embodiment of FIG. 19c, the corresponding base layer blocks of the current block 28, namely 92a, 92b, 94a, 94b, are determined in such a way that these base layer blocks are associated with the enhancement layer adjacent blocks 92a, 92b, 94a, 94b using the same adjacent determination rules for determining these base layer adjacent blocks. In particular, FIG. 19c shows the predefined sample positions of the juxtaposed block 108, namely the upper left, lower left, and upper right sample positions 601. Based on these sample positions, the four adjacent blocks of block 108 are determined in the same way as described for the enhancement layer adjacent blocks 92a, 92b, 94a, 94b with respect to the predefined sample positions 581, 583, 585 of the current block 28. The four base layer adjacent blocks 603a, 603b, 605a, 605b are found in this way. 603a clearly corresponds to the enhancement layer adjacent block 92a. The base layer block 603b corresponds to the enhancement layer adjacent block 92b. The base layer block 605a corresponds to the enhancement layer adjacent block 94a. The base layer block 605b corresponds to the enhancement layer adjacent block 94b. In the same way as previously described, the base layer motion parameters M1 to M4 of the base layer blocks 903a, 903b, 905a, 905b and their distances to the base layer motion parameter m of the juxtaposed base layer block 108 are used to order the motion parameter candidates within the list 528 / 512 formed from the motion parameters M1 to M4 of the enhancement layer blocks 92a, 92b, 94a, 94b.

[0117] In accordance with the possibilities of FIG. 20, the formation 562 of the final motion parameter candidate list for the enhancement layer signal for part 28 is optionally performed only, as outlined with respect to FIGS. 18 and / or 19. That is, the formation 562 is 528 or 512 or 528 / 512´. Reference numeral 564 is used in FIG. 20. In accordance with FIG. 20, the index 566 pointed out within the motion parameter candidate list 564 is determined depending on, for example, the index 567 into the motion parameter candidate list 568 used for encoding / decoding the base layer signal with respect to the juxtaposed block 108. For example, when reconstructing the base layer signal at block 108, the list 568 of motion parameter candidates is for a block 108 having the same adjacency relationship as the adjacency relationship between the predetermined adjacent enhancement layer blocks 92, 94 and the current block 28, and may be determined based on the motion parameters 548 of the adjacent block 548 of the block 108 having an adjacency relationship (adjacent to the left, adjacent to the top, A1, A2, B1, B2, B0, or for another example, see FIGS. 36 to 38). Here, the determination 572 of the list 567 potentially uses the same configuration rules as those used within the formation 562 such as the ordering within the list members of the lists 568 and 564. More generally, the index 566 for the enhancement layer is determined in such a way that its adjacent enhancement layer blocks 92, 94 are pointed out by the index 566 juxtaposed with the indexed base layer candidate, that is, the base layer block 548 related to what the index 567 points out. As a result, the index 567 can function as an important prediction of the index 566. The enhancement layer motion parameters are then determined using the index 566 into the motion parameter candidate list 564, and the motion compensation prediction for block 28 is performed using the determined motion parameters.

[0118] For the motion parameters mentioned in FIG. 20, the same applies as described above with respect to FIGS. 18 and 19.

[0119] Regarding the following figures, as described above for FIGS. 1-4, how a scalable video encoder or decoder can be implemented to form an embodiment of the present application according to another example of the application is described. The detailed implementation of the example described below is described with reference to Example V below.

[0120] This example relates to residual coding in the enhancement layer. In particular, FIG. 21 exemplarily shows the image 22b of the enhancement layer signal 400 and the image 22a of the base layer signal 200 in a temporally registered manner. FIG. 21 shows a method of reconstructing in a scalable video decoder or encoding in a scalable video encoder, and shows the enhancement layer signal, and focuses on a predetermined transform coefficient block of transform coefficients 402 representing the enhancement layer signal 400 and a predetermined portion 404. In other words, the transform coefficient block 402 represents the spatial decomposition of the portion 404 of the enhancement layer signal 400. According to the encoding / decoding ordering already described above, the corresponding portion 406 of the base layer signal 200 can already be decoded / encoded when decoding / encoding the transform coefficient block 402. As far as the base layer signal 200 is concerned, predictive encoding / decoding can be used, including the signaling of the base layer residual signal in an encoded data stream such as the base layer substream 6a.

[0121] According to the embodiment described with respect to FIG. 21, the scalable video decoder / encoder utilizes the fact that the evaluation 408 of the base layer signal or the base layer residual signal can result in an advantageous selection of the sub - division of the transform coefficient block 402 into sub - blocks 412 in the portion 406 collocated with the portion 404. In particular, several possible sub - divisions of the transform coefficient block 402 into sub - blocks are supported by the scalable video decoder / encoder. These possible sub - divisions of the transform coefficient block 402 can divide the transform coefficient block 402 into regularly rectangular sub - blocks 412. That is, the transform coefficients 414 of the transform coefficient block 402 can be arranged in rows and columns, and according to the possible sub - divisions of the sub - blocks, these transform coefficients 414 are densely arranged within the sub - blocks 412, so that the sub - blocks 412 themselves are arranged in rows and columns. The evaluation 408 uses the sub - division of the sub - blocks thus selected to enable the encoding of the transform coefficient block 402 in such a way that the ratio between the number of columns and the number of rows of the sub - blocks 412, that is, the ratio between their width and height, can be set to be the most efficient. If, for example, it can be determined that the reconstructed base layer signal 200 within the collocated portion 406, or at least the base layer residual signal within the corresponding portion 406, is mainly composed of horizontal edges in the spatial domain, then the transform coefficient block 402 will likely have significance, that is, the transform coefficient level is non - zero, that is, the quantized transform coefficients are near the zero - horizontal - frequency side of the transform coefficient block 402. In the case of vertical edges, the transform coefficient block 402 will likely have a non - zero transform coefficient level at a position near the zero - vertical - frequency side of the transform coefficient block 402. Therefore, first, the sub - blocks 412 should be selected to be longer along the vertical direction and smaller along the horizontal direction. And second, the sub - blocks should be made longer along the horizontal direction and smaller along the vertical direction. The latter case is schematically shown in FIG. 40.

[0122] That is, the scalable video decoder / encoder selects a sub-division of one sub-block within a set of possible sub-divisions of sub-blocks, based on the base layer residual signal or the base layer signal. At that time, the encoding 414 or decoding of the transform coefficient block 402 will be performed by applying the sub-division of the selected sub-block. In particular, since the positions of the transform coefficients 414 are traversed within the unit of the sub-block 412, all positions within one sub-block are traversed in such a way that they immediately follow in sequence to the next sub-block within the sub-block ordering defined within the sub-block. The syntax element, such as the currently visited sub-block 412 having the reference sign 112, which is exemplarily shown at 22 in FIG. 40, is signaled within a data stream such as the enhancement layer sub-stream 6b indicating whether the currently visited sub-block has significant transform coefficients. In FIG. 21, the syntax element 416 is described for two exemplary sub-blocks. If, for each sub-block, each syntax element indicates non-significant transform coefficients, then nothing else needs to be transmitted within the data stream or the enhancement layer sub-stream 6b. Rather, the scalable video decoder sets the transform coefficients within that sub-block to zero. However, if, for each sub-block, the syntax element 416 indicates that this sub-block has significant transform coefficients, then other information related to the transform coefficients within that sub-block is signaled within the data stream or the sub-stream 6b. On the decoding side, the scalable video decoder decodes the syntax element 418 indicating the levels of the transform coefficients within each sub-block from the data stream or the sub-stream 6b. The syntax element 418 indicates the scan order within these transform coefficients within each sub-block and, optionally, the positions of the significant transform coefficients within that sub-block according to the scan order within the transform coefficients within each sub-block.

[0123] FIG. 22 shows various possibilities that each exist to perform a selection within a sub - division of possible sub - blocks within the evaluation 408. FIG. 22 shows again the portion 404 of the enhancement layer signal to which the transform coefficient block 402 is related, where the latter represents the spectral decomposition of the portion 404. For example, the transform coefficient block 402 represents the spectral decomposition of the enhancement layer residual signal with a scalable video decoder / encoder that predictively encodes / decodes the enhancement layer signal. In particular, the encoding / decoding of the transform is used by the scalable video decoder / encoder to encode the enhancement layer residual signal. The encoding / decoding of the transform is performed in a block - by - block manner, i.e., within the blocks into which the image 22b of the enhancement layer signal is sub - divided. FIG. 22 shows the corresponding or juxtaposed portion 406 of the base layer signal. Here, the scalable video decoder / encoder applies predictive encoding / decoding to the base layer signal while using the encoding / decoding of the transform with respect to the prediction residual of the base layer signal, i.e., with respect to the base layer residual signal. In particular, the block - by - block transform is used for the base layer residual signal. That is, the base layer residual signal is transformed block - by - block in the individually transformed blocks shown by the dotted lines in FIG. 22. As shown in FIG. 22, the block boundaries of the base layer transform blocks do not need to coincide with the outer shape of the juxtaposed portion 406.

[0124] Nevertheless, to perform the evaluation 408, one or a combination of the following options A - C can be used.

[0125] In particular, a scalable video decoder / encoder can perform a transform 422 on a base layer residual signal or a reconstructed base layer signal within portion 406 in order to obtain a transform coefficient block 424 of transform coefficients that matches in size the transform coefficient block 402 to be encoded / decoded. Examination of the distribution of the values of the transform coefficients within transform coefficient blocks 424, 426 can be used to appropriately set the dimensions of sub-blocks 412 along the direction 428 of horizontal frequency, i.e., and along the direction of vertical frequency, i.e., 432.

[0126] Additionally, or alternatively, a scalable video decoder / encoder can examine all transform coefficient blocks of a base layer transform block 434 indicated by different hatching in FIG. 22 that at least partially overlap juxtaposed portions 406. In the exemplary case of FIG. 22, there are four base layer transform blocks, and then their transform coefficient blocks are examined. In particular, all of these base layer transform blocks are of different sizes from each other and further different in size with respect to transform coefficient block 412. Scaling 436 can be performed on these transform coefficient blocks that overlap the base layer transform block 434 in order to result in an approximation of a transform coefficient block 438 of the spectral decomposition of the base layer residual signal within portion 406. The distribution of the values of the transform coefficients within that transform coefficient block 438, i.e., 442, can be used within evaluation 408 to appropriately set the dimensions 428 and 432 of the sub-blocks. As a result, a sub-division of the sub-blocks of transform coefficient block 402 is selected.

[0127] A further alternative means that may additionally or alternatively be used to perform evaluation 408 is to examine the base layer residual signal within the spatial domain using edge detection 444 or determination of the main gradient direction, and to determine, for example, based on the extension direction of the detected edges, or the gradient determined within juxtaposed portions 406, so as to appropriately set the dimensions 428 and 432 of the sub-blocks.

[0128] Although not explicitly described above, in traversing the position of the transform coefficient and the unit of sub-block 412, it is preferable to traverse sub-block 412 in order starting from the zero-frequency corner of the transform coefficient block, i.e., the upper left corner of FIG. 21, and reaching the highest-frequency corner of block 402, i.e., the lower right corner of FIG. 21. Further, entropy coding can be used to signal syntax elements within data stream 6b. That is, syntax elements 416 and 418 can be entropy codings that are convenient for encoding, such as arithmetic coding, variable-length coding, or another form of entropy coding. The ordering for traversing sub-block 412 can depend on the sub-block shape selected according to 408. For sub-blocks selected to be wider than their height, the traversal order can traverse the sub-block in row units first and then proceed to the next row, etc. Beyond this, it should be noted again that the base layer information used to select the dimensions of the sub-block can be the base layer residual signal or the base layer signal itself reconstructed by itself.

[0129] In the following, various embodiments that can be combined with the embodiments described above are described. The embodiments described below relate to many different examples or means for making scalable video coding more efficient. In part, the above embodiments are described in more detail below while maintaining the general concept to present another derived embodiment thereof. These descriptions presented below can be used to obtain alternatives or extensions of the above embodiments / examples. However, most of the embodiments described below are optionally combinable with the embodiments already described above. That is, they can be implemented together with the above embodiments within a single scalable video decoder / encoder, but relate to sub-aspects that are not necessarily required.

[0130] To more easily understand the foregoing description, more detailed embodiments for implementing a suitable scalable video encoder / decoder incorporating any of the embodiments or combinations thereof are presented next. The various examples described below are enumerated by the use of alphanumeric symbols. Some descriptions of these aspects refer to elements in the figures being described now, according to one embodiment, in which these aspects can be commonly implemented. However, it should be noted that, as far as individual examples are concerned, the presence of all elements within the implementation of the scalable video decoder / encoder is not necessary as far as all examples are concerned. Depending on the problem situation, some elements and some interconnections can be omitted in the drawings described next. Only the elements referred to for each example are present to perform the work or function mentioned in the description of each example. However, in particular, when several elements are listed with respect to one function, there may be alternatives.

[0131] However, to provide an overview of the functions of the scalable video decoder / encoder, the examples described next can be implemented. The elements shown in the following figures are now briefly described.

[0132] FIG. 23 shows a scalable video decoder for decoding an encoded data stream 6 of a video in such a way that a suitable sub - portion of the encoded data stream 6, namely 6a, represents the video at a first resolution or quality level. An additional portion 6b of the encoded data stream corresponds to the representation of the video at an increasing resolution or quality level. To keep the data volume of the encoded data stream 6 low, the inter - layer redundancy between sub - streams 6a and 6b is utilized when forming sub - stream 6b. Some of the examples described below are such that sub - stream 6a is directed towards inter - layer prediction from the associated base layer, and sub - stream 6b is directed towards the associated enhancement layer.

[0133] The scalable video decoder includes two block-based prediction decoders 80, 60 operating in parallel, and receives sub-streams 6a and 6b respectively. As shown in the figure, the demultiplexer 40 can provide the decoding stages 80 and 60 separately, together with the corresponding sub-streams 6a and 6b.

[0134] The internal structures of the block-based prediction encoding stages 80 and 60 can be the same as shown in the figure. From each input of the decoding stages 80, 60, the entropy encoding modules 100; 320, the inverse transformers 560; 580, the adders 180; 340, any filters 120; 300 and 140; 280 are connected in series in this order of description. As a result, at the end of this series connection, the reconstructed base layer signal 600 and the reconstructed enhancement layer signal 360 can be obtained respectively. On the other hand, the outputs of the adders 180, 340 and the filters 120, 140, 300, 280 provide various versions of the reconstruction of the base layer signal and the enhancement layer signal respectively. Then, each prediction provider 160; 260 receives a subset or all of these versions and provides a prediction signal to the residual input of the adder 180; 340 based on them. The entropy decoding stages 100; 320 decode from the respective input signals 6a and 6b, and the transform coefficient blocks enter the inverse transformers 560; 580 and encode the parameters including the prediction parameters for the prediction providers 160; 260 respectively.

[0135] Therefore, the prediction providers 160 and 260 predict the blocks of the video frame at their respective resolution / quality levels. And for this purpose, the prediction providers 160 and 260 can be selected within a predetermined prediction mode such as a spatial intra prediction mode or a temporal inter prediction mode. Both modes are intra-layer prediction modes, that is, prediction modes that only depend on the data within the sub-stream containing each level.

[0136] However, in order to utilize the aforementioned interlayer redundancy, the enhancement layer decoding stage 60 additionally includes a coding parameter interlayer predictor 240, a resolution / quality improver 220, and / or a predictor 260 to be compared with the predictor 160. Further, or also, the enhancement layer decoding stage 60 assists an interlayer prediction mode that can provide an enhancement layer prediction signal 420 based on data obtained from the internal state of the base layer decoding stage 80. The resolution / quality improver 220 subjects either the reconstructed base layer signals 200a, 200b, 200c or the base layer residual signal 480 to resolution or quality improvement in order to obtain an interlayer prediction signal 380. The coding parameter interlayer predictor 240 is to predict coding parameters such as prediction parameters and motion parameters in some form. The predictor 260 assists the interlayer prediction mode, for example, further according to the reconstructed portions of the base layer signals such as 200a, 200b, 200c. Alternatively, the reconstructed portions of the potentially improved base layer residual signal 640 at increasing resolution / quality levels are used as a reference / basis.

[0137] As described above, the decoding stages 60 and 80 can operate in a block-based manner. That is, a video frame can be sub-divided into parts such as blocks. The various roughness levels can be used to assign prediction modes such as those performed by the prediction providers 160, 260, local transforms by the inverse transformers 560, 580, filter coefficient selection by the filters 120, 140, and prediction parameter settings for the prediction modes by the prediction providers 160, 260. That is, sub-dividing a frame into prediction blocks can, in turn, be continued with sub-division of the frame into blocks for which a prediction mode is selected (e.g., so-called coding units or prediction units). The sub-division of a frame into blocks for transform coding (so-called transform units) can be different from the partitioning into prediction units. Some of the inter-layer prediction modes used by the prediction provider 260 are described below with respect to the examples. The prediction provider 260 is also applied with respect to some intra-layer prediction modes, that is, prediction modes that internally derive the respective prediction signals input to the respective adders 180, 340, that is, prediction modes that are based solely on the state associated with the encoding stages 60, 80 of the current level, respectively.

[0138] Some other details of the illustrated blocks will become apparent from the description of the individual examples below. It should be noted that these descriptions are equally generally transferable to the description of other examples and figures, unless such descriptions are explicitly related to the examples provided.

[0139] In particular, the embodiment for the scalable video decoder of FIG. 23 represents a possible implementation of the scalable video decoder according to FIGS. 2 and 4. Although the scalable video decoder according to FIG. 23 has been described above, FIG. 23 shows the corresponding scalable video encoder, and the same reference numerals are used for the internal elements of the predictive coding / decoding method in FIGS. 23 and 24. The reason is as described above. Also, for the purpose of maintaining a common prediction basis between the encoder and the decoder, the reconstructable versions of the base and enhancement layer signals are used in the encoder and, in order to obtain a reconstructable version of the scalable video, the already encoded portions up to this point are reconstructed. Thus, the only difference from the description of FIG. 23 is that, similar to the inter-layer predictor 240 for the encoding parameters, the prediction provider 160 and the prediction provider 260 determine the prediction parameters within some ratio / distortion optimization process rather than receiving the prediction parameters from the data stream. Rather, the providers transmit the prediction parameters thus determined to the entropy decoders 19a and 19b. The entropy decoders 19a and 19b transmit the respective base layer sub-stream 6a and enhancement layer sub-stream 6b in turn via the multiplexer 16 for inclusion in the data stream 6. In a similar manner, rather than these entropy encoders 19a and 19b outputting such entropy decoding results of the residuals, the reconstructed base layer signal 200 and the reconstructed enhancement layer signal 400 and the prediction residuals between the original base layer and enhancement layer versions 4a, 4b are received such that they are obtained via the subsequent subtractors 720 and 722 by the conversion modules 724, 726. However, otherwise, the structure of the scalable video encoder of FIG. 24 is consistent with the structure of the scalable video decoder of FIG. 23. Thus, with regard to these issues, the above description of FIG. 23 is referenced. Here, as just outlined, any derivation from any data stream must be changed to the respective determination of the respective elements having subsequent insertion into the respective data stream.

[0140] The techniques for intra-coding of enhancement layer signals used in the embodiments described below include a plurality of ways to generate an intra-prediction signal (using base layer data) for an enhancement layer block. These ways are provided in addition to ways to generate an intra-prediction signal based only on samples of the reconstructed enhancement layer.

[0141] Intra-prediction is part of the process of reconstructing an intra-coded block. The final reconstructed block is obtained by adding the residual signal (which may be zero) coded by transform to the intra-prediction signal. The residual signal is generated by inverse quantization (scaling) of the transform coefficient levels transmitted in the bitstream followed by inverse transform.

[0142] The following description applies to scalable coding with a quality enhancement layer (where the enhancement layer represents an input video with the same resolution as the base layer but with higher quality or fidelity) and to scalable coding with a spatial enhancement layer (where the enhancement layer has a higher resolution than the base layer, i.e., more samples than the base layer). In the case of a quality enhancement layer, upsampling of the base layer signal is not necessary, for example, within block 220, but is applied to a filter 500 etc. of the samples of the reconstructed base layer. Processing In the case of a spatial enhancement layer, generally, upsampling of the base layer signal is required, for example, within block 220.

[0143] The following-described embodiments support different ways for using the reconstructed base layer samples (compared to 200) or base layer residual samples (compared to 640) for intra prediction of enhancement layer blocks. One or more of the methods described below can be supported in addition to intra layer intra coding (where only the reconstructed enhancement layer samples (compared to 400) are used for intra prediction). The use of a particular method can be signaled at the level of the largest supported block size (such as the macroblock of H.264 / AVC or the block / maximum coding unit of the coding tree of HEVC). Or, it can be signaled at all supported block sizes. Alternatively, it can be signaled for a subset of the supported block sizes.

[0144] For all of the methods described below, the prediction signal can be used directly as the reconstruction signal for the block. That is, no residual is transmitted at all. Or, the selected method for inter-layer intra prediction can be combined with residual coding. In a particular embodiment, the residual signal is transmitted via transform coding. That is, the quantized transform coefficients (transform coefficient levels) are transmitted using an entropy coding technique (e.g., variable length coding or arithmetic coding (compared to 19b)). And the residual is obtained by inverse quantizing (scaling) the transmitted transform coefficient levels and applying an inverse transform (compared to 580). In a particular version, the complete residual block corresponding to the block in which the inter-layer intra prediction signal is generated is transformed using one transform (compared to 726). That is, the entire block is transformed using one transform of the same size as the prediction block. In another embodiment, the prediction block can be further sub-divided into smaller blocks (e.g., using hierarchical decomposition). And for each of the smaller blocks (which can have different block sizes), a separate transform is applied. In another embodiment, the coding unit can be divided into smaller prediction blocks. And for zero or more of the prediction blocks, the prediction signal is generated using one of the methods for inter-layer intra prediction. And then, the residual of the entire coding unit is transformed using one transform (compared to 726). Or, the coding unit is sub-divided into different transform units. Here, the sub-division for forming the transform unit (the block to which one transform is applied) is different from the sub-division for decomposing the coding unit into prediction blocks.

[0145] In certain embodiments, the (upsampled / filtered) reconstructed base layer signal (compared to 380) is used directly as a prediction signal. Multiple methods for using the base layer to intra-predict the enhancement layer include the following. The (upsampled / filtered) reconstructed base layer signal (compared to 380) is used directly as an enhancement layer prediction signal. This method is similar to the well-known H.264 / SVC inter-layer intra prediction mode. In this method, prediction blocks for the enhancement layer are extracted (compared to 220) to match the corresponding sample positions in the enhancement layer and are formed by collocated samples of the base layer reconstruction signal that are optionally filtered before and after extraction. In contrast to the SVC inter-layer intra prediction mode, this mode is supported not only at the macroblock level (or the largest supported block size), but also at any block size. That is, the mode can be signaled not only for the largest supported block size, but also the blocks of the largest supported block size (macroblocks in MPEG4, H.264 and coded tree blocks / largest coded units in HEVC) can be hierarchically sub-divided into smaller blocks / coded units, and the use of the inter-layer intra prediction mode can be signaled at any supported block size (for the corresponding blocks). In certain embodiments, this mode is only supported for the selected block size. Next, the syntax element signaling the use of this mode can be sent only for the corresponding block size. Or, the value of the syntax element signaling the use of this mode (within another coding parameter) can be correspondingly restricted for another block size. Another difference from the inter-layer intra prediction mode in the SVC extension of H.264 / AVC is that the inter-layer intra prediction mode is supported not only when the collocated region in the base layer is intra-coded, but also when the collocated base layer region is inter-coded or partially inter-coded.

[0146] In certain embodiments, spatial intra prediction of the differential signal (see Example A) is performed. The multiplexing method includes the following method. The reconstructed base layer signal (compared to 380), which is (potentially upsampled / filtered), is combined with the spatial intra prediction signal. Therein, the spatial intra prediction (compared to 420) is derived based on differential samples for adjacent blocks (compared to 260). The differential samples represent the difference between the reconstructed enhancement layer signal (compared to 400) and the reconstructed base layer signal (compared to 380), which is (potentially upsampled / filtered).

[0147] FIG. 25 shows the generation of an inter-layer intra prediction signal by the sum 732 of a base layer reconstruction signal 380 (BL Reco) (which is either upsampled or filtered) and a spatial intra prediction using a difference signal 734 (EH Diff) of an already encoded adjacent block 736. Therein, the difference signal (EH Diff) for the already encoded block 736 is generated by subtracting 738 the (upsampled or filtered) base layer reconstruction signal 380 (BL Reco) from the reconstructed enhancement layer signal (compared to 400). If the already encoded / decoded part is indicated by hatching, the current encoded / decoded block / region / part is 28. That is, the inter-layer intra prediction method described in FIG. 25 uses two overlapping input signals to generate a prediction block. For this method, the difference signal 734 is required. The difference signal 734 is the difference between the reconstructed base layer signal 200 juxtaposed with the reconstructed enhancement layer signal 400. The base layer signal 200 is upsampled 220 to match the corresponding sample positions of the enhancement layer and can optionally be filtered before or after upsampling (if it is the case of quality scalable coding, if upsampling is not applied, it can be filtered). In particular, for spatial scalable coding, usually, the difference signal 734 mainly contains high-frequency components. The difference signal 734 is available for all already reconstructed blocks (i.e., all already encoded / decoded enhancement layer blocks). The difference signal 734 for the adjacent samples 742 of the already encoded / decoded block 736 is used as an input to a spatial intra prediction technique (such as the spatial intra prediction mode specified in H.264 / AVC or HEVC). The spatial intra prediction indicated by the arrow 744 generates prediction signals 746 for the various components of the block 28 to be predicted.In certain embodiments, any clipping functions in the process of spatial intra prediction (as known from H.264 / AVC or HEVC) are modified or disabled to match the dynamic range of the differential signal 734. The actually used intra prediction method (one of the provided methods, which can include planar intra prediction, DC intra prediction, or directional intra prediction 744 with any specific angles) is signaled within the bitstream 6b. It is possible to use spatial intra prediction techniques different from those provided in H.264 / AVC and HEVC (methods for generating a prediction signal using samples of already encoded adjacent blocks). The obtained prediction block 746 (using the differential samples of adjacent blocks) is the first part of the final prediction block 420.

[0148] A second portion of the prediction signal is generated using the collocated region 28 within the reconstructed signal 200 of the base layer. For the quality enhancement layer, the collocated base layer samples can be used directly or optionally filtered, for example, by a low-pass filter or a filter 500 that attenuates high-frequency components. For the spatial enhancement layer, the collocated base layer samples are upsampled. For upsampling 220, a FIR filter or a set of FIR filters is used. An IIR filter can also be used. Optionally, the samples 200 of the reconstructed base layer can be filtered prior to upsampling. Alternatively, the base layer prediction signal (the signal obtained after upsampling the base layer) can be filtered after the upsampling stage. The process of reconstructing the base layer can include one or more additional filters such as a non-blocking filter (compared to 120) or an adaptive loop filter (compared to 140). The base layer reconstruction 200 used for upsampling is the reconstructed signal before any loop filter (compared to 200c). Alternatively, it can be the reconstructed signal after the non-blocking filter but before any other filter (compared to 200b). Alternatively, it can be the reconstructed signal after a particular filter or the reconstructed signal after applying all the filters used in the base layer decoding process (compared to 200a).

[0149] Two generated portions of the prediction signal (the spatially predicted difference signal 746 and the potentially filtered / upsampled base layer reconstruction 380) are added 732 sample by sample to form the final prediction signal 420.

[0150] Merely transferring the embodiments just outlined to the embodiments of FIGS. 6 to 10 means that the possibility just outlined of predicting the current block of the enhancement layer signal is supported by each scalable video decoder / encoder as an alternative to the prediction scheme outlined with respect to FIGS. 6 to 10. The mode used is signaled within the enhancement layer substream 6b via respective prediction mode identifiers not shown in FIG. 8.

[0151] In a particular embodiment, intra prediction follows interlayer residual prediction (see Embodiment B). A plurality of methods for generating an intra prediction signal using base layer data includes the following methods. A conventional spatial intra prediction signal (obtained using samples of an adjacent reconstructed enhancement layer) is combined with a base layer residual signal (inverse transform of base layer transform coefficients, or the difference between base layer reconstruction and base layer prediction) (extracted / filtered).

[0152] FIG. 26 shows the generation of an interlayer intra prediction signal 420 by the sum 752 of a base layer residual signal 754 (BL Resi) (upsampled / filtered) and a spatial intra prediction 756 using samples 758 (EH Reco) of a reconstructed enhancement layer of an already encoded adjacent block depicted by a dotted line 762.

[0153] The concept shown in FIG. 26 overlays two prediction signals to form prediction block 420. Therein, one prediction signal 764 is generated from the already reconstructed enhancement layer samples 758, and the other prediction signal 754 is generated from the base layer residual samples 480. The first portion 764 of prediction signal 420 is obtained by applying spatial intra prediction 756 using the reconstructed enhancement layer samples 758. The spatial intra prediction 756 can be one of the methods specified within H.264 / AVC. Or, one of the methods specified within HEVC. Alternatively, it can be another spatial intra prediction technique that generates prediction signal 764 for current block 18 that forms samples 758 of adjacent block 762. The actually used intra prediction method 756 (one of the provided methods, which can include planar intra prediction, DC intra prediction, or directional intra prediction with any particular angle) is signaled within bitstream 6b. It is possible to use a spatial intra prediction technique (a method for generating a prediction signal using samples of already encoded adjacent blocks) different from the methods provided for H.264 / AVC and HEVC. The second portion 754 of prediction signal 420 is generated using the co-located residual signal 480 of the base layer. For the quality enhancement layer, the residual signal can be used such that it is reconstructed within the base layer. Or, the residual signal can additionally be filtered. For the spatial enhancement layer 480, the residual signal is upsampled 220 (to map the base layer sample positions to the enhancement layer sample positions) before it is used as the second portion of the prediction signal. Also, the base layer residual signal 480 can be filtered either before or after the upsampling stage. A FIR filter can be applied to upsample 220 the residual signal. The upsampling process can be configured in a way that it is not filtered across the transform block boundaries within the base layer that are applied for the purpose of upsampling.

[0154] The base layer residual signal 480 used for inter-layer prediction is a residual signal from which the transform coefficient levels of the base layer are obtained by scaling and inverse transformation. Alternatively, the base layer residual signal 480 can be the difference between the reconstructed base layer signal 200 (before or after non-blocking and additional filtering Processing or during any filtering Processing operation) and the prediction signal 660 used within the base layer.

[0155] Two generated signal elements (spatial intra prediction signal 764 and inter-layer residual prediction signal 754) are added 752 to form the final enhancement layer intra prediction signal.

[0156] This means that the prediction mode just outlined with respect to FIG. 26 forms an alternative prediction mode with respect to the currently encoded / decoded portion 28, which can be used or supported by any scalable video decoder / encoder according to FIGS. 6 - 10.

[0157] In a particular embodiment, a weighted prediction of spatial intra prediction and base layer reconstruction (see Example C) is used. This represents a specification that actually discloses a particular implementation of the embodiments outlined above with respect to FIGS. 6 - 10. Thus, the description of such weighted prediction is to be interpreted not only as an alternative to the above embodiments, but also as a description of the possibility of implementing the embodiments outlined above with respect to FIGS. 6 - 10 in a different manner in a particular aspect, different from the given examples.

[0158] Multiple methods for generating an intra prediction signal using base layer data include the following. The (upsampled / filtered) reconstructed base layer signal is combined with a spatial intra prediction signal, where the spatial intra prediction is obtained based on samples of the reconstructed enhancement layer of adjacent blocks. The final prediction signal is obtained by weighting (compared to 41) the spatial prediction signal and the base layer prediction signal in a way that different frequency components use different weightings. This can be achieved, for example, by filtering (compared to 62) the base layer prediction signal (compared to 38) with a low-pass filter and filtering (compared to 64) the spatial intra prediction signal (compared to 34) with a high-pass filter, and then adding (compared to 66) the resulting filtered signals. Or it can be achieved by transforming (compared to 72, 74) the base layer prediction signal (compared to 38) and the enhancement layer prediction signal (compared to 34) based on frequency, and adding (compared to 52) the resulting transformed blocks (compared to 76, 78). Different weighting coefficients (compared to 82, 84) are used for different frequency positions. Then, the resulting transformed block (compared to 42 in FIG. 10) is inverse-transformed (compared to 84) and used as the enhancement layer prediction signal (compared to 54). The resulting transformed block (compared to 42 in FIG. 10) is then inverse-transformed (compared to 84) and used as the enhancement layer prediction signal (compared to 54), or the resulting transformation coefficients are added (compared to 52) to the scaled transmission transformation coefficient levels (compared to 59), and then inverse-transformed (compared to 84) to obtain the reconstructed block (compared to 54) before deblocking and in-loop processing.

[0159] FIG. 27 shows the generation of an inter-layer intra prediction signal by frequency-weighted summation of the (upsampled / filtered) base layer reconstruction signal (BL Reco) and the spatial intra prediction using samples of the reconstructed enhancement layer of already encoded adjacent blocks (EH Reco).

[0160] The concept of FIG. 27 uses two overlapping signals 772, 774 to form the prediction block 420. The first portion 774 of the signal 420 is obtained by applying the spatial intra prediction 776 corresponding to 30 in FIG. 6 using the reconstructed samples 778 of adjacent blocks already configured within the enhancement layer. The second portion 772 of the prediction signal 420 is generated using the collocated reconstructed signal 200 of the base layer. For the quality enhancement layer, the collocated base layer samples 200 can be used directly. Or they can optionally be filtered, for example, by a low-pass filter or a filter that attenuates high-frequency components. For the spatial enhancement layer, the collocated base layer samples are upsampled 220. For upsampling, a FIR filter or a set of FIR filters can be used. It is also possible to use an IIR filter. Optionally, the reconstructed base layer samples can be filtered before upsampling. Alternatively, the base layer prediction signal (the signal obtained after upsampling the base layer) is filtered after the upsampling stage. The process of reconstructing the base layer can include one or more additional filters such as the deblocking filter 120 and the adaptive loop filter 140. The base layer reconstruction 200 used for upsampling can be the reconstructed signal 200c before any of the loop filters 120, 140. Alternatively, it can be the reconstructed signal 200b after the deblocking filter 120 but before another filter. Alternatively, it can be the reconstructed signal 200a after a particular filter, or the reconstructed signal after applying all the filters 120, 140 used in the base layer decoding process.

[0161] When the reference signs used in FIGS. 23 and 24 are compared with the reference signs used in connection with FIGS. 6 to 10, block 220 corresponds to reference sign 38 used in FIG. 6. 39 corresponds to the portion of 380. At least insofar as it relates to the portion juxtaposed to the current portion 28, 420 juxtaposed to the current portion 28 corresponds to 42. The spatial prediction 776 corresponds to 32.

[0162] Two prediction signals (potentially upsampled / filtered base layer reconstruction 386 and enhancement layer intra prediction 782) are combined to form the final prediction signal 420. The method for combining these signals can have the property that different weighting factors are used for different frequency components. In a particular embodiment, the upsampled base layer reconstruction is filtered with a low-pass filter (compared to 62) (it is also possible to filter the base layer reconstruction before upsampling 220). The intra prediction signal (compared to 34 obtained by 30) is filtered with a high-pass filter (compared to 64). The signals filtered by both are added 784 (compared to 66) to form the final prediction signal 420. Although a pair of a low-pass filter and a high-pass filter can represent an orthogonal mirror filter pair, this is not necessarily required.

[0163] In another specific embodiment (compared to FIG. 10), the process of combining the two prediction signals 380 and 782 is realized via a spatial transformation. Both the base layer reconstruction 380 (potentially upsampled / filtered) and the intra prediction signal 782 are transformed (compared to 72, 74) using a spatial transformation. Next, the transform coefficients of both signals (compared to 76, 78) are scaled with appropriate weighting coefficients (compared to 82, 84), and then added (compared to 90) to form the transform coefficient block of the final prediction signal (compared to 42). In one version, the weighting coefficients (compared to 82, 84) are selected such that for each transform coefficient position, the sum of the weighting coefficients for the components of both signals equals 1. In another version, for some or all transform coefficient positions, the sum of the weighting coefficients can be unequal to 1. In a particular version, the weighting coefficients are selected such that for the transform coefficients representing low-frequency components, the weighting coefficient for the base layer reconstruction is greater than the weighting coefficient for the enhancement layer intra prediction signal, and for the transform coefficients representing high-frequency components, the weighting coefficient for the base layer reconstruction is less than the weighting coefficient for the enhancement layer intra prediction signal.

[0164] In one embodiment, the resulting transform coefficient block (compared to 42) (obtained by combining the weighted transform signals for both components) is inverse-transformed (compared to 84) to form the final prediction signal 420 (compared to 54). In another embodiment, the prediction is made directly in the transform domain. That is, the encoded transform coefficient levels (compared to 59) are scaled (i.e., inverse quantized) to the transform coefficients of the prediction signal (compared to 42) (obtained by adding the weighted transform signals for both components), and then, although not shown in FIG. 10, (potentially non-blocking 120 and further in-loop filtering ProcessingBefore stage 140, it is added (compared to 52) to the block resulting from the conversion coefficients that are inverse-transformed (compared to 84) to obtain the reconstructed signal 420 for the current block. In other words, in the first embodiment, the transformed block obtained by adding the weighted transform signals for both components is inverse-transformed and used as the enhancement layer prediction signal. Alternatively, in the second embodiment, the obtained transform coefficients can be added to the scaled transmitted transform coefficient levels and inverse-transformed to obtain the reconstructed block before deblocking and in-loop processing.

[0165] Selection between the base layer reconstruction and the residual signal (see Example D) can also be used. For the method of using the reconstructed base layer signal (described above), the following versions can be used.

[0166] · The reconstructed base layer samples 200c before deblocking 120 and further in-loop processing 140 (such as an adaptive offset filter or an adaptive loop filter as samples). · The reconstructed base layer samples 200b after deblocking 120, but before another in-loop processing 140 (such as an adaptive offset filter or an adaptive loop filter as samples). · The reconstructed base layer samples 200a after deblocking 120 and further in-loop processing 140 (such as an adaptive offset filter or an adaptive loop filter as samples), or between multiple in-loop processing steps.

[0167] The selection of the corresponding base layer signals 200a, b, c can be fixed for the implementation of a particular decoder (and encoder). Or, it can be signaled within the bitstream 6. For the latter case, different versions can be used. The use of a particular version of the base layer signal can be signaled at the sequence level, or at the picture level, or at the slice level, or at the maximum coding unit level, or at the coding unit level, or at the prediction block level, or at the transform block level, or at any other block level. In another version, the selection can depend on other coding parameters (such as the coding mode) or on the characteristics of the base layer signal.

[0168] In another embodiment, multiple versions of a method of using the (upsampled / filtered) base layer signal 200 can be used. For example, two different modes of directly using the upsampled base layer signal, i.e., 200a, can be provided. Therein, the two modes can use different interpolation filters, or one mode can use an additional filter Processing 500 of the (upsampled) base layer reconstruction signal. Similarly, multiple different versions for the other modes described above can be provided. The upsampled / filtered base layer signal 380 employed for different versions of the mode can differ within the interpolation filter used (including the interpolation filter that also filters the integer sample positions). Or, the upsampled / filtered base layer signal 380 for a second version can be obtained by filtering the extracted / filtered base layer for the first version with 500. One selection of the various versions can be signaled at the sequence, picture, slice, maximum coding unit, coding unit level, prediction block level, or transform block level. Or, it can be inferred from the characteristics of the corresponding reconstructed base layer signal or the transmitted coding parameters.

[0169] The same applies to the mode using the reconstructed base layer residual signal via 480. Here, the interpolation filter or additional filter used Processing Different versions from the step can also be used.

[0170] Different filters can be used to upsample / filter the reconstructed base layer signal and the base layer residual signal. This means that different approaches are used for the upsampling of the base layer residual signal than for the upsampling of the base layer reconstruction signal.

[0171] For the base layer block, the residual signal is zero (i.e., the transform coefficient level is not sent to the block at all). The corresponding base layer residual signal can be replaced with another signal obtained from the base layer. For example, this can be a high-pass filter version of the reconstructed base layer block, or a different similar signal obtained from samples of the reconstructed base layer residuals of adjacent blocks.

[0172] As far as the samples used for spatial intra prediction in the enhancement layer (see Example H) are concerned, the following special processing can be provided. For the mode using spatial intra prediction, the adjacent samples not available in the enhancement layer (the adjacent blocks are encoded after the current block, so the adjacent samples are not available) can be replaced with the corresponding samples of the upsampled / filtered base layer signal.

[0173] As far as the coding of the intra prediction mode (see Example X) is concerned, the following special modes and functionality can be provided. For a mode that uses spatial intra prediction as in 30a, the coding of the intra prediction mode can (if available) change the information about the intra prediction mode in the base layer in a way that is used to more efficiently code the intra prediction mode in the enhancement layer. This can be used, for example, for parameter 56. If the collocated region in the base layer (compared to 36) is intra-coded using a particular spatial intra prediction mode, then a similar intra prediction mode is likely to be used within the enhancement layer block (compared to 28). The intra prediction mode is usually signaled in a way that within the set of possible intra prediction modes, one or more modes are classified as the most likely modes. There, it is signaled with a shorter codeword. Alternatively, the shorter the arithmetic coding, the fewer bits the alternative brings. Within the intra prediction of HEVC, (if available) the intra prediction mode of the upper block and (if available) the intra prediction mode of the left block are included within the set of the most likely modes. In addition to these modes, one or more additional modes (often used) are included within the list of the most likely modes. There, the actual additional modes depend on the usefulness of the intra prediction modes of the block above the current block and the block to the left of the current block. Within HEVC, three modes are accurately classified as the most likely modes. Within H.264 / AVC, one mode is classified as the most likely mode. This mode is obtained based on the intra prediction modes used for the block above the current block and the block to the left of the current block. Any other concept (different from H.264 / AVC and HEVC) for classifying the intra prediction mode is possible and can be used for the following extensions.

[0174] To use base layer data for efficient coding of intra prediction modes within an enhancement layer, the concept of using one or more most likely modes is modified (if the corresponding base layer block is intra coded) in a way that the most likely modes include the intra prediction modes used within the collocated base layer blocks. In certain embodiments, the following approach is used. Given a current enhancement layer block, the collocated base layer block is determined. In a particular version, the collocated base layer block is the base layer block covering the collocated position of the top left sample of the enhancement block. In another version, the collocated base layer block is the base layer block covering the collocated position of the sample at the center of the enhancement block. In other versions, another sample within the enhancement layer block can be used to determine the collocated base layer block. If the determined collocated base layer block is intra coded and the base layer intra prediction mode specifies an angular intra prediction mode and the intra prediction mode obtained from the enhancement layer block to the left of the current enhancement layer block does not use the angular intra prediction mode, then the intra prediction mode obtained from the left enhancement layer block is replaced with the corresponding base layer intra prediction mode. Otherwise, if the determined collocated base layer block is intra coded, the base layer intra prediction mode specifies an angular intra prediction mode, and the intra prediction mode obtained from the enhancement layer block above the current enhancement layer block does not use the angular intra prediction mode, then the intra prediction mode obtained from the above enhancement layer block is replaced with the corresponding base layer intra prediction mode. In other versions, different approaches are used to modify the list of most likely modes (consisting of a single element) using the base layer intra prediction mode.

[0175] Intercoding techniques for the spatial and quality enhancement layers are then provided.

[0176] In the state-of-the-art hybrid video coding standards (such as H.264 / AVC or the upcoming HEVC), pictures of a video sequence are partitioned into blocks of samples. The block size is fixed or the coding technique can provide a hierarchical structure that allows the block to be further sub-divided into blocks with smaller block sizes. The reconstruction of a block is usually obtained by generating a prediction signal for the block and adding the transmitted residual signal. The residual signal is usually transmitted using transform coding. This means that a quantization index list for the transform coefficients (also called transform coefficient levels) is transmitted using entropy coding techniques. And on the decoder side, these transmitted transform coefficient levels are scaled and then inverse-transformed to obtain the residual signal that is added to the prediction signal. The residual signal is generated either by intra prediction (using only data already transmitted for the current time instant) or by inter prediction (using data already transmitted for different time instants).

[0177] In inter prediction, the predicted block is obtained by motion-compensated prediction using samples of a previously reconstructed frame. This can be done by uni-directional prediction (using one reference picture and a set of motion parameters). Alternatively, the prediction signal can be generated by multi-hypothesis prediction. In the latter case, two or more prediction signals are superimposed. That is, for each sample, a weighted average is constructed to form the final prediction signal. The multiple prediction signals (that are superimposed) can be generated using different motion parameters for different hypotheses (e.g., different reference pictures or motion vectors). Also, for uni-directional prediction, it is possible to multiply samples of the motion-compensated prediction signal with a constant coefficient and add a constant offset to form the final prediction signal. Also, such scaling and offset correction can be used for all or selected hypotheses within multi-hypothesis prediction.

[0178] Even within scalable video coding, base layer information can be utilized to support the inter-prediction process for the enhancement layer. The SVC extension of the state-of-the-art video coding standard for scalable coding, H.264 / AVC, has one additional mode for improving the coding efficiency of the inter-prediction process within the enhancement layer. This mode is signaled at the macroblock level (a block of 16×16 luma samples). Within this mode, the reconstructed residual samples in the lower layer are used to improve the motion-compensated prediction signal in the enhancement layer. This mode is also called inter-layer residual prediction. If this mode is selected for a macroblock within the quality enhancement layer, the inter-layer prediction signal is assembled by the collocated samples of the reconstructed lower layer residual signal. If the inter-layer residual prediction mode is selected within the spatial enhancement layer, the prediction signal is generated by upsampling the collocated reconstructed base layer residual signal. For upsampling, an FIR filter is used. However, the filter Processing is not applied across the transform block boundaries. The prediction signal generated from the samples of the reconstructed base layer residual is added to the conventional motion-compensated prediction signal to form the final prediction signal for the enhancement layer block. Generally, for the inter-layer residual prediction mode, an additional residual signal is transmitted by transform coding. The transmission of the residual signal can be omitted (inferred to be equal to zero) if it is correspondingly signaled within the bitstream. The final reconstructed signal is obtained by adding the reconstructed residual signal, which is obtained by scaling the transmitted transform coefficient levels and applying the inverse spatial transform, to the prediction signal (where the inter-layer residual prediction signal is obtained by adding it to the motion-compensated prediction signal).

[0179] Next, a technique for inter-coding of enhancement layer signals is described. This section describes a method for employing the base layer signal in addition to the already reconstructed enhancement layer signal for inter-predicting the enhancement layer signal to be encoded within a scalable video coding scenario. By employing the base layer signal for inter-predicting the enhancement layer signal to be encoded, the prediction error can be sufficiently suppressed. This results in an overall bit transmission rate saving for the encoding of the enhancement layer. The main focus of this section is to increase the block-based motion compensation of the samples of the enhancement layer using samples of the already encoded enhancement layer with additional signals from the base layer. The following description provides possibilities for using various signals from the encoded base layer. Although quadtree block partitioning is generally adopted as a preferred embodiment, the presented examples can be applied to a general block-based hybrid coding approach without assuming any specific block partitioning. The use of the base layer reconstruction of the current time index, the base layer residual of the current time index, or even the base layer reconstruction of the already encoded picture for inter-predicting the enhancement layer blocks to be encoded is described. Also, a method is described for combining the base layer signal with the already encoded enhancement layer signal to obtain a better prediction for the current enhancement layer. One of the state-of-the-art main techniques is the inter-layer residual prediction within H.264 / SVC. The inter-layer residual prediction within H.264 / SVC can be employed for all inter-coded macroblocks, regardless of whether they are coded or not, using the SVC macroblock types signaled by their use of either the base mode flag or the conventional macroblock type. The flag is added to the macroblock syntax that signals the usage of the inter-layer residual prediction for the spatial and quality enhancement layers. When this residual prediction flag is equal to 1, the residual signal of the corresponding region within the reference layer is upsampled block-by-block using a bilinear filter and used as a prediction for the residual signal of the enhancement layer macroblock. As a result, only the corresponding difference signal needs to be coded within the enhancement layer. For the description in this section, the following notations are used. t 0 := the temporal index of the current picture t 1 := the temporal index of the already reconstructed picture EL := enhancement layer BL := base layer EL(t 0 ):= the current enhancement layer picture to be coded EL_reco := enhancement layer reconstruction BL_reco := base layer reconstruction BL_resi := base layer residual signal (inverse transform of the base layer transform coefficients, or the difference between the base layer reconstruction and the base layer prediction) EL_diff := the difference between the enhancement layer reconstruction and the upsampled / filtered base layer reconstruction The different base layer signals and enhancement layer signals are used within the description explained in Figure 28.

[0180] For the specification, the following characteristics of the filter are used. · Linearity: Although many of the filters referenced within the specification are linear, non-linear filters are also used. · Number of output samples: In the upsampling operation, the number of output samples is greater than the number of input samples. Here, the filter of the input data Processing creates more samples than the input value. Conventional filters Processing have an output sample count equal to the number of input samples. Such a filter Processing operation can be used, for example, within high-quality scalable coding. · Phase delay: For samples at integer positions in the filter Processing the phase delay is usually zero (or a delay of an integer value within the sample). For the generation of samples at fractional positions (e.g., half-pel or quarter-pel positions), a filter with a fractional delay (within the unit of the sample) is usually applied to the samples of the integer grid.

[0181] Conventional motion compensation prediction used in all hybrid video coding standards (e.g., MPEG-2, H.264 / AVC, or the upcoming HEVC standard) is illustrated in Figure 29. To predict the signal of the current block, the region of the already reconstructed image is replaced and used as the prediction signal. For the signaling of the replacement, motion vectors are usually encoded within the bitstream. For motion vectors with integer sample accuracy, the reference region in the reference image can be directly copied to form the prediction signal. However, it is also possible to transmit motion vectors with fractional sample accuracy. In this case, the prediction signal is filtered with a filter having a fractional sample delay from the reference signal Perform processingIt is obtained by. The reference image used can usually be specified by including a reference image index in the bitstream syntax. Generally, it is also possible to superimpose two or more prediction signals to form a final prediction signal. The concept is supported, for example, within a B slice with two motion hypotheses. In this case, the multiple prediction signals are generated using different motion parameters (e.g., different reference images or motion vectors) for different hypotheses. For single-direction prediction, it is also possible to multiply samples of a motion-compensated prediction signal with a certain coefficient and add a certain offset to form the final prediction signal. Such scaling and offset correction can also be used for all or selected hypotheses within multi-hypothesis prediction.

[0182] The following description applies to scalable coding with a quality enhancement layer (the enhancement layer has the same resolution as the base layer but represents an input video with higher quality or fidelity) and scalable coding with a spatial enhancement layer (the enhancement layer has a higher resolution than the base layer, i.e., a larger number of samples). For the quality enhancement layer, upsampling of the base layer signal is not necessary, but filtering of the samples of the reconstructed base layer Processing is applied. In the case of the spatial enhancement layer, upsampling of the base layer signal is generally necessary.

[0183] Embodiments support various methods for using reconstructed base layer samples or base layer residual samples for inter prediction of enhancement layer blocks. It is possible to support one or more of the methods described below by adding conventional inter prediction and intra prediction. The usage of a particular method can be signaled at the level of the largest supported block size (such as a macroblock in H.264 / AVC, or a coded tree block / maximum coded unit in HEVC, etc.). It can be signaled for all supported block sizes that are signaled. Or, it can be signaled for a subset of the supported block sizes.

[0184] For all the methods described below, the prediction signal can be used directly as the reconstruction signal for the block. Alternatively, the selected method for inter-layer prediction can be combined with residual coding. In certain embodiments, the residual signal is transmitted via transform coding. That is, the quantized transform coefficients (transform coefficient levels) are transmitted using entropy coding techniques (e.g., variable length coding or arithmetic coding), and the residual is obtained by inverse quantizing (scaling) the transmitted transform coefficient levels and applying an inverse transform. In a particular version, the complete residual block corresponding to the block where the inter-layer prediction signal is generated is transformed using a single transform (i.e., the entire block is transformed using a single transform of the same size as the prediction block). In another embodiment, the prediction block can be further sub-divided into smaller blocks (e.g., using hierarchical decomposition). For each of the smaller blocks having different block sizes, separate transforms are applied. In another embodiment, the coding unit can be divided into smaller prediction blocks. For zero or more of the prediction blocks, the prediction signal is generated using one of the methods for inter-layer prediction. Then, the residual of the entire coding unit is transformed using a single transform. Or, the coding unit is sub-divided into different transform units. There, the sub-division for forming the transform unit (the block to which a single transform is applied) is different from the sub-division for decomposing the coding unit into prediction blocks.

[0185] Below, the possibility of performing prediction using the base layer residual and enhancement layer reconstruction is explained. The multiple methods include the following. The conventional inter-prediction signal (obtained by motion-compensated interpolation of the already reconstructed enhancement layer image) is combined with the base layer residual signal (the inverse transform of the base layer transform coefficients, or the difference between the base layer reconstruction and the base layer prediction, upsampled / filtered). This method is also called the BL_resi mode (illustrated in FIG. 30).

[0186] TIFF0007682959000001.tif238170

[0187] The multiple method for prediction using base layer reconstruction and enhancement layer difference signals (see Example J) includes the following method. The reconstructed base layer signal (upsampled / filtered) is combined with the motion compensated prediction signal. There, the motion compensated prediction signal is obtained by the motion compensated difference image. The difference image represents the difference between the reference image and the difference between the reconstructed enhancement layer signal and the reconstructed base layer signal (upsampled / filtered). This method is also called the BL_reco mode.

[0188] TIFF0007682959000002.tif23170

[0189] TIFF0007682959000003.tif19170

[0190] For the EL difference signal, the following version is used. · The difference between the EL reconstruction and the upsampled / filtered BL reconstruction, or, · A loop filter (such as non-blocking, SAO, ALF) Processing The difference between the EL reconstruction before or during the stage and the upsampled / filtered BL reconstruction.

[0191] The usage of a specific version can be fixed in the decoder, or it can be signaled at the sequence level, image level, slice level, maximum coding unit level, coding unit level, or another partition level. Alternatively, it can depend on another coding parameter.

[0192] When the EL differential signal is defined using the difference between the EL reconstruction and the BL reconstruction that has been extracted and filtered, this makes it possible to save the EL and BL reconstructions and calculate the EL differential signal of the block using the prediction mode in place. As a result, the memory required to store the EL differential signal can be saved. However, it incurs a slight computational overhead.

[0193] The MCP filter used on the EL differential image can be of integer or fractional sample precision. · For the MCP of the differential image, an interpolation filter different from the MCP of the reconstructed image can be used. · For the MCP of the differential image, the interpolation filter can be selected based on the characteristics of the corresponding region within the differential image (or based on the encoding parameters within the bitstream, or based on the transmitted information).

[0194] The motion vector MV(x,y,t) is defined to point to a specific position within the EL differential image. The parameters x and y refer to the spatial position within the image, and the parameter t is used to specify the time index of the differential image.

[0195] The integer part of the MV is used to obtain a set of samples from the differential image, and the fractional part of the MV is used to select the MCP filter from a set of filters. The obtained differential samples are filtered to produce the filtered differential samples.

[0196] The dynamic range of the difference image can theoretically exceed the dynamic range of the original image. Assuming an 8-bit representation of the image within the range [0 255], the difference image can have a range of [-255 255]. However, in practice, most of the amplitudes are distributed around the vicinity of ±0. In a preferred embodiment for storing the difference image, a fixed offset of 128 is added, and the result is clipped to the range [0 255] and stored as a normal 8-bit image. Then, within the encoding and decoding processes, the offset of 128 is subtracted from the difference amplitude read from the difference image.

[0197] For the method of using the reconstructed BL signal, the following versions can be used. This can be fixed, or it can be signaled at the sequence level, image level, slice level, maximum coding unit level, coding unit level, or another partition level. Alternatively, it can depend on another coding parameter. · Reconstructed base layer samples before non-blocking and further in-loop processing (such samples as adaptive offset filters or adaptive loop filters). · Reconstructed base layer samples after non-blocking but before further in-loop processing (such samples as adaptive offset filters or adaptive loop filters). · Reconstructed base layer samples after non-blocking and further in-loop processing (such samples as adaptive offset filters or adaptive loop filters), or reconstructed base layer samples between multiple in-loop processing steps.

[0198] To calculate the EL prediction component from the current BL reconstruction, the region in the BL image juxtaposed to the region considered in the EL image is identified. Then, the reconstruction signal is taken from the identified BL region. The definition of the juxtaposed region can be made such that it results in the same EL resolution as an integer scaling factor of the BL resolution (e.g., 2× scalability), or a fractional scaling factor of the BL resolution (e.g., 1.5× scalability), or the BL resolution (e.g., SNR scalability). In the case of SNR scalability, the juxtaposed blocks in the BL image have the same coordinates as the EL blocks to be predicted.

[0199] The final EL prediction is obtained by adding the filtered EL difference samples and the filtered BL reconstruction samples.

[0200] (Upsampled / filtered) Some possible Variation modes of combining the base layer reconstruction signal and the motion-compensated enhancement layer difference signal are described below. · Multiple versions of the method using the (upsampled / filtered) BL signal can be used. The upsampled / filtered BL signal employed for these versions can be different from the interpolation filter (including the interpolation filter applied to the integer sample positions used), or the upsampled / filtered BL signal for the second version can be obtained by filtering the upsampled / filtered BL signal for the first version. One selection of different versions can be signaled in sequence, in the image, in the slice, at the maximum coding unit, at the coding unit level, or at another level of the image partition. Alternatively, it can be inferred from the characteristics of the corresponding reconstructed BL signal or the transmitted coded parameters. · Various filters can be used for the upsampled / filtered BL reconstruction signal in the BL_reco mode and the BL residual signal in the BL_resi mode. · The upsampled / filtered BL signal can also be combined into two or more hypotheses of the motion-compensated difference signal. This is illustrated in Figure 32.

[0201] Considering the above, the prediction can be performed using a combination of base layer reconstruction and enhancement layer reconstruction (see Example C). One major difference from the above description regarding Figures 11, 12, and 13 is the coding mode for obtaining the intra layer prediction 34 that is performed more temporally rather than spatially. That is, instead of the spatial prediction 30, the temporal prediction 32 is used to form the intra layer prediction signal 34. Therefore, some of the examples described below can be easily applied to the respective embodiments above Figures 6 - 10 and Figures 11 - 13. The multiple methods include the following. The reconstructed base layer signal (upsampled / filtered) is combined into the inter prediction signal. There, the inter prediction is obtained by motion-compensated prediction using the reconstructed enhancement layer image. The final prediction signal is obtained by weighting the inter prediction signal and the base layer prediction signal in a way that different frequency components use different weightings. For example, this can be achieved by any of the following. · Filtering the base layer prediction signal with a low-pass filter, filtering the inter prediction signal with a high-pass filter, and adding the resulting filtered signals. ·Converting the base layer prediction signal and the inter prediction signal, and superimposing the obtained conversion blocks. There, various weighting coefficients are used for various frequency positions. Next, the obtained conversion block is inverse-transformed and can be used as the enhancement layer prediction signal. Alternatively, the obtained conversion coefficients can be added to the scaled transmitted conversion coefficient levels and inverse-transformed to obtain a block reconstructed before deblocking and in-loop processing.

[0202] This mode can also be called the BL_comb mode described in FIG. 33.

[0203] TIFF0007682959000004.tif18167

[0204] In a preferred embodiment, the weighting is made depending on the ratio of the EL resolution to the BL resolution. For example, when BL is to be scaled up by a coefficient within the range [1 1.25], a predetermined set of weightings for EL and BL reconstruction can be used. When BL is to be scaled up by a coefficient within the range [1.25 1.75], various sets of weightings can be used. When BL is to be scaled up by a coefficient of 1.75 or more, another different set of weightings and the like can be used.

[0205] Rendering a specific weighting that depends on the scaling factor separation base layer and the enhancement layer is also possible in another embodiment regarding spatial intra layer prediction.

[0206] In another preferred embodiment, the weighting is made depending on the EL block size to be predicted. For example, for a 4×4 block in EL, a weighting matrix that specifies the weighting for the EL reconstruction conversion coefficients is defined. And another weighting matrix that specifies the weighting for the BL reconstruction conversion coefficients can be defined. The weighting matrix for the BL reconstruction conversion coefficients can be, for example, as follows. 64,63,61,49, 63,62,57,40, 61,56,44,28, 49,46,32,15, And the weighting matrix for the EL reconstruction conversion coefficients can be formed, for example, as follows. 0,2,8,24, 3,7,16,32, 9,18,20,26, 22,31,30,23,

[0207] Similarly, for block sizes such as 8×8, 16×16, 32×32, etc., a separate weighting matrix can be defined.

[0208] The actual conversion used for frequency domain weighting can be the same as, or different from, the conversion used to encode the prediction residual. For example, the integer approximation for DCT can be used both for frequency domain weighting and for calculating the conversion coefficients of the prediction residual to be encoded within the frequency domain.

[0209] In another preferred embodiment, the maximum conversion size is defined for frequency domain weighting in order to limit the computational complexity. If the EL block size under consideration is larger than the maximum conversion size, then the EL reconstruction and the BL reconstruction are spatially separated into a series of adjacent sub-blocks. The frequency domain weighting is performed on the sub-blocks, and the final prediction signal is formed by assembling the weighted results.

[0210] Moreover, the weighting can be performed on the luminance and chrominance components, or a selected subset of the color components.

[0211] The following describes various possibilities for obtaining enhancement layer coding parameters. The coding (or prediction) parameters to be used to reconstruct the enhancement layer blocks are obtained by multiple methods from the collocated coding parameters in the base layer. The base layer and the enhancement layer can have different spatial resolutions, or they can have the same spatial resolution.

[0212] In the scalable video extension of H.264 / AVC, inter-layer motion prediction is performed for macroblock types signaled by the syntax element base mode flag. If the base mode flag is equal to 1 and the corresponding reference macroblock in the base layer is inter-coded, then the enhancement layer macroblock is also inter-coded, and all motion parameters are inferred from the collocated base layer blocks. Otherwise (the base mode flag is equal to 0), for each motion vector, a so-called motion prediction flag, a syntax element is sent, and it is specified whether the base layer motion vector is used as a motion vector predictor. If the motion prediction flag is equal to 1, the motion vector predictor of the collocated reference block in the base layer is scaled according to the resolution ratio and used as a motion vector predictor. If the motion prediction flag is equal to 0, the motion vector predictor is calculated as specified within H.264 / AVC.

[0213] The following describes a method for obtaining enhancement layer coding parameters. The sample array associated with the base layer image is decomposed into blocks, and each block is associated with coding (or prediction) parameters. In other words, all sample positions within a particular block have a particular associated coding (or prediction) parameter. The coding parameters can include parameters for motion compensation prediction, including the number of motion hypotheses, reference index lists, motion vectors, motion vector predictor identifiers, and merge identifiers. The coding parameters can also include intra prediction parameters such as the intra prediction direction.

[0214] That the blocks in the enhancement layer are encoded using collocated information from the base layer can be signaled within the bitstream.

[0215] For example, the derivation of the enhancement layer coding parameters (see Example T) can be made as follows. For the N×M blocks of the enhancement layer signaled using collocated base layer information, the coding parameters related to the sample positions within the block are derived based on the coding parameters related to the collocated sample positions within the base layer sample array.

[0216] In certain embodiments, this process is done by the following steps. 1. Derivation of the coding parameters for each sample position within the N×M enhancement layer block based on the base layer coding parameters. 2. Derivation of the partitioning of the N×M enhancement layer blocks in the sub-block such that all sample positions within a particular sub-block have the same related coding parameters.

[0217] Also, the second step can be omitted.

[0218] TIFF0007682959000005.tif20169

[0219] TIFF0007682959000006.tif75170

[0220] TIFF0007682959000007.tif19170

[0221] TIFF0007682959000008.tif26169

[0222] TIFF0007682959000009.tif13169

[0223] Since each sample position is related to the prediction parameters after step 1, the samples of each enhancement layer can be predicted after step 1. Nevertheless, in step 2, block partitions are obtained to perform the prediction operation of samples of larger blocks or to transform and encode prediction residuals within the blocks of the derived partitions.

[0224] Step 2 can be performed by classifying the enhancement layer sample positions into square or rectangular blocks. Each is decomposed into one of a set of possible decompositions within the sub-blocks. The square or rectangular blocks correspond to the leaves in a quadtree structure where they can exist at different levels represented in FIG. 34.

[0225] The level and decomposition of each square or rectangular block can be determined by performing the following ordered steps. a) Set the highest level to the level corresponding to a block of size N×M. Set the current level to the lowest level (i.e., the level where the square or rectangular block contains a single block of the smallest block size). Go to step b). b) For each square or rectangular block at the current level, if there exists an allowed decomposition of the square or rectangular block, all sample positions within each sub-block are related to the same coding parameters or (according to some difference magnitude) related to the coding parameters with a small difference. That decomposition is a candidate decomposition. Among all candidate decompositions, select the one that decomposes the square or rectangular block into the smallest number of sub-blocks. If the current level is the highest level, go to step c). Otherwise, set the current level to the next higher level and go to step b). c) End

[0226] The function can be selected in such a way that at a certain level within step b), there always exists at least one candidate decomposition.

[0227] Grouping of blocks having the same encoding parameters is not limited to square blocks, but the blocks can be grouped into rectangular blocks. Further, the grouping is not limited to a quadtree structure. It is also possible to use a decomposition structure in which a block is decomposed into two rectangular blocks of the same size or two rectangular blocks of different sizes. It is also possible to use a decomposition structure that uses quadtree decomposition up to a certain level and then uses decomposition into two rectangular blocks. Also, any other block decomposition is possible.

[0228] In contrast to the motion parameter prediction mode between SVC layers, the described mode is supported not only at the macroblock level (or the most widely supported block size), but also at any block size. It means that the mode can be signaled not only for the most widely supported block size, but also that the blocks of the most widely supported block size (macroblocks in MPEG4, H.264, and coding tree blocks / largest coding units in HEVC) are hierarchically subdivided into smaller blocks / coding units, and the usage of the inter-layer motion mode can be signaled for any supported block size (for the corresponding blocks). In a particular embodiment, this mode supports only the selected block size. Then, the syntax element signaling the usage of this mode can be sent only for the corresponding block size. Or, the value of the syntax element signaling the usage of this mode (within another coding parameter) can be restricted corresponding to another block size. Also, the difference from the inter-layer motion parameter prediction mode in the SVC extension of H.264 / AVC is that the blocks encoded in this mode are not fully inter-encoded. The block can include intra-encoded sub-blocks depending on the collocated base layer signal.

[0229] One of several methods for reconstructing samples of an enhancement layer block using the encoded parameters obtained by the method described above can be signaled within the bitstream. Such a method for predicting an enhancement layer block using the obtained encoded parameters can include the following. · For motion compensation, obtaining a prediction signal for the enhancement layer block using the obtained motion parameters and the reconstructed enhancement layer reference image. · A combination of (a) the (upsampled / filtered) base layer reconstruction for the current image and (b) the motion compensation signal, generated by subtracting the (upsampled / filtered) base layer reconstruction from the reconstructed enhancement layer image, using the obtained motion parameters and the enhancement layer reference image. · A combination of (a) the (upsampled / filtered) base layer residual for the current image (the difference between the reconstructed signal and the prediction, or the inverse transform of the encoded transform coefficient values) and (b) the motion compensation signal, using the obtained motion parameters and the reconstructed enhancement layer reference image, generated by subtracting the (upsampled / filtered) base layer reconstruction from the reconstructed enhancement layer image.

[0230] While another sub-block is classified as being inter-coded, a process for obtaining partitions within a smaller block for the current block and obtaining coding parameters for the sub-block can be classified into several sub-blocks as being intra-coded. For an inter-coded sub-block, motion parameters are obtained from the collocated base layer block. However, if the collocated base layer block is intra-coded, the corresponding sub-block in the enhancement layer can be classified as being intra-coded. For samples of such intra-coded sub-blocks, enhancement layer signals can be predicted by using information from the base layer. For example, · The (upsampled / filtered) version of the corresponding base layer reconstruction is used as the intra-prediction signal. · The obtained intra-prediction parameters are used for spatial intra-prediction within the enhancement layer.

[0231] The following embodiments for predicting an enhancement layer block using a weighted combination of prediction signals include a method for generating a prediction signal for an enhancement layer block by combining (a) an enhancement layer intra-prediction signal obtained by spatial or temporal (i.e., motion compensated) prediction using samples of the reconstructed enhancement layer, and (b) a base layer prediction signal that is the (upsampled / filtered) base layer reconstruction for the current image. The final prediction signal is obtained by weighting the enhancement layer prediction signal and the base layer prediction signal within the enhancement layer in a manner where a weight according to a weighting function is used for each sample.

[0232] TIFF0007682959000010.tif47170

[0233] Different weighting functions can be used for different block sizes of the current block to be predicted. Also, the weighting function can be varied according to the temporal distance of the reference images from which the inter-prediction hypotheses are obtained.

[0234] In the case of an enhancement layer intra-prediction signal, the weighting function is realized using different weightings that depend, for example, on the position within the current block to be predicted.

[0235] In a preferred embodiment, a method for obtaining enhancement layer coding parameters is used. And step 2 of the method uses a set of possible decompositions of a square block as described in FIG. 35.

[0236] TIFF0007682959000011.tif23170

[0237] TIFF0007682959000012.tif51170

[0238] Also, combinations of the above embodiments are possible.

[0239] In another embodiment, for an enhancement layer block for which it has been signaled that collocated base layer information is used, associate an initial set of motion parameters with the enhancement layer sample positions having derived intra-prediction parameters so that the block can merge with the block containing these samples (i.e., copy the default set of motion parameters). The initial set of motion parameters consists of an indicator for using one or two hypotheses, a list of reference indices referring to the first image in the list of reference images, and a motion vector with zero padding.

[0240] In another embodiment, using the juxtaposed base layer information, for the enhancement layer blocks to be signaled, the enhancement layer samples having the obtained motion parameters are first predicted and reconstructed in a certain order. Thereafter, the samples having the obtained intra prediction parameters are predicted in the intra reconstruction order. As a result, intra prediction can use the already reconstructed sample values from (a) adjacent inter prediction blocks and (b) the previous adjacent intra prediction blocks within the intra reconstruction order.

[0241] In another embodiment, for the merged (i.e., taking motion parameters obtained from another inter prediction block) enhancement layer blocks, the list of merge candidates additionally includes candidates from the corresponding base layer blocks. And if the enhancement layer has a higher spatial sampling rate than the base layer, additionally, the maximum four candidates obtained from the base layer candidates are included by improving only the available adjacent values within the enhancement layer for the spatial replacement components.

[0242] In another embodiment, the magnitude of the difference used in step 2b) asserts that there is a small difference within the sub - block only if the difference disappears completely. That is, the sub - block is formed only when all the included sample positions have the same obtained coded parameters.

[0243] In another embodiment, if (a) all included sample positions have the obtained motion parameters and pairs of sample positions within a block do not have derived motion parameters that differ by more than a specific value according to the vector norm applied to the corresponding motion vectors, or (b) all included sample positions have the obtained intra prediction parameters and pairs of sample positions within a block do not have obtained intra prediction parameters that differ by more than a specific angle of the intra prediction in direction, the magnitude of the difference used in step 2b) asserts that there are small differences within the sub-block. The parameters resulting for the sub-block are calculated by an average or median operation. In another embodiment, the partition obtained by inferring the coding parameters from the base layer can be further refined based on side information signaled within the bitstream. In another embodiment, the residual coding for a block for which the coding parameters are inferred from the base layer is independent of the partition into the block inferred from the base layer. For example, it means that although the inference of the coding parameters from the base layer partitions the block into several sub-blocks each having a separate set of coding parameters, a single transform can be applied to the block. Or, the block for which the partition and coding parameters for the sub-blocks are inferred from the base layer can be split into smaller blocks for the purpose of transform coding the residual. There, the splitting into transform blocks is independent of the inferred partition within the block having different coding parameters.

[0244] In another embodiment, the residual coding for the blocks whose coding parameters are inferred from the base layer depends on the partitions within the blocks inferred from the base layer. For example, for transform coding, it means that the partitioning of the blocks within the transform block depends on the partitions inferred from the base layer. In one version, a single transform can be applied to each sub-block having different coding parameters. In another version, the partitions can be refined based on side information included in the bitstream. In another version, some sub-blocks can be grouped into larger blocks so that the residual signals are signaled in the bitstream for the purpose of transform coding.

[0245] Also, embodiments obtained by combinations of the above-described embodiments are also possible.

[0246] In connection with enhancement layer motion vector coding, the following part describes a method for reducing motion information in a scalable video coding application by providing multiple enhancement layer predictors and using the motion information coded within the base layer to efficiently code the motion information of the enhancement layer. This idea is suitable for scalable video coding including spatial, temporal and quality scalability.

[0247] In the scalable video extension of H.264 / AVC inter-layer motion prediction, motion prediction is performed for macroblock types signaled by a syntax element base mode flag. If the base mode flag is equal to 1 and the corresponding reference macroblock in the base layer is inter-coded, then the enhancement layer macroblock is also inter-coded. And all motion parameters are inferred from the collocated base layer blocks. Otherwise (if the base mode flag is equal to 0), the syntax elements of each motion vector, so-called motion prediction flag, are transmitted and specified regardless of whether the base layer motion vector is used as a motion vector predictor. If the motion prediction flag is equal to 1, the motion vector predictor of the collocated reference blocks in the base layer is scaled according to the resolution ratio and used as a motion vector predictor. If the motion prediction flag is equal to 0, the motion vector predictor is calculated as defined in H.264 / AVC. In HEVC, motion parameters are predicted by applying advanced motion vector prediction (AMVP). AMVP features two competing spatial motion vector predictors and one temporal motion vector predictor. The spatial candidates are selected from the positions of adjacent prediction blocks located to the left or above the current prediction block. The temporal candidate is selected from among the collocated positions of the previously encoded pictures. The positions of all spatial and temporal candidates are shown in Figure 36.

[0248] TIFF0007682959000013.tif114170

[0249] TIFF0007682959000014.tif43170

[0250] The following section describes a method for using multiple enhancement layer predictors that include predictors obtained from the base layer to encode motion parameters of the enhancement layer. Motion information that has already been encoded for the base layer can be used to significantly reduce the motion data rate while encoding the enhancement layer. This method includes the possibility of directly obtaining all motion data of the prediction block from the base layer, in which case additional motion data need not be encoded. In the following description, the term prediction block refers to a prediction unit within HEVC, an M×N block within H.264 / AVC, and can be understood as a general set of samples within an image.

[0251] The first part of the current section relates to extending the list of motion vector prediction candidates by the base layer motion vector predictor (see Example K). The base layer motion vectors are added to the motion vector predictor list during enhancement layer encoding. This is achieved by inferring one or multiple motion vector predictors of the collocated prediction blocks from the base layer and using them as candidates within the list of predictors for motion compensation prediction. The collocated prediction blocks of the base layer are located at the center, left, top, right, or bottom of the current block. If the predicted block of the base layer at the selected position does not contain motion-related data or is outside the current range and thus not currently accessible, alternative positions can be used to infer the motion vector predictor. These alternative positions are shown in Figure 38.

[0252] TIFF0007682959000015.tif149170

[0253] TIFF0007682959000016.tif155170

[0254] TIFF0007682959000017.tif157169

[0255] TIFF0007682959000018.tif129169

[0256] The following relates to the enhancement layer coding of the conversion coefficients.

[0257] In the state-of-the-art video and image coding, the residual of the prediction signal is previously transformed and the resulting quantized conversion coefficients are signaled within the bitstream. This coefficient coding follows a fixed scheme.

[0258] Depending on the conversion size (for luma residuals: 4×4, 8×8, 16×16, and 32×32), different scanning directions are defined. Given the first and last positions in the scanning order, these scans uniquely determine which coefficient positions can be significant and, as a result, need to be coded. Within all scans, the last position has to be signaled within the bitstream, but the first coefficient is set to be the DC coefficient at position (0,0). The bitstream is done by coding the (horizontal) x and (vertical) y positions within the conversion block. Starting from the last position, the signaling of the significant coefficients is done in reverse scanning order until the DC position is reached.

[0259] For conversion sizes 16×16 and 32×32, only one scan, namely, the "diagonal scan", is defined. However, for conversion blocks of sizes 2×2, 4×4, and 8×8, "vertical" and "horizontal" scans can also be used. However, the use of vertical and horizontal scans is restricted to the residuals of the intra prediction coding units. And the scan actually used is obtained from the direction mode of that intra prediction. Direction modes with indices in the range of 6 to 14 result in a vertical scan, while direction modes with indices in the range of 22 to 30 result in a horizontal scan. All remaining direction modes result in a diagonal scan.

[0260] Figure 39 shows the diagonal scan, vertical scan, and horizontal scan defined for a 4×4 transform block. Larger transform coefficients are divided into subgroups of 16 coefficients. These subgroups enable hierarchical coding of significant coefficient positions. Subgroups signaled as non-significant do not contain significant coefficients. Figures 40 and 41 show the transforms for 8×8 and 16×16 respectively, along with the associated subgroup partitioning. The large arrows represent the scan order of the coefficient subgroups.

[0261] In the zigzag scan, for blocks of size larger than 4×4, the subgroups consist of 4×4 pixel blocks scanned in the zigzag scan. The subgroups are scanned in a zigzag manner. Figure 42 shows the vertical scan for 16×16 transform as proposed within JCTVC-G703.

[0262] The following paragraphs describe extensions for transform coefficient coding. These include the introduction of new scan modes, a method of assigning scans to transform blocks, and a modified coding of significant coefficient positions. These extensions enable better adaptation to different coefficient distributions within the transform block, and as a result, achieve coding gain in terms of rate distortion.

[0263] New realizations for vertical and horizontal scan patterns are introduced for 16×16 and 32×32 transform blocks. In contrast to previously proposed scan patterns, the size of the scan subgroups is 16×1 for horizontal scans and 1×16 for vertical scans respectively. Also, subgroups with sizes of 8×2 and 2×8 can be selected respectively. The subgroups themselves are scanned in the same way.

[0264] Vertical scan is effective for transform coefficients located within the horizontal spread. This can be found in images containing horizontal edges.

[0265] Horizontal scan is efficient for transform coefficients found within the vertical spread. This can be found in images containing vertical edges.

[0266] Figure 43 shows the realization of vertical and horizontal scanning for a 16×16 transform block. The coefficient subgroups are defined as one row or one column respectively. The vertical-horizontal scanning is the introduced scanning pattern. The scanning pattern enables the encoding of coefficients within a column by row-wise scanning. For a 4×4 block, the first column is scanned following the remainder of the first row, then the remainder of the second column is scanned, then the remainder of the coefficients of the second row is scanned. Next, the remainder of the third column is scanned, and finally the remainder of the fourth row and column is scanned.

[0267] For larger blocks, the block is divided into 4×4 subgroups. These 4×4 blocks are scanned by vertical-horizontal scanning, and the subgroups are scanned by the vertical-horizontal scanning itself.

[0268] The vertical-horizontal scanning can be used when the coefficients are located in the first row and column within the block. In this way, the coefficients are scanned earlier than when using another scanning, such as diagonal scanning. This can be found for an image including both horizontal edges and vertical edges.

[0269] Figure 44 shows the vertical and horizontal scanning for a 16×16 transform block.

[0270] Other scans are similarly possible. For example, all combinations between scans and subgroups can be used. For example, horizontal scanning for a 4×4 block and diagonal scanning for the subgroup are used. An appropriate selection of scans can be applied by selecting different scans for each subgroup.

[0271] It should be stated that various scans can be realized in the way that the transform coefficients are rearranged after quantization on the encoder side and conventional encoding is used. On the decoder side, the transform coefficients are rearranged after conventional decoding and before scaling and inverse transformation (or before inverse transformation after scaling).

[0272] Different parts of the base layer signal can be utilized to obtain coding parameters from the base layer signal. The base layer signal contains the following: · Collocated reconstructed base layer signals · Collocated residual base layer signals · Estimated residual signals of the enhancement layer obtained by subtracting the enhancement layer prediction signal from the reconstructed base layer signal · Image partitions of the base layer frame

[0273] Gradient parameters The gradient parameters are obtained as follows: For each pixel of the investigated block, a gradient is calculated. From these gradients, the magnitude and the angle are calculated. The angle that occurs most frequently within the block is the one associated with the block (block angle). The angle is rounded to use only three directions: horizontal (0°), vertical (90°), and diagonal (45°).

[0274] Edge detection The edge detector can be applied to the investigated blocks as follows: First, the block is smoothed by an n×n smoothing filter (e.g., Gaussian). A gradient matrix of size m×m is used to calculate the gradient of each pixel. The magnitude and the angle of every pixel are calculated. The angle is rounded and only three directions are used: horizontal (0°), vertical (90°), and diagonal (45°). For every pixel having a magnitude greater than a predetermined threshold 1, the adjacent pixels are checked. If an adjacent pixel has a magnitude greater than threshold 2 and has the same angle as the current pixel, the counter for this angle is incremented. For the entire block, the counter with the highest value is selected as the angle of the block.

[0275] Obtaining base layer coefficients by previous transformations For a particular TU, to obtain encoding parameters from the frequency domain of the base layer signal, the examined and collocated signals (reconstructed base layer signal / residual base layer signal / estimated enhancement layer signal) are transformed in the frequency domain. Preferably, this is performed using the same transform that is used by that particular enhancement layer TU. The resulting base layer transform coefficients may or may not be quantized. Rate-distortion quantization with a modified lambda can be used to obtain a coefficient distribution comparable to that of the enhancement layer block.

[0276] Scanning effectiveness score for a particular distribution and scan The scanning effectiveness score for a particular significant coefficient distribution can be defined as follows: Let each position of the examined block be represented by an index in the order of the examined scan. Then, the sum of the index values of the significant coefficient positions is defined as the effectiveness score of this scan. As a result, the smaller the score of the scan, the better the efficiency represented by the particular distribution.

[0277] Selection of a matching scan pattern for transform coefficient encoding If several scans are available for a particular TU, a rule for uniquely selecting one of the scans needs to be defined.

[0278] Method for scan pattern selection The selected scan can be obtained directly from the already decoded signal (without transmitting any additional data). This is possible either based on the characteristics of the collocated base layer signal or by using only the enhancement layer signal. The scan pattern can be obtained from the EL signal as follows. · The aforementioned state-of-the-art derivation rules. · Using the scan pattern for the chrominance residual selected for the collocated luminance residual. ·Define a fixed mapping between the symbolization mode and the used scanning pattern. ·Obtain the scanning pattern from the last significant coefficient position (in relation to the estimated fixed scanning pattern). ·In a preferred embodiment, the scanning pattern is selected depending on the last position already decoded as follows:

[0279] The last position is represented as x and y coordinates within the transform block and is already decoded (for the last encoding depending on the scan, a fixed scanning pattern is estimated for the decoding process of the last position. It can be the leading scanning pattern of that TU). Let T be a defined threshold depending on the specific transform size. If neither the x - coordinate nor the y - coordinate of the last significant position exceeds T, a diagonal scan is selected.

[0280] Otherwise, x is compared with y. If x exceeds y, a horizontal scan is selected and a vertical scan is not selected. The preferred value of T for a 4×4 TU is 1. The preferred value of T for a TU larger than 4×4 is 4.

[0281] In another preferred embodiment, the derivation of the scanning pattern described within the previous embodiment is restricted to only for TUs of sizes 16×16 and 32×32. It can be further restricted to only the luminance signal.

[0282] Also, the scanning pattern is obtained from the BL signal. Any of the encoding parameters described above can be used to obtain the scanning pattern selected from the base layer signal. In particular, the gradient of the collocated base layer signal can be calculated and compared with a predefined threshold and / or potentially discovered edges can be utilized.

[0283] In a preferred embodiment, the scanning direction is obtained depending on the block gradient angle as follows. For a gradient quantized in the horizontal direction, vertical scanning is used. For a gradient quantized in the vertical direction, horizontal scanning is used. Otherwise, diagonal scanning is selected.

[0284] In another preferred embodiment, the scanning pattern is obtained as described in the previous embodiment. However, it is obtained only for those transform blocks where the number of occurrences of the block angle exceeds a threshold. The remaining transform units are decoded using the leading scanning pattern of the TU.

[0285] If the base layer coefficients of the juxtaposed blocks are valid and are clearly signaled within the base layer data stream or are calculated by forward transformation, the base layer coefficients can be utilized in the following ways. · For each available scan, the cost for encoding the base layer coefficients can be evaluated. The scan with the lowest cost is used for decoding the enhancement layer coefficients. · The effective score for each available scan is calculated for the base layer coefficient distribution. The scan with the minimum score is used for decoding the enhancement layer coefficients. · The distribution of the base layer coefficients within the transform block is classified into one of a predefined set of distributions associated with a specific scanning pattern. · The scanning pattern is selected depending on the last significant base layer coefficient.

[0286] If the juxtaposed base layer blocks are predicted using intra prediction, the intra prediction direction can be used to obtain the enhancement layer scanning pattern.

[0287] Moreover, the transform size of the juxtaposed base layer blocks can be utilized to obtain the scanning pattern.

[0288] In a preferred embodiment, the scanning pattern is obtained from the BL signal only for the TUs representing the residuals of the INTRA_COPY mode prediction blocks. And their collocated base layer blocks are intra-predicted. For those blocks, a modified state-of-the-art scan selection is used. In contrast to the state-of-the-art scan selection, the intra-prediction direction of the collocated base layer blocks is used to select the scanning pattern.

[0289] Signaling of the scan pattern index in the bitstream (see Example R). The scan pattern of the transform block can be selected by the encoder in terms of rate distortion and then signaled in the bitstream.

[0290] A particular scan pattern can be encoded by signaling an index to a list of available scan pattern candidates. This list is either a fixed list of scan patterns defined for a particular transform size or can be actively filled during the decoding process. Actively filling the list allows for a suitable selection of those scan patterns. The scan pattern probably most efficiently encodes a particular coefficient distribution. By doing so, the number of available scan patterns for a particular TU can be reduced. And as a result, signaling the index into that list is less costly. If the number of scan patterns in a particular list is reduced to one, signaling is not necessary. For a particular TU, the process of selecting scan pattern candidates may utilize any of the encoding parameters described above and / or follow a predetermined rule that utilizes particular characteristics of that particular TU. Among them are the following. · The TU represents the residual of the luminance / chrominance signal. · The TU has a particular size. · The TU represents the residual of a particular prediction mode. · The last significant position within the TU is known by the decoder and belongs within a particular sub-division of the TU. ·TU is part of one I / B / P - slice. ·The coefficients of TU are quantized using specific quantization parameters.

[0291] In a preferred embodiment, the list of scan pattern candidates includes three scans for all TUs: "diagonal scan", "vertical scan" and "horizontal scan".

[0292] Another embodiment can be obtained by including any combination of scan patterns in the candidate list.

[0293] In a particular preferred embodiment, the list of scan pattern candidates can include any one of the "diagonal scan", "vertical scan" and "horizontal scan".

[0294] However, the scan pattern selected by the state - of - the - art scan derivation (as described above) is initially set to be within the list. Another candidate is added to the list only if a particular TU has a size of 16×16 or 32×32. The order of the remaining scan patterns depends on the last significant coefficient position.

[0295] (Note: The diagonal scan is always the first pattern in the list for estimating 16×16 and 32×32 conversions)

[0296] If the magnitude of the x - coordinate exceeds the magnitude of the y - coordinate, the horizontal scan is then selected. And the vertical scan is placed in the last position. Otherwise, the vertical scan is placed in the second position followed by the horizontal scan.

[0297] Another preferred embodiment is obtained by further restricting the conditions because there is one or more candidates in the list.

[0298] In another embodiment, if the coefficients of the transform block represent the residual of the luminance signal, the vertical and horizontal scans are only added to the candidate list of 16×16 and 32×32 transform blocks.

[0299] In another embodiment, if both the x and y coordinates of the last significant position are greater than a specific threshold, the vertical and horizontal scans are added to the candidate list of transform blocks. This threshold depends on the mode and / or TU size. The preferred threshold is 3 for all sizes larger than 4×4, and 1 for 4×4 TUs.

[0300] In another embodiment, if either the x or y coordinate of the last significant position is greater than a specific threshold, the vertical and horizontal scans are only added to the candidate list of transform blocks. This threshold depends on the mode and / or TU size. The preferred threshold is 3 for all sizes larger than 4×4, and 1 for 4×4 TUs.

[0301] In another embodiment, if both the x and y coordinates of the last significant position are greater than a specific threshold, the vertical and horizontal scans are only added to the candidate list of 16×16 and 32×32 transform blocks. This threshold depends on the mode and / or TU size. The preferred threshold is 3 for all sizes larger than 4×4, and 1 for 4×4 TUs.

[0302] In another embodiment, if either the x or y coordinate of the last significant position is greater than a specific threshold, the vertical and horizontal scans are only added to the candidate list of 16×16 and 32×32 transform blocks. This threshold depends on the mode and / or TU size. The preferred threshold is 3 for all sizes larger than 4×4, and 1 for 4×4 TUs.

[0303] For any of the described embodiments, a specific scan pattern is signaled within the bitstream. The signaling itself can be done at various signaling levels. In particular, the signaling can be done at any node of the residual quad-tree (all sub-TUs of that node, which use the signaled scan and the same candidate list index), at the CU / LCU level, or at the slice level, for each TU (which decreases within a subgroup of TUs having the signaled scan pattern).

[0304] The indexes within the candidate list can be transmitted using fixed-length coding, variable-length coding, arithmetic coding (including context-adaptive binary arithmetic coding), or PIPE coding. If context-adaptive coding is used, the context can be obtained based on adjacent blocks, the coding modes described above, and / or parameters of the characteristics of the specific TU itself.

[0305] In a preferred embodiment, context-adaptive coding is used to signal the indexes within the scan pattern candidate list of the TU. However, the context model is obtained based on the transform size and / or position of the last significant position within the TU. Any of the methods described above for obtaining the scan pattern can also be used to obtain a context model for signaling the explicit scan pattern for a specific TU.

[0306] For coding the last significant scan position, the following enhancements can be used within the enhancement layer. · Separate context models are used for all or a subset of the coding modes using the base layer information. It is also possible to use different context models for different modes having the base layer information. · The context model can depend on the data within the juxtaposed base layer blocks (e.g., the transform coefficient distribution within the base layer, the gradient information of the base layer, the last cancel position within the juxtaposed base layer blocks). · The last scan position can be coded as the difference from the last base layer scan position. · If the last scan position is coded by signaling its x and y positions within the TU, the context model of the second signaled coordinate can depend on the value of the first signaling. · To obtain a scan pattern independent of the last significant position, any of the above methods can be used both to signal to the last significant position and to obtain a context model.

[0307] In a particular version, scan pattern derivation depends on the last significant position: · If the last scan position is coded by signaling its x and y positions within the TU, the context model of the second coordinate can depend on those scan patterns that are still possible candidates when the first coordinate is already known. · If the last scan position is coded by signaling its x and y positions within the TU, the context model of the second coordinate can depend on whether the scan pattern has already been uniquely selected when the first coordinate is already known.

[0308] In another version, scan pattern derivation is independent of the last significant position. · The context model can depend on the scan pattern used within a particular TU. · Any of the methods described above to obtain the scan pattern can be used both to signal to the last significant position and to obtain a context model.

[0309] To code the significant positions and significant flags (subgroup flags and / or significant flags for one transform coefficient) within the TU, respectively, the following changes can be used within the enhancement layer: · Separate context models are used for all or a subset of the coding modes that use base layer information. It is also possible to use different context models for different modes with base layer information. · The context model can depend on data within juxtaposed base layer blocks (e.g., the number of significant transform coefficients for a particular frequency position). · Any of the methods described above for obtaining a scan pattern can be used to obtain a context model for signaling significant positions and / or their levels. · A generalized template can be used that evaluates both the significant number of already - encoded transform coefficient levels within the spatial neighborhood of the coefficient to be encoded and the number of significant transform coefficients within the juxtaposed base layer signals at similar frequency positions. · A generalized template can be used that evaluates both the significant number of already - encoded transform coefficient levels within the spatial neighborhood of the coefficient to be encoded and the levels of the significant transform coefficients within the juxtaposed base layer signals at similar frequency positions. · Context modeling for subgroup flags can depend on the scan pattern used and / or the particular transform size.

[0310] Different usage methods of context initialization tables for the base layer and the enhancement layer can be used. The initialization of the context model for the enhancement layer can be changed in the following ways. · The enhancement layer uses separate sets of initialization values. · The enhancement layer uses separate sets of initialization values for different operation modes (spatial / temporal or quality scalability). · An enhancement layer context model having corresponding parts within the base layer can use the states of those corresponding parts as the initialization state. · An algorithm for obtaining the initial state of the context model can depend on the base layer QP and / or the delta QP.

[0311] Next, using the base layer data, the possibility of encoding the appropriate subsequent enhancement layer is explained. The following part explains a method of generating an enhancement layer prediction signal within a scalable video coding system. The method uses a decoded base layer of image sample information to infer the values of prediction parameters. The values of the prediction parameters are not transmitted within the encoded video bitstream, but are used to form a prediction signal for the enhancement layer. Accordingly, the overall bitrate required to encode the enhancement layer signal is reduced.

[0312] State-of-the-art hybrid video encoders typically decompose the original (source) image into blocks of different sizes according to a hierarchical structure. For each block, the video signal is predicted from spatially adjacent blocks (intra prediction) or from previously temporally encoded images (inter prediction). The difference between the prediction and the actual image is transformation and quantization. The resulting prediction parameters and transformation coefficients are entropy encoded to form an encoded video bitstream. A compliant decoder follows the steps in reverse order... A scalable video encoding the bitstream is composed of different layers: a base layer that provides a fully decodable video and an enhancement layer that can be additionally used for decoding. The enhancement layer can provide higher spatial resolution (spatial scalability), temporal resolution (temporal scalability) or quality (SNR scalability). In previous standards such as H.264 / AVC SVC, syntax elements such as motion vectors, reference image indices or intra prediction modes are directly predicted from the corresponding syntax elements within the encoded base layer. Within the enhancement layer, a mechanism exists to switch between them using a prediction signal obtained from the base layer syntax elements at the block level, or predicted from another enhancement layer syntax element or decoded enhancement layer samples.

[0313] In the following part, the base layer data is used to obtain the enhancement layer parameters on the decoder side.

[0314] Method 1: Motion parameter candidate derivation For a block (a) of the spatial or quality enhancement layer image, the corresponding block (b) of the base layer image is determined. It covers the same image area. The inter-prediction signal for the block (a) of the enhancement layer is formed using the following method: 1. A set of motion compensation parameter candidates is determined, for example, from temporally or spatially adjacent enhancement layer blocks or derivatives thereof. 2. Motion compensation is performed for each set of motion compensation parameters of each candidate to form an inter-prediction signal within the enhancement layer. 3. The best set of motion compensation parameters is selected by minimizing the magnitude of the error between the prediction signal for the enhancement layer block (a) and the reconstruction signal of the base layer block (b). For spatial scalability, the base layer block (b) can be spatially upsampled using an interpolation filter.

[0315] A set of motion compensation parameters includes a specific combination of motion compensation parameters.

[0316] The motion compensation parameters can be a motion vector, a reference image index, a selection between one and two predictions and another parameter.

[0317] In an alternative embodiment, candidates for the motion compensation parameter set are used from the base layer block. Also, inter prediction is performed within the base layer (using the base layer reference picture). To apply the error magnitude, the base layer block (b) reconstruction signal can be used directly without being upsampled. The selected optimal motion compensation parameter set is applied to the enhancement layer reference picture to form the prediction signal for block (a). When motion vectors are applied within the spatial enhancement layer, the motion vectors are scaled according to the resolution change. Both the encoder and the decoder can perform the same prediction steps to select an optimal motion compensation parameter set within the available candidates and create the same prediction signal. These parameters are not signaled within the encoded video bitstream.

[0318] The selection of the prediction method is signaled within the bitstream and can be encoded using entropy coding. Within a hierarchical block sub-division structure, this coding method can be selected at any sub-level or alternatively only for a subset of the coding hierarchy. In an alternative embodiment, the encoder can send an improved motion parameter set prediction signal to the decoder. The improved signal includes differentially encoded values of the motion parameters. The improved signal can be entropy encoded.

[0319] In an alternative embodiment, the decoder generates a list of best candidates. The index of the motion parameter set used is signaled within the encoded video bitstream. The index can be entropy encoded. In an example, the list can be ordered by increasing error magnitude.

[0320] The example uses the adaptive motion vector prediction (AMVP) candidate list of HEVC to generate candidates for the motion compensation parameter set. Another example uses the merge mode candidate list of HEVC to generate candidates for the motion compensation parameter set.

[0321] Method 2: Motion Vector Derivation For a block (a) of an image in a spatial or quality enhancement layer, a corresponding block (b) of the image in the base layer covering the same image region is determined.

[0322] An inter-prediction signal for the block (a) of the enhancement layer is formed using the following method: 1. A motion vector predictor is selected. 2. An estimation of motion for a defined set of search positions is performed on a reference image of the enhancement layer. 3. For each search position, a magnitude of error is determined and a motion vector having the smallest error is selected. 4. A prediction signal for the block (a) is formed using the selected motion vector.

[0323] In an alternative embodiment, the search is performed on the reconstructed base layer signal. For spatial scalability, the selected motion vector is scaled according to the spatial resolution change before generating the prediction signal in step 4.

[0324] The search positions can be at full resolution or sub-pel resolution. Also, the search can be performed in multiple steps, for example, first determining the best full-pel position followed by another set of candidates based on the selected full-pel position. For example, the search can end early when the magnitude of error is below a defined threshold.

[0325] Both the encoder and decoder can perform the same prediction steps to select the optimal motion vector within the candidates to generate the same prediction signal. These vectors are not signaled in the encoded video bitstream.

[0326] The selection of the prediction method can be signaled within the bitstream and encoded using entropy coding. Within a hierarchical block sub-division structure, this coding method can be selected within any sub-level or only for a subset of alternative coding levels. In an alternative embodiment, the encoder can send an improved motion vector prediction signal to the decoder. The improved signal can be entropy coded.

[0327] An implementation example uses the algorithm described in Method 1 to select a motion vector predictor.

[0328] Another implementation example uses the Adaptive Motion Vector Prediction (AMVP) method of HEVC to select a motion vector predictor from temporally or spatially adjacent blocks in the enhancement layer.

[0329] Method 3: Intra Prediction Mode Derivation For each block (a) in the enhancement layer (n) image, a corresponding block (b) that covers the same region in the reconstructed base layer (n-1) image is determined.

[0330] In a scalable video decoder, for each base layer block (b), an intra prediction signal is formed using an intra prediction mode (p) inferred by the following algorithm. 1) The intra prediction signal is generated for each available intra prediction mode according to the rules for intra prediction in the enhancement layer, but using sample values from the base layer. 2) The best prediction mode (p best ) is determined by minimizing the magnitude of the error (e.g., the sum of absolute differences) between the intra prediction signal and the decoded base layer block (b). 3) The prediction (p best ) mode selected in step 2) is used to generate a prediction signal for the enhancement layer block (a) according to the intra prediction rules for the enhancement layer.

[0331] Both the encoder and the decoder can execute the same steps to select the best prediction mode (p best ) and form a consistent prediction signal. Therefore, the actual intra prediction mode (p best ) is not signaled in the encoded video bitstream.

[0332] The selection of the prediction method can be signaled in the bitstream and encoded using entropy coding. In a hierarchical block sub-division structure, this coding mode can be selected within any sub-level or, alternatively, only for a subset of the coding hierarchy. An alternative embodiment generates an intra prediction signal using samples from the enhancement layer in step 2). For a spatially scalable enhancement layer, the base layer can be upsampled using an interpolation filter to apply the magnitude of the error.

[0333] An alternative embodiment divides the enhancement layer block into a plurality of blocks of a smaller block size (a i ) (e.g., a 16×16 block (a) is divided into 16 4×4 blocks (a i )). The algorithms described above are applied to each sub-block (a i ) and the corresponding base layer block (b i ). After prediction of the block (a i ), residual coding is applied, and the result is used to predict the block (a i+1 ).

[0334] An alternative embodiment determines the predicted intra prediction mode (p i ) using the sample values around (b) or (b best ). For example, when a 4×4 block (a i ) of the spatial enhancement layer (n) has a corresponding 2×2 base layer block (b i ), the samples around (b i ) are used to determine the predicted intra prediction mode (p bestThe 4x4 block (c) used for the determination of i is used to form

[0335] In an alternative embodiment, the encoder can send an improved intra prediction direction signal to the decoder. For example, in a video codec such as HEVC, most intra prediction modes correspond to the angles at which the boundary pixels are used to form the prediction signal. The offset to the optimal mode can be sent as the difference to the prediction mode (determined as described above) p best ). The improved mode can be entropy coded.

[0336] Intra prediction modes are usually coded depending on their probabilities. In H.264 / AVC, one maximum likelihood mode is determined based on the modes used in the (spatial) neighborhood of the block. In the HEVC list, a list of maximum likelihood modes is created. These maximum likelihood modes are selected using symbols in the bitstream that are less than the total number of modes required. An alternative embodiment uses the predicted intra prediction mode (p best ) for the block (a) (determined as described in the above algorithm) as the maximum likelihood mode, or as a member of the list of maximum likelihood modes.

[0337] Method 4: Intra prediction using the boundary region In a scalable video decoder for forming an intra prediction signal for a block (a) (see FIG. 45) of a scalable or quality enhancement layer, the lines of samples (b) from the surrounding region of the same layer are used to be filled within the block region. These samples are obtained from the already encoded region (usually, however, not necessary on the upper and left boundaries).

[0338] The following alternative variations for selecting these pixels can be used. a) If the pixels in the surrounding region are not yet encoded, the pixel values are not used to predict the current block. b) If the pixels in the surrounding area have not yet been encoded, the pixel values are obtained from the adjacent pixels that have already been encoded (e.g., by iteration). c) If the pixels in the surrounding area have not yet been encoded, the pixel values are obtained from the pixels in the corresponding area of the decoded base layer image.

[0339] To form the intra prediction of block (a), the adjacent lines of the pixels (b) obtained as described above are used as a template to fill within each line (a j ) of block (a).

[0340] The lines (a j ) of block (a) are filled stepwise along the x-axis. To achieve the best possible prediction signal, the columns of the template samples (b) are shifted along the y-axis to form the prediction signal (b´ j ) for the relevant line (a j ).

[0341] To find the optimal prediction within each line, the shift offset (o j ) is determined by minimizing the magnitude of the error between the resulting prediction signal (a j ) and the sample values of the corresponding line in the base layer.

[0342] If (o j ) is a non-integer value, an interpolation filter is used to map the value of (b) to the integer sample position of (a 7 ), as shown within (b´ j ).

[0343] If spatial scalability is used, the interpolation filter can be used to create a matching number of sample values of the corresponding line in the base layer.

[0344] The filling direction (x-axis) can be horizontal (left - right), vertical (up - down), diagonal, or any other angle. The samples used for the template line (b) are the samples directly adjacent to the block along the x-axis. The template line (b) is shifted along the y-axis which forms an angle of 90° with respect to the x-axis.

[0345] To find the optimal direction of the x-axis, a filling intra prediction signal is generated for the block (a). An angle with the minimum error magnitude between the prediction signal and the corresponding base layer block is selected. The number of possible angles can be limited.

[0346] Both the encoder and the decoder execute the same algorithm to determine the best prediction angle and offset. No explicit angle information or offset information needs to be signaled within the bitstream. In an alternative embodiment, only the samples of the base layer image are used to determine the offset (o i ).

[0347] In an alternative embodiment, an improvement of the predicted offset (o i ), e.g., a difference value, is signaled within the bitstream. Entropy coding can be used to code the improved offset value.

[0348] In an alternative embodiment, an improvement of the predicted direction, e.g., a difference value, is signaled within the bitstream. Entropy coding can be used to code the improved direction value.

[0349] If a line (b´ j ) is used for prediction, an alternative embodiment uses a threshold for selection. If the error magnitude for the optimal offset (o j ) is less than the threshold, the line (c i ) is used to determine the value of the block line (a j ). If the optimal offset (oj ) If the magnitude of the error for is greater than or equal to the threshold value, the (upsampled) base layer signal is used to determine the value of block line a j ).

[0350] Method 5: Another prediction parameter Another prediction information is inferred in the same manner as in Methods 1 to 3. For example, a block is divided into sub-blocks.

[0351] For a block (a) of a spatial or quality enhancement layer image, a corresponding block (b) of the base layer image is determined, which covers the same image area.

[0352] A prediction signal for block (a) of the enhancement layer is formed using the following method. 1) The prediction signal is generated for each possible value of the tested parameter. 2) The best prediction mode p best ) is determined by minimizing the error magnitude (e.g., the sum of absolute differences) between the prediction signal and the decoded base layer block (b). 3) The prediction (p best ) mode selected in step 2) is used to generate the prediction signal for block (a) of the enhancement layer.

[0353] Both the encoder and the decoder can execute the same prediction steps to select the optimal prediction mode within the possible candidates and generate the same prediction signal. The actual prediction mode is not signaled in the encoded video bitstream.

[0354] The selection of the prediction method can be signaled in the bitstream and encoded using entropy coding. In a hierarchical block sub-division structure, this coding method can be alternatively selected within any sub-level or only for a subset of the coding hierarchy.

[0355] The following description briefly summarizes some of the above embodiments.

[0356] Enhancement layer coding with multiple methods for generating an intra prediction signal using samples of a reconstructed base layer Main example: For encoding blocks within an enhancement layer, multiple methods for generating an intra prediction signal using samples of a reconstructed base layer are provided in addition to a method for generating a prediction signal based only on samples of a reconstructed enhancement layer.

[0357] Sub - example: · The multiple methods include the following. The reconstructed base layer signal (upsampled / filtered) is used directly as the enhancement layer prediction signal. · The multiple methods include the following. The reconstructed base layer signal (upsampled / filtered) is combined with a spatial intra prediction signal. There, the spatial intra prediction is obtained based on differential samples for adjacent blocks. The differential samples represent the difference between the reconstructed enhancement layer signal and the reconstructed base layer signal (upsampled / filtered) (see Example A). · The multiple methods include the following. A conventional spatial intra prediction signal (obtained using samples of an adjacent reconstructed enhancement layer) is combined with an (upsampled / filtered) base layer residual signal (the inverse transform of the base layer transform coefficients, or the difference between the base layer reconstruction and the base layer prediction) (see Example B). · The multiple methods include the following. The reconstructed base layer signal (upsampled / filtered) is combined with a spatial intra prediction signal. There, the spatial intra prediction is obtained based on samples of a reconstructed enhancement layer of adjacent blocks. The final prediction signal is obtained by weighting the spatial prediction signal and the base layer prediction signal in a way that different frequency components use different weightings (see Example C1). This can be achieved, for example, by any of the following. ○ Filter the base layer prediction signal by a low-pass filter, filter the spatial intra prediction signal by a high-pass filter, and add the signals filtered thereby (see Embodiment C2). ○ Convert the base layer prediction signal and the enhancement layer prediction signal, and overlay the resulting conversion blocks. Different weighting coefficients are used for different frequency positions therein (see Embodiment C3). The resulting conversion blocks can be inverse-transformed and used as the enhancement layer prediction signal. Alternatively, the resulting conversion coefficients are added to the scaled transmitted conversion coefficient levels and then inverse-transformed to obtain a block reconfigured prior to deblocking and in-loop processing (see Embodiment C4). · The following versions can be used for the method of using the reconstructed base layer signal. This can be fixed, or it can be signaled at the sequence level, picture level, slice level, maximum coding unit level, coding unit level. Alternatively, it can be created depending on other coding parameters. ○ Samples of the reconstructed base layer before deblocking and further in-loop processing (such as sample adaptive offset filters or adaptive loop filters). ○ Samples of the reconstructed base layer after deblocking but before further in-loop processing (such as sample adaptive offset filters or adaptive loop filters). ○ Samples of the reconstructed base layer after deblocking and in-loop processing (such as sample adaptive offset filters or adaptive loop filters), or samples of the reconstructed base layer during multiple in-loop processing steps (see Embodiment D). · Multiple versions of a method using the (upsampled / filtered) base layer signal are used. The upsampled / filtered base layer signals employed for these versions can differ within the interpolation filter used (including an interpolation filter that filters integer sample positions). Alternatively, the upsampled / filtered base layer signal for the second version is obtained by filtering the upsampled / filtered base layer signal for the first version. One selection of the different versions can be signaled at the slice level, picture level, slice level, maximum coding unit level, coding unit level. It can be inferred from the characteristics of the corresponding reconstructed base layer signal or the transmitted coding parameters (see Example E). · Different filters can be used to upsample / filter the reconstructed base layer signal (see Example E) and the base layer residual signal (see Example F). · For a base layer block where the residual signal is zero, it is replaced by another signal obtained from the base layer, e.g., a version filtered through a high-pass filter of the reconstructed base layer block (see Example G). · For a mode using spatial intra prediction, adjacent samples not available within the enhancement layer (due to a particular coding order) can be replaced by the corresponding samples of the upsampled / filtered base layer signal (see Example H). · For a mode using spatial intra prediction, the coding of the intra prediction mode can be changed. The list of most likely modes includes the intra prediction modes of the collocated base layer signal. ·In a specific version, the image of the enhancement layer is decoded in a two-stage process. In the first stage, only the blocks that use only the base layer signal for prediction (without using adjacent blocks) or the inter-prediction signal are decoded and reconstructed. In the second stage, the residual blocks that use adjacent samples for prediction are reconstructed. For the blocks reconstructed in the second stage, the spatial intra-prediction concept can be extended. (Refer to Example I) Based on the usefulness of the already reconstructed blocks, not only the samples adjacent to the upper side and the left side of the current block but also the samples adjacent to the lower side and the right side can be used for spatial intra-prediction.

[0358] Enhancement layer coding with multiple methods for generating an inter-prediction signal using the reconstructed base layer samples Main example: For encoding blocks in the enhancement layer, a multiple method for generating an inter-prediction signal using the samples of the reconstructed base layer is provided in addition to the method of generating a prediction signal based only on the samples of the reconstructed enhancement layer.

[0359] Sub-example: ·The multiple method includes the following methods. The conventional inter-prediction signal (obtained by motion-compensated interpolation of the already reconstructed image of the enhancement layer) is combined with the (upsampled / filtered) base layer residual signal (the inverse transform of the base layer transform coefficients or the difference between the base layer reconstruction and the base layer prediction). ·The multiple method includes the following methods. The reconstructed base layer signal (upsampled / filtered) is combined with the motion-compensated prediction signal. There, the motion-compensated prediction signal is obtained by the motion-compensated difference image. The difference image represents the difference between the reconstructed enhancement layer signal and the (upsampled / filtered) reconstructed base layer signal with respect to the reference image (refer to Example J). · The multiple method includes the following methods. The reconstructed base layer signal (upsampled / filtered) is combined with the inter-prediction signal. Therein, the inter-prediction is obtained by motion-compensated prediction using the reconstructed enhancement layer image. The final prediction signal is obtained by weighting the inter-prediction signal and the base layer prediction signal in a way that different frequency components use different weightings (see Example C). This can be achieved, for example, by any of the following. ○ Filter the base layer prediction signal with a low-pass filter, filter the inter-prediction signal with a high-pass filter, and add the resulting filtered signals. ○ Transform the base layer prediction signal and the inter-prediction signal, and overlay the resulting transformed blocks. Therein, different weighting coefficients are used for different frequency positions. The resulting transformed blocks are inverse-transformed to obtain a reconstructed block before deblocking and in-loop processing and can be used as the enhancement layer prediction signal, or the resulting transform coefficients are added to the scaled transmitted transform coefficient levels and then inverse-transformed. · For the method using the reconstructed base layer signal, the following versions can be used. This can be fixed, or it can be signaled at the sequence level, picture level, slice level, maximum coding unit level, coding unit level. Or it can be created depending on other coding parameters. ○ Samples of the reconstructed base layer before deblocking and further in-loop processing (such as sample adaptive offset filter or adaptive loop filter). ○ Samples of the reconstructed base layer after deblocking and before further in-loop processing (such as sample adaptive offset filter or adaptive loop filter). ○ Samples of the reconstructed base layer after deblocking and in-loop processing (such as sample adaptive offset filter or adaptive loop filter), or samples of the reconstructed base layer during multiple in-loop processing steps (see Example D). ·For a base layer block where the residual signal is zero, it is replaced by another signal obtained from the base layer, e.g., a version passed through a high-pass filter of the reconstructed base layer block (see Example G). ·Multiple versions of the method of using the (upsampled / filtered) base layer signal can be used. The upsampled / filtered base layer signals employed for these versions can differ within the interpolation filter used (including the interpolation filter that filters integer sample positions). Or, the upsampled / filtered base layer signal for the second version is obtained by filtering the upsampled / filtered base layer signal for the first version. One selection of the different versions is signaled at the sequence level, picture level, slice level, maximum coding unit level, coding unit level. It can be inferred from the corresponding reconstructed base layer signal or the characteristics of the transmitted coding parameters (see Example E). ·Different filters can be used to upsample / filter the reconstructed base layer signal (see Example E) and the base layer residual signal (see Example F). ·For motion-compensated prediction of the difference image (the difference between the enhancement layer reconstruction and the upsampled / filtered base layer signal) (see Example J), different interpolation filters can be used than for motion-compensated prediction of the reconstructed image. ·For motion-compensated prediction of the difference image (the difference between the enhancement layer reconstruction and the upsampled / filtered base layer signal) (see Example J), the interpolation filter is selected based on the characteristics of the corresponding region within the difference image (or based on coding parameters or information transmitted in the bitstream).

[0360] Enhancement layer motion parameter coding Main embodiment: Use of a plurality of enhancement layer predictors and at least one predictor obtained from the base layer for encoding enhancement layer motion parameters.

[0361] Sub - embodiment: · Add the (scaled) base layer motion vector to the motion vector predictor list (see Example K). ○ Use of the base layer block covering the collocated samples at the center position of the current block (possible alternative derivation). ○ Scale motion vectors according to the resolution ratio. · Add the motion data of the collocated base layer blocks to the merge candidate list (see Example K). ○ Use of the base layer block covering the collocated samples at the center position of the current block (possible alternative derivation). ○ Scale motion vectors according to the resolution ratio. ○ If, within the base layer, the merge_flag is equal to 1, do not add. · Re - ordering of the merge candidate list based on base layer merge information (see Example L) ○ If the collocated base layer block is merged to a specific candidate, the corresponding enhancement layer candidate is used as the first entry in the enhancement layer merge candidate list. · Re - ordering of the motion predictor candidate list based on base layer motion predictor information (see Example L) ○ If the collocated base layer block uses a specific motion vector predictor, the corresponding enhancement layer motion vector predictor is used as the first entry in the enhancement layer motion vector predictor candidate list. ·Derivation of the merge index (i.e., candidates for which the current block is to be merged) is based on the base layer information within the collocated blocks (see Example M). As an example, if the base layer block is merged into a particular adjacent block and it is signaled within a bitstream where the enhancement layer block is also merged, then the merge index is not transmitted at all. Instead, the enhancement layer block is merged into the same adjacent blocks (but within the enhancement layer) as the collocated base layer block.

[0362] Enhancement Layer Partitioning and Motion Parameter Inference Main Example: Inference of enhancement layer partitioning and motion parameters based on base layer partitioning and motion parameters (perhaps this example needs to be combined with any of the sub-examples).

[0363] Sub-Examples: ·Obtain motion parameters for N×M sub-blocks of the enhancement layer based on collocated base layer motion data. Group blocks having the same obtained parameters (or parameters with a small difference) into larger blocks. Determine the prediction and coding units. (See Example T) ·The motion parameters may include the number of motion hypotheses, reference index lists, motion vectors, motion vector predictor identifiers, and merge identifiers. ·Signal one of multiple methods for generating the enhancement layer prediction signal. Such methods can include the following. ○ Motion compensation using the obtained motion parameters and the reconstructed reference image of the enhancement layer. ○Combining (a) the (upsampled / filtered) base layer reconstruction for the current image, (b) the motion compensation signal using the obtained motion parameters, and the reference image of the enhancement layer generated by subtracting the (upsampled / filtered) base layer reconstruction from the reconstructed enhancement layer image. ○Combining (a) the (upsampled / filtered) base layer residual (the difference between the reconstructed signal and the prediction, or the inverse transform of the encoded transform coefficient values) for the current image, (b) the motion compensation signal using the obtained motion parameters, and the reference image of the reconstructed enhancement layer. ·If the juxtaposed blocks within the base layer are intra-coded, then the corresponding enhancement layer M×N blocks (or CUs) are also intra-coded. Therein, the intra-prediction signal is obtained using the base layer information (see Example U). For example, ○The (upsampled / filtered) version of the corresponding base layer reconstruction is used as the intra-prediction signal (see Example U). ○The intra-prediction mode is obtained based on the intra-prediction mode used within the base layer. And this intra-prediction mode is used for spatial intra-prediction within the enhancement layer. ·If the juxtaposed base layer blocks for the M×N enhancement layer blocks (sub-blocks) are merged with previously encoded base layer blocks (or have the same motion parameters), then the M×N enhancement layer (sub-) blocks are also merged with the enhancement layer blocks corresponding to the base layer blocks used for merging within the base layer (i.e., the motion parameters are copied from the corresponding enhancement layer blocks) (see Example M).

[0364] Coding / Context Modeling of Transform Coefficient Levels Main embodiment: Encoding the transform coefficients using different scanning patterns. Modeling the context for the enhancement layer based on the encoding mode and / or base layer data, and performing different initializations for the context mode.

[0365] Sub - embodiment: · Introducing one or more additional scanning patterns, e.g., horizontal and vertical scanning patterns. Redefining sub - blocks for the additional scanning patterns. Instead of 4×4 sub - blocks, for example, 16×1 or 1×16 sub - blocks can be used. Or 8×2 or 2×8 sub - blocks can be used. The additional scanning pattern can be introduced only for blocks larger than or equal to a specific size, e.g., 8×8 or 16×16 (see Embodiment V). · (If the encoded block flag is equal to 1,) the selected scanning pattern is signaled in the bit - stream (see Embodiment N). A fixed context can be used to signal the corresponding syntax element. Or the context derivation for the corresponding syntax element can depend on any of the following. ○ The gradient of the juxtaposed reconstructed base layer signal or the reconstructed base layer residual. Or the edges detected in the base layer signal. ○ The transform coefficient distribution within the juxtaposed base layer blocks. · The selected scan can be obtained directly from the base layer signal (without transmitting any additional data) based on the characteristics of the juxtaposed base layer signal (see Embodiment N). ○ The gradient of the juxtaposed reconstructed base layer signal or the reconstructed base layer residual. Or the edges detected in the base layer signal. ○ The transform coefficient distribution within the juxtaposed base layer blocks. · Different scans can be realized in a way that the transform coefficients are reordered after quantization on the encoder side and the conventional encoding is used. On the decoder side, the transform coefficients are decoded as usual and reordered before scaling and inverse transform (or before scaling and inverse transform after scaling). · To encode important flags (subgroup flags and / or important flags for a single conversion factor), the following changes can be used within the enhancement layer. ○ The separate context model is used for all or a subset of the encoding modes that use base layer information. It is also possible to use different context models for different modes with base layer information. ○ Context modeling can depend on the data of juxtaposed base layer blocks (e.g., the number of important conversion factors at a specific frequency position) (see Example O). ○ A generalized template that evaluates both the number of already encoded important conversion factor levels within the spatial neighborhood of the coefficients to be encoded and the number of important conversion factors in the juxtaposed base layer signals at similar frequency positions can be used (see Example O). · To encode the last important scan position, the following changes can be used within the enhancement layer. ○ The separate context model is used for all or a subset of the encoding modes that use base layer information. It is also possible to use different context models for different modes with base layer information (see Example P). ○ Context modeling can depend on the data within the juxtaposed base layer blocks (e.g., the conversion factor distribution in the base layer, the gradient information of the base layer, the last scan position in the juxtaposed base layer blocks). ○ The last scan position can be encoded as the difference relative to the last base layer scan position (see Example S). · How to use different context initialization tables for the base layer and the enhancement layer.

[0366] Encoding of the backward adaptation enhancement layer using base layer data Main example: Use of base layer data to obtain enhancement layer encoding parameters.

[0367] Sub - example: ·Obtaining merge candidates based on (potentially upsampled) base layer reconstruction. Within the enhancement layer, only the use of merge is signaled. However, in practice, the candidates used to merge the current block are obtained based on the reconstructed base layer signal. Therefore, for all merge candidates, the magnitude of the error between the (potentially upsampled) base layer signal for the current enhancement layer block and the corresponding predicted signal (obtained using motion parameters for the merge candidate) is evaluated for all merge candidates (or a subset thereof). And the merge candidate associated with the smallest error magnitude is selected. Also, the error magnitude is calculated within the base layer using the reconstructed base layer signal and the base layer reference picture (see Example Q). ·Obtaining motion vectors based on (potentially upsampled) base layer reconstruction. The motion vector difference is not encoded but is inferred based on the reconstructed base layer. For the current block, determine a motion vector predictor and evaluate a set of search positions defined around the motion vector predictor. For each search position, determine the magnitude of the error between the (potentially upsampled) base layer signal for the current enhancement layer block and the replaced reference frame (the replacement is given by the search position). Select the search position / motion vector that results in the smallest error magnitude. The search is divided into several stages. For example, the best full pel search is performed first. Subsequently, a half pel search is performed around the best full pel vector. Subsequently, a quarter pel search is performed around the best full / half pel vector. Also, the search is performed within the base layer using the reconstructed base layer signal and the base layer reference picture. The found motion vector is then scaled according to the resolution change between the base layer and the enhancement layer (see Example Q). ·Obtaining an intra prediction mode based on a base layer reconstruction (which may be potentially upsampled). The intra prediction mode is not encoded but is inferred based on the reconstructed base layer. For each possible intra prediction mode (or a subset thereof), determining the magnitude of the error between the (potentially upsampled) base layer signal and the intra prediction signal for the current enhancement layer block (using the tested prediction mode). Selecting the prediction mode that results in the smallest error magnitude. Also, the calculation of the error magnitude can be done within the base layer using the base layer signal and the intra prediction signal reconstructed within the base layer. Further, implicitly, an intra block can be decomposed into 4×4 blocks (or another block size). And for each 4×4 block, a separate intra prediction mode can be determined (see Example Q). ·The intra prediction signal can be determined by aligning the reconstructed base layer signal in units of columns or rows of boundary samples. To obtain a shift between adjacent samples and the current line / row, the error magnitude is calculated between the shifted line / row of adjacent samples and the reconstructed base layer signal. And the shift that results in the smallest error magnitude is selected. As adjacent samples, (upsampled) base layer samples or enhancement layer samples can be used. Also, the error magnitude can be calculated directly within the base layer (see Example W). ·Using a backward adaptation technique for deriving other coding parameters such as block partitioning.

[0368] A more concise overview of the above embodiment is presented below. In particular, the above embodiment is described.

[0369] A1) A scalable video decoder reconstructs (80) base layer signals (200a, 200b, 200c) from an encoded data stream (6), reconstructs (60) an enhancement layer signal (360), The reconstruction (60) To obtain the inter-layer prediction signal (380), the reconstructed base layer signals (200a, 200b, 200c) are subjected to resolution or quality improvement (220), Calculate (260) the difference signal between the already reconstructed part (400a or 400b) of the enhancement layer signal and the inter-layer prediction signal (380), To obtain the spatial intra prediction signal, spatially predict (260) the difference signal from a second part (460) of the difference signal that is spatially adjacent to the first part (440, exemplified in FIG. 46) juxtaposed to the part of the enhancement layer signal (360) to be currently reconstructed and belongs to the already reconstructed part of the enhancement layer signal (360), To obtain the enhancement layer prediction signal (420), combine (260) the inter-layer prediction signal (380) and the spatial intra prediction signal, It is configured to include predictively reconstructing (320, 580, 340, 300, 280) the enhancement layer signal (360) using the enhancement layer prediction signal (420). According to Example A1, the base layer signal can be reconstructed by the base layer decoding stage 80 from the coded data stream 6 or sub-stream 6a, respectively, by a prediction method based on the aforementioned block having inverse transform decoding, as long as, for example, the base layer residual signal 640 / 480 is relevant. However, other alternative reconstructions are also possible. As far as the reconstruction of the enhancement layer signal 360 by the enhancement layer decoding stage 60 is concerned, the resolution or quality improvement that the reconstructed base layer signals 200a, 200b or 200c undergo means, for example, upsampling in the case of resolution improvement, or copying in the case of quality improvement, or tone mapping from n bits to m bits (m > n) in the case of bit depth improvement. The calculation of the difference signal can be done on a pixel-by-pixel basis. That is, pixels with the enhancement layer signal on one side and the prediction signal 380 on the other side juxtaposed are subtracted from each other. And this is done for each pixel position. Spatial prediction of differential signals can be done in some way such that within the encoded data stream 6, or within sub-stream 6b, it transmits intra prediction parameters such as the intra prediction direction, and along this intra prediction direction within the current portion of the enhancement layer signal, it copies / interpolates the already reconstructed pixels adjacent to the portion of the enhancement layer signal 360 to be currently reconstructed. The combination means addition, weighted sum or a more sophisticated combination, such as a combination that differently weights the contributions in the frequency domain. Predictive reconstruction of the enhancement layer signal 360 using the enhancement layer prediction signal 420 can, as shown in the figure, mean entropy decoding and inverse transformation of the enhancement layer residual signal 540, and the enhancement layer prediction signal 420 and the combination 340 of the latter.

[0370] B1) The scalable video decoder decodes (100) the base layer residual signal (480) from the encoded data stream (6), reconstructs (60) the enhancement layer signal (360), The reconstruction (60) subjects the reconstructed base layer residual signal (480) to an improvement in resolution or quality (220) in order to obtain an inter-layer residual prediction signal (380), spatially predicts (260) the portion of the enhancement layer signal (360) to be currently reconstructed from the already reconstructed portion of the enhancement layer signal (360) in order to obtain an intra prediction signal within the enhancement layer, combines (260) the inter-layer residual prediction signal and the intra prediction signal within the enhancement layer in order to obtain an enhancement layer prediction signal (420), and is configured to predictively reconstruct (340) the enhancement layer signal (360) using the enhancement layer prediction signal (420). Decoding of the base layer residual signal from the symbolized data stream can be performed using entropy decoding and inverse transformation, as shown in the figure. Further, the scalable video decoder optionally obtains the base layer prediction signal 660 and reconstructs the base layer signal itself by predictively decoding by combining this signal with the base layer residual signal 480. As just mentioned, this is merely optional. As far as the reconstruction of the enhancement layer signal is concerned, an improvement in resolution or quality can be performed as indicated above with respect to Example A). Also, as far as the spatial prediction of parts of the enhancement layer signal is concerned, this spatial prediction can be performed as exemplarily outlined in A) for different signals. Similar considerations are valid as far as combinations and predictive reconstructions are concerned. However, it is noted that the base layer residual signal 480 within Example B) is not restricted to be equal to the explicitly signaled version of the base layer residual signal 480. Rather, it is possible for the scalable video decoder to subtract any reconstructed base layer signal version 200 having the base layer prediction signal 660. As a result, a base layer residual signal 480 is obtained that deviates from what is explicitly signaled by the offset that stalls from the filter function such as filter 120 or 140. Also, the latter situation is valid for another example where the base layer residual signal is involved in inter-layer prediction.

[0371] C1) The scalable video decoder reconstructs (80) the base layer signal (200a, 200b, 200c) from the encoded data stream (6), reconstructs (60) the enhancement layer signal (360), The reconstruction (60) causes the reconstructed base layer signal (200) to undergo an improvement in resolution or quality (220) in order to obtain the inter-layer prediction signal (380), To obtain an enhancement layer prediction signal, a portion of the enhancement layer signal (360) that is to be currently reconstructed is spatially or temporally predicted (260) from the already reconstructed portion of the enhancement layer signal (360) (400a,b in the case of "spatial"; 400a,b,c in the case of "temporal"), To obtain an enhancement layer prediction signal (420), a weighted average of the inter-layer prediction signal and the enhancement layer intra-prediction signal (380) is formed (260) in the portion that is to be currently reconstructed, such that the weighting by which the inter-layer prediction signal and the enhancement layer intra-prediction signal (380) contribute to the enhancement layer prediction signal (420) varies across different spatial frequency components. It is configured to include predictively reconstructing (340) the enhancement layer signal (360) using the enhancement layer prediction signal (420).

[0372] C2) Here, forming the weighted average (260) includes filtering the inter-layer prediction signal (380) with a low-pass filter (260) and filtering the enhancement layer intra-prediction signal with a high-pass filter (260) in the portion that is to be currently reconstructed to obtain filtered signals, and then summing the obtained filtered signals. C3) Here, forming the weighted average (260) includes transforming the inter-layer prediction signal and the enhancement layer intra-prediction signal (260) in the portion that is to be currently reconstructed to obtain transformation coefficients, and then superimposing the obtained transformation coefficients using different weighting coefficients for different spatial frequency components to obtain superimposed transformation coefficients, and then inverse-transforming the superimposed transformation coefficients to obtain the enhancement layer prediction signal. C4) Here, using the enhancement layer prediction signal (420), the predictive reconstruction (320, 340) of the enhancement layer signal extracts the transform coefficient levels for the enhancement layer signal from the encoded data stream (6) (320), and performs the sum of the transform coefficients superimposed on the transform coefficient levels to obtain a transformed version of the enhancement layer signal (340), and subjects the transformed version of the enhancement layer signal to an inverse transform (i.e., the inverse transform T in the figure) to obtain the enhancement layer signal (360). -1 This includes being placed downstream of the adder 340, at least for its encoding mode. As far as the reconstruction of the base layer signal is concerned, the reference is generally made with respect to the figure and as described for Examples A) and B) above. The same applies to the improvement in resolution or quality mentioned in C), similar to spatial prediction. The temporal prediction mentioned in C) may involve the motion prediction parameters obtained from the encoded data stream 6 and the sub - stream 6a respectively for the prediction provider 160. The motion parameters can include motion vectors, reference frame indices. Alternatively, the motion parameters can include a combination of motion sub - division information and motion vectors for each sub - block of the currently reconstructed part. As described above, the formation of the weighted average can end within the spatial domain or the transform domain. Therefore, the addition in the adder 340 may be performed within the spatial domain or the transform domain. In the latter case, the inverse transformer 580 applies the inverse transform to the weighted average.

[0373] D1) A scalable video decoder reconstructs the base layer signals (200a, 200b, 200c) from the encoded data stream (6) (80), reconstructs the enhancement layer signal (380) (60), The reconstruction (60) causes the reconstructed base layer signal to undergo resolution or quality improvement (220) to obtain the inter - layer prediction signal (380). Predictively reconstruct (320, 340) the enhancement layer signal (360) using the inter-layer prediction signal (380), wherein the reconstruction (60) of the enhancement layer signal is performed such that the inter-layer prediction signal (380) evolves, and for different parts of the video represented on a scale by the base layer signal and the enhancement layer signal respectively, non-blocking and in-loop filtering Processing (140) is none (200a), or is controlled via side information in the encoded bitstream 360) from one or all (200b, 200c) of different ones.

[0374] As far as the reconstruction of the base layer signal is concerned, the reference is generally made to the figures and as for Examples A) and B) as described above. The same applies to the improvement of resolution or quality. The predictive reconstruction referred to in D) can, as described above, involve the prediction provider 160. And the predictive reconstruction involves spatially or temporally predicting (260) the part of the enhancement layer signal (360) to be currently reconstructed from the already reconstructed part of the enhancement layer signal (380) to obtain an in-enhancement layer prediction signal, and involves combining (260) the inter-layer prediction signal (380) and the in-enhancement layer prediction signal to obtain an enhancement layer prediction signal (420). The fact that the inter-layer prediction signal (380) evolves is controlled via side information in the encoded bitstream (360) from none (200a) of non-blocking, or from one or all (200b, 200c) of different ones, and that for different parts, the in-loop filter (140) is applied means the following. Of course, the base layer sub-stream 6a itself can (optionally) signal the use of different means, the use of only deblocking, or the use of only in-loop filtering, or the use of both deblocking and in-loop filtering, to provide the final base layer signal 600 so as to bypass all filters 120, 140. Even the filter transfer function can be signaled / changed by side information within the base layer sub-stream 6a. These changes (variations) The size defining the different parts where these are made can be defined by any of the aforementioned coding units, prediction blocks, or any other size. As a result, a scalable video decoder (encoding stage 80) applies these changes (variations) if, hypothetically, only the base layer signal is to be reconstructed. However, independently therefrom (i.e., independent of the side information just mentioned within the base layer signal 6a), the sub-stream 6b contains side information signaling new changes (variations) wherein a combination of filtering is used to bypass all filters 120, 140 for use in the predictive reconstruction of the enhancement signal, the use of only deblocking, or the use of only in-loop filtering, or the use of both deblocking and in-loop filtering. That is, even the filter transfer function is signaled / changed by side information within the sub-stream 6b. The size defining the different parts where these changes (variations) are made can be defined by any of the aforementioned coding units, or prediction blocks, or any other size, and this signaling can be different from the size used within the base layer signal 6a.

[0375] E1) A scalable video decoder reconstructs (80) base layer signals (200a, 200b, 200c) from an encoded data stream (6), reconstructs (60) an enhancement layer signal (360), the reconstruction (60) To obtain the inter-layer prediction signal (380), the reconstructed base layer signal is subjected to an improvement in resolution or quality (220), The enhancement layer signal (60) is predictively reconstructed (320, 340) using the inter-layer prediction signal (380), wherein the reconstruction (60) of the enhancement layer signal (360) is performed such that the inter-layer prediction signal evolves and is controlled by different filter transfer functions for the upsampling interpolation filter (220) for different portions of the video scaled by the base layer signal and the enhancement layer signal respectively via side information in the encoded bitstream (6) or depending on the signaling.

[0376] As far as the reconstruction of the base layer signal is concerned, the reference is made as previously described, generally with respect to the figures and as in Examples A) and B). The same applies to the improvement in resolution or quality. The predictive reconstruction mentioned can be related to the prediction provider 160 as described above. And the predictive reconstruction To obtain the enhancement layer intra prediction signal, a portion of the enhancement layer signal (360) to be currently reconstructed is spatially or temporally predicted (260) from the already reconstructed portion of the enhancement layer signal (360), The obtaining of the enhancement layer prediction signal (420) can involve combining (260) the inter-layer prediction signal (380) and the intra prediction signal within the enhancement layer. The fact that the inter-layer prediction signal evolves means that it is controlled by different filter transfer functions for the upsampling interpolation filter (220) for different portions of the following video means via side information in the encoded bitstream (6) or depending on the signaling. Of course, the base layer sub-stream 6a itself can (optionally) be signaled to use different means, only use non-blocking, or only use in-loop filtering, or use both non-blocking and in-loop filtering, so as to provide the final base layer signal 600 to bypass all filters 120, 140. Even the filter transfer function can be signaled / changed by the side information within the base layer sub-stream 6a. The sizes that define the different parts where these changes (variations) are made can be defined by any of the aforementioned coding units, prediction blocks, or any other size. As a result, if only the base layer signal is reconstructed, the scalable video decoder (encoding stage 80) applies these changes (variations) However, independently therefrom (i.e., independent of the side information just mentioned within the base layer signal 6a), the sub-stream 6b can include side information that additionally signals changes in the filter transfer function used in the improver 220 to obtain an improved signal 380. The sizes that define the different parts where these changes (variations) are made can be defined by any of the aforementioned coding units, or prediction blocks, or any other size, and can be different from the aforementioned size of the base layer signal 6a. (variations) As mentioned above, the changes to be used can be inferred signal-dependently from the base layer signal, or the base layer residual signal, or the coding parameters within the sub-stream 6a, regardless of the presence or absence of additional side information. (variations)

[0377] F1) The scalable video decoder decodes (100) the base layer residual signal (480) from the coded data stream, To obtain the inter-layer residual prediction signal (380), the enhanced layer signal (360) is reconstructed (60) by subjecting the reconstructed base layer residual signal (480) to an improvement in resolution or quality (220), and the enhanced layer signal (360) is predictively reconstructed (320, 340, and optionally 260) using the inter-layer residual prediction signal (380). Here, the reconstruction (60) of the enhanced layer signal (360) is performed such that the inter-layer residual prediction signal unfolds and is configured to be controlled from different filter transfer functions for different parts of the video scaled by the base layer signal and the enhanced layer signal respectively via side information in the encoded bitstream (6) or depending on signaling. As far as the reconstruction of the base layer residual signal is concerned, references are made as previously done, generally with respect to the figures and as for Example B). The same applies to the improvement in resolution or quality. The predictive reconstruction mentioned can involve the prediction provider 160 as described above. And the predictive reconstruction is To obtain the intra prediction signal of the enhanced layer, a part of the enhanced layer signal (360) to be currently reconstructed is spatially or temporally predicted (260) from the already reconstructed part of the enhanced layer signal (360). The enhanced layer residual signal is decoded (320) from the encoded data stream. To obtain the enhanced layer signal (360), it can involve combining the in-layer prediction signal, the inter-layer residual prediction signal (380), and the enhanced layer residual signal (involving 340 and 260). The fact that the inter-layer residual prediction signal unfolds means that it is controlled from different filter transfer functions for different parts of the following video means via side information ...

Claims

Claim 1 A scalable video decoder configured to reconstruct (80) base layer signals (200a, 200b, 200c) from an encoded data stream (6), and predictively reconstruct (60) an enhancement layer signal (360) by reconstruction, the reconstruction of the enhancement layer signal (360) comprising: subjecting the reconstructed base layer signal to resolution improvement or quality improvement (220) to obtain an inter-layer prediction signal (380); and predictively reconstructing (320, 340) the enhancement layer signal (360) using the inter-layer prediction signal (380), wherein the reconstruction of the enhancement layer signal (60) is performed such that the inter-layer prediction signal (380) is subjected to different ones of zero (200a), one, or all (200b, 200c) of non-blocking and filtering processes (140) for different parts of the video, which are scalably represented by the base layer signal and the enhancement layer signal, under control via side information in the encoded data stream (6), and when only the base layer signals (200a, 200b, 200c) are reconstructed from the encoded data stream (6), the scalable video decoder is configured to reconstruct (80) the base layer signals (200a, 200b, 200c) from the encoded data stream (6) using different ones of zero (200a), one, or all (200b, 200c) of non-blocking and filtering processes (140) for different parts of the base layer signals (200a, 200b, 200c) under control via first side information in a base layer sub-stream (6a) of the encoded data stream (6), and when an enhancement layer signal (360) is reconstructed from the encoded data stream (6), the scalable video decoder is configured to reconstruct (80) the base layer signals (200a, 200b, 200c) from the encoded data stream (6) using different ones of zero (200a), one, or all (200b, 200c) of non-blocking and filtering processes (140) for different parts of the base layer signals (200a, 200b, 200c) under control via first side information in a base layer sub-stream (6a) of the encoded data stream (6), and when an enhancement layer signal (360) is reconstructed from the encoded data stream (6), The scalable video decoder, under the control via second side information in the enhancement layer sub-stream (6b) of the encoded data stream (6), for different parts of the base layer signals (200a, 200b, 200c), respectively, uses 0 (200a) or 1 or different ones of all (200b, 200c) of non-blocking and filtering (140) to reconstruct another version of the base layer signals (200a, 200b, 200c) from the encoded data stream (6), reconstructs (60) the enhancement layer signal (360) from the enhancement layer sub-stream (6b), is configured as wherein the reconstruction (60) of the enhancement layer signal (360) subjects the another version of the reconstructed base layer signal to resolution improvement or quality improvement (220) to obtain the inter-layer prediction signal (380), and uses the inter-layer prediction signal (380) to predictively reconstruct the enhancement layer signal (360) (320, 340) is configured to include is characterized in that the first side information is configured to signal a filter transfer function, and the second side information is configured to additionally signal variations of the filter transfer function used when subjecting the reconstructed base layer signal to resolution improvement or quality improvement (220) to obtain the inter-layer prediction signal (380), scalable video decoder. **Claim 2** The scalable video decoder according to claim 1, configured to reconstruct (80) the base layer signals (200a, 200b, 200c) predictively in a block-based manner from the encoded data stream (6) using inverse transform decoding of the base layer residual signals. **Claim 3** The scalable video decoder according to claim 1 or claim 2, configured to perform the resolution improvement or quality improvement by spatial upsampling. **Claim 4** The scalable video decoder according to any one of claims 1 to 3, configured to perform the resolution improvement or quality improvement by tone mapping from n bits to m bits where m > n to achieve an improvement in bit depth. **Claim 5** Predicting (260) spatially or temporally a portion of the enhancement layer signal (360) to be reconstructed now from a portion of the enhancement layer signal (360) already reconstructed to obtain an intra-enhancement layer prediction signal, and Combining (260) the inter-layer prediction signal (380) and the intra-enhancement layer prediction signal to obtain an enhancement layer prediction signal (420) A scalable video decoder according to any one of claims 1 to 4, configured to perform said predictive reconstruction thereby.

6. The scalable video decoder is configured to derive the inter-layer prediction signal by bypassing non-blocking and filtering processing (140), using both non-blocking and filtering processing (140), or using only non-blocking or only filtering processing and further responding to the side information regarding the filter transfer function, in response to the side information in the encoded data stream (6). A scalable video decoder according to any one of claims 1 to 5.

7. Modifying the transfer function of the non-blocking filtering process and / or the filtering process (140) depending on the first side information for the reconstruction (80) of the base layer signal (200a, 200b, 200c), and depending on the second side information for the reconstruction (80) of the other version of the base layer signal (200a, 200b, 200c), A scalable video decoder according to claim 1, configured to modify the transfer function of the non-blocking filtering process and / or the filtering process (140).

8. Reconstructing (80) a base layer signal (200a, 200b, 200c) from an encoded data stream (6), and Reconstructing (60) an enhancement layer signal (360) A scalable video decoding method comprising: The step (60) of reconstructing the enhancement layer signal (360) comprises: Subjecting the reconstructed base layer signal to resolution improvement or quality improvement (220) to obtain an inter-layer prediction signal (380), and Predictively reconstructing (320, 340) the enhancement layer signal (360) using the inter-layer prediction signal (380) Including The reconstruction (60) of the enhancement layer signal is characterized in that the inter-layer prediction signal (380) is subjected to non-blocking and filtering (140) for different parts of the video that are scalable represented by the base layer signal and the enhancement layer signal respectively under the control via the side information of the encoded bit stream (6), where different ones of 0 (200a) or 1 or all (200b, 200c) of non-blocking and filtering are applied, When only the base layer signal (200a, 200b, 200c) is reconstructed from the encoded data stream (6), the scalable video decoding method includes the step of reconstructing (80) the base layer signal (200a, 200b, 200c) from the encoded data stream (6) using different ones of 0 (200a) or 1 or all (200b, 200c) of non-blocking and filtering for different parts of the base layer signal (200a, 200b, 200c) under the control via the first side information in the base layer sub-stream (6a) of the encoded data stream (6), When reconstructing the enhancement layer signal (360) from the encoded data stream (6), the scalable video decoding method includes the step of reconstructing another version of the base layer signal (200a, 200b, 200c) from the encoded data stream (6) using different ones of 0 (200a) or 1 or all (200b, 200c) of the non-blocking and the filtering (140) for different parts of the base layer signal (200a, 200b, 200c) under the control via the second side information in the enhancement layer sub-stream (6b) of the encoded data stream (6), and the step of reconstructing (60) the enhancement layer signal (360) from the enhancement layer sub-stream (6b), wherein the step of reconstructing (60) the enhancement layer signal (360) includes the step of subjecting the another version of the reconstructed base layer signal to resolution improvement or quality improvement (220) to obtain the inter-layer prediction signal (380), and Predictively reconstructing the enhancement layer signal (360) using the inter-layer prediction signal (380) (steps 320, 340) comprising wherein the first side information signals a filter transfer function, and the second side information additionally signals a variation of the filter transfer function used when subjecting the reconstructed base layer signal to resolution improvement or quality improvement (220) to obtain an inter-layer prediction signal (380) A scalable video decoding method.

9. Encoding a base layer signal into an encoded data stream (6) Encoding an enhancement layer signal (360) A scalable video encoder configured to The encoding of the enhancement layer signal (360) comprises Subjecting the reconstructed base layer signal to resolution improvement or quality improvement to obtain an inter-layer prediction signal (380), and Predictively encoding the enhancement layer signal (360) using the inter-layer prediction signal (380) comprising wherein the encoding of the enhancement layer signal (60) is performed such that the inter-layer prediction signal (380) is subjected to different ones of non-blocking and filtering (140), namely, 0 (200a) or 1 or all (200b, 200c) for different parts of the video that are scalably represented by the base layer signal and the enhancement signal under the control via side information in the encoded data stream (6), respectively wherein the scalable video encoder Encodes the base layer signal into the encoded data stream (6) using a different one of non-blocking and filtering (140), namely, 0 (200a) or 1 or all (200b, 200c) for different parts of the base layer signal (200a, 200b, 200c) under signaling by first side information in the base layer sub-stream (6a) of the encoded data stream (6) Signaling second side information within the enhancement layer sub-stream (6b) of the encoded data stream (6), the second side information signaling a different one of zero (200a) or one or all (200b, 200c) of non-blocking and filtering processing (140) for different parts of the base layer signals (200a, 200b, 200c). For the different parts of the base layer signals (200a, 200b, 200c), obtaining a different version of the base layer signal using a different one of zero (200a) or one or all (200b, 200c) of non-blocking and filtering processing (140). Subjecting the different version of the base layer signal to resolution improvement or quality improvement (220) to obtain the inter-layer prediction signal (380). Predictively encoding the enhancement layer signal (360) using the inter-layer prediction signal (380) (320, 340). Thereby, it is characterized in that it is configured to encode the enhancement layer signal (360) into the encoded data stream (6). The first side information is configured to signal a filter transfer function, and the second side information is configured to additionally signal variations of the filter transfer function used when subjecting the reconstructed base layer signal to resolution improvement or quality improvement (220) to obtain the inter-layer prediction signal (380). Scalable video encoder. [

10. ] Encoding a base layer signal into an encoded data stream (6). Encoding an enhancement layer signal (360). A scalable video encoding method including: The step of encoding the enhancement layer signal includes Subjecting the reconstructed base layer signal to resolution improvement or quality improvement to obtain an inter-layer prediction signal (380). Predictively encoding the enhancement layer signal (360) using the inter-layer prediction signal (380). Including The step (60) of encoding the enhancement layer signal is characterized in that the inter-layer prediction signal (380) is non-blocked and filtered (140) for different parts of the video that are scalable represented by the base layer signal and the enhancement signal, respectively, under the control via side information in the encoded data stream (6), and is applied to different ones of 0 (200a) or 1 or all (200b, 200c) of them. The scalable video encoding method encoding the base layer signal (200a, 200b, 200c) into the encoded data stream (6) using different ones of 0 (200a) or 1 or all (200b, 200c) of non-blocking and filtering (140) for different parts of the base layer signal (200a, 200b, 200c) respectively, under signaling by first side information in the base layer sub-stream (6a) of the encoded data stream (6); signaling second side information in the enhancement layer sub-stream (6b) of the encoded data stream (6), the second side information signaling different ones of 0 (200a) or 1 or all (200b, 200c) of non-blocking and filtering (140) for different parts of the base layer signal (200a, 200b, 200c) respectively; obtaining another version of the base layer signal using different ones of 0 (200a) or 1 or all (200b, 200c) of non-blocking and filtering (140) for different parts of the base layer signal (200a, 200b, 200c) respectively; subjecting the another version of the base layer signal to resolution improvement or quality improvement (220) to obtain the inter-layer prediction signal (380); predictively encoding the enhancement layer signal (360) using the inter-layer prediction signal (380) (320, 340); thereby, the step (60) of encoding the enhancement layer signal (360) into the encoded data stream (6); is characterized by including. The first side information signals a filter transfer function, and the second side information additionally signals a variation of the filter transfer function used when subjecting the reconstructed base layer signal to resolution improvement or quality improvement (220) to obtain an interlayer prediction signal (380). Scalable video encoding method.

11. A computer program having the program code, wherein when the program code is executed on a computer, the computer executes the scalable video decoding method according to claim 8 or the scalable video encoding method according to claim 10.

Citation Information

Patent Citations

  • Quality-scalable encoding method

    JP2010507941A

  • Image encoding-decoding system and related techniques

    US20070223582A1