Scalable video coding using inter-layer prediction contribution to enhancement layer prediction

By combining inter-layer and intra-enhancement layer prediction with spatially varying weights and utilizing base layer information, the method enhances coding efficiency in scalable video coding, achieving a more accurate prediction signal and reduced signaling overhead.

JP2025102790AActive Publication Date: 2025-07-08DOLBY VIDEO COMPRESSION LLC
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2025035195
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2012-10-01
Filing Date
2025-03-06
Publication Date
2025-07-08
Estimated Expiration
2033-10-01

AI Technical Summary

Technical Problem

Existing scalable video coding techniques do not achieve optimal coding efficiency, particularly in utilizing inter-layer prediction for enhancement layers.

Method used

A method for scalable video coding that forms the enhancement layer prediction signal by combining inter-layer and intra-enhancement layer prediction signals with spatially varying weights based on different frequency components, and utilizes base layer information for improved motion compensation and sub-block division to enhance coding efficiency.

Benefits of technology

This approach increases the coding efficiency by providing a more accurate enhancement layer prediction signal, reducing the need for signaling overhead, and improving the compression ratio.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025102790000002
    Figure 2025102790000002
  • Figure 2025102790000003
    Figure 2025102790000003
  • Figure 2025102790000004
    Figure 2025102790000004
Patent Text Reader

Abstract

To provide a scalable video decoder that achieves a higher coding efficiency.SOLUTION: A scalable video decoder is configured to: reconstruct a base layer signal from a coded data stream to obtain a reconstructed base layer signal; reconstruct an enhancement layer signal comprising spatially or temporally predicting a portion of an enhancement layer signal, currently to be reconstructed, from an already reconstructed portion of the enhancement layer signal to obtain an enhancement layer internal prediction signal; form, at the portion currently to be reconstructed, a weighted average of an inter-layer prediction signal obtained from the reconstructed base layer signal, and the enhancement layer internal prediction signal to obtain an enhancement layer prediction signal such that a weighting between the inter-layer prediction signal and the enhancement layer internal prediction signal varies over different spatial frequency components; and predictively reconstruct the enhancement layer signal using the enhancement layer prediction signal.SELECTED DRAWING: None
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to scalable video coding.

Background Art

[0002] In non-scalable coding, intra coding refers to a coding technique that uses only data of already-coded portions of the current image (e.g., reconstructed samples, coding modes, or symbol statistics), rather than reference data of already-coded images. For example, an intra-coded image (or intra image) is used within a broadcast bitstream as a so-called random access point for the decoder to synchronize to the bitstream. Also, intra images are used to limit error propagation in error-prone environments. Generally, since the images that can be used as reference images are not available here, the first image of an encoded video sequence must be coded as an intra image. Often, intra images are also used at scene cuts where they usually cannot provide a prediction signal suitable for temporal prediction.

[0003] Furthermore, the intra coding mode is used for specific regions / blocks within so-called inter images. There, they perform better than the inter coding mode with respect to rate-distortion efficiency. This is often the case within flat regions, as well as in regions where temporal prediction is performed quite poorly (occlusions, objects that are partially dissolved or faded).

[0004] In scalable coding, the concept of intra coding (encoding of intra pictures and intra blocks within inter pictures) can be extended to all pictures belonging to the same access unit or time instance. Thus, the intra coding mode for spatial or quality enhancement layers can instantaneously increase the coding efficiency while making use of inter-layer prediction from the lower layer pictures. This means that not only the already encoded parts within the picture of the current enhancement layer can be used for intra prediction, but also the lower layer pictures already encoded at the same time instance. Also, the latter concept is also referred to as inter-layer intra prediction.

[0005] In state-of-the-art hybrid video coding standards (such as H.264 / AVC or HEVC), the pictures of a video sequence are partitioned into sample blocks. The size of the blocks can be fixed or the coding method can provide a hierarchical structure that allows the blocks to be further sub-divided into smaller block sizes. Usually, the reconstruction of a block is obtained by generating a prediction signal for the block and adding the transmitted residual signal. Usually, the residual signal is transmitted using transform coding. This means that the quantization index for the transform coefficients (also referred to as transform coefficient levels) is transmitted using entropy coding techniques. And on the decoder side, these transmitted transform coefficient levels are scaled, inverse-transformed to obtain the residual signal that is added to the prediction signal. The residual signal is generated either by intra prediction (using only the data already transmitted for the current time instance) or by inter prediction (using the data already transmitted for different time instances).

[0006] If inter prediction is used, the prediction block is derived by motion compensated prediction using samples from a frame that has already been reconstructed. This can be done by uni-directional prediction (using one reference image and one set of motion parameters). Alternatively, the prediction signal can be generated by multi-hypothesis prediction. In the latter case, two or more prediction signals are superimposed. That is, for each sample, a weighted average is constructed to form the final prediction signal. The multiple prediction signals (that are superimposed) can be generated using different motion parameters for different hypotheses (e.g., different reference images or motion vectors). Also, for uni-directional prediction, it is possible to multiply the samples of the motion compensated prediction signal by a constant factor and add a constant offset to form the final prediction signal. Also, such scaling and offset correction is used for all hypotheses or for selected hypotheses in multi-hypothesis prediction.

[0007] In current state-of-the-art video coding techniques, the intra prediction signal for a block is obtained by predicting samples from the spatial neighborhood of the current block (which are blocks reconstructed before the current block in the block processing order). In the latest standards, various prediction techniques that perform prediction in the spatial domain are utilized. A refined granular directional prediction mode, with or without filtering of the samples of adjacent blocks, is extended to specific angles to generate the prediction signal. Further, there are also plane-based and DC-based prediction modes that use samples of adjacent blocks to generate a flat prediction plane or a DC prediction block.

[0008] In old video coding standards (e.g., H.263, MPEG-4), intra prediction was performed within the transform domain. In this case, the transmitted coefficients were inverse quantized. Then, for a subset of the transform coefficients, the transform coefficient values were predicted using the corresponding reconstructed transform coefficients of adjacent blocks. The inverse quantized transform coefficients were added to the predicted transform coefficient values, and the reconstructed transform coefficients were used as input to the inverse transform. The output of the inverse transform formed the final reconstructed signal for the block.

[0009] In scalable video coding as well, the base layer information can be utilized to support the prediction process for the enhancement layer. In the state-of-the-art video coding standard for scalable coding (SVC extension of H.264 / AVC), there is one additional mode to improve the coding efficiency of the intra prediction process within the enhancement layer. This mode is signaled at the macroblock level (a block of 16×16 luma samples). This mode is supported only when the collocated samples in the lower layer are encoded using the intra prediction mode. If this mode is selected for a macroblock within the quality enhancement layer, the prediction signal is assembled by the collocated samples of the reconstructed lower layer signal before the non-blocking filter operation. If the inter-layer intra prediction mode is selected within the spatial enhancement layer, the prediction signal is generated by upsampling the collocated reconstructed base layer signal (after the non-blocking filter operation). An FIR filter is used for upsampling. Generally, for the inter-layer intra prediction mode, an additional residual signal is transmitted by transform coding. Also, if it is correspondingly signaled in the bitstream, the transmission of the residual signal can be omitted (inferred to be equal to zero). The final reconstructed signal is obtained by adding the reconstructed residual signal (obtained by scaling the transmitted transform coefficient levels and applying the inverse spatial transform) to the prediction signal. SUMMARY OF THE INVENTION

Problems to be Solved by the Invention

[0010] However, in scalable video coding, it is preferable to achieve higher coding efficiency.

[0011] Therefore, an object of the present invention is to provide a concept for scalable video coding that realizes higher coding efficiency.

Means for Solving the Problems

[0012] This object is achieved by the subject matter of the independent claims enclosed.

[0013] One embodiment of the present invention is that within scalable video coding, a better predictor for predicting the enhancement layer signal obtains the enhancement layer prediction signal by forming the enhancement layer prediction signal from the inter-layer prediction signal and the intra-enhancement layer prediction signal in a method of different weighting for different spatial frequency components (i.e., by forming a weighted average of the inter-layer prediction signal and the intra-enhancement layer prediction signal in the portion to be currently reconstructed), and the weights by which the inter-layer prediction signal and the intra-enhancement layer prediction signal contribute to the enhancement layer prediction signal are such that the enhancement layer prediction signal is obtained by varying different spatial frequency components. For this reason, it is possible to interpret the enhancement layer prediction signal from the inter-layer prediction signal and the intra-enhancement layer prediction signal in a method optimized for the spectral characteristics of the individual contributing components, i.e., on the one hand the inter-layer prediction signal and on the other hand the intra-enhancement layer prediction signal. For example, based on the improvement of resolution or quality, the inter-layer prediction signal is obtained from the reconstructed base layer signal. The inter-layer prediction signal can be more accurate at lower frequencies compared to higher frequency waves. As for the intra-enhancement layer prediction signal, the characteristics can be the opposite. That is, its accuracy can increase for higher frequencies compared to lower frequencies. In this example, at low frequencies, the contribution of the inter-layer prediction signal to the enhancement layer prediction signal exceeds the contribution of the intra-enhancement layer prediction signal to the enhancement layer prediction signal with their respective weightings. And as far as high frequencies are concerned, it does not exceed the contribution of the intra-enhancement layer prediction signal to the enhancement layer prediction signal. For this reason, a more accurate enhancement layer prediction signal can be achieved. As a result, the coding efficiency increases, resulting in a higher compression ratio.

[0014] Various embodiments are described for incorporating different possibilities for incorporating the concepts just outlined into any scalable video encoding based on the concepts. For example, the formation of the weighted average can be performed either in the spatial domain or in the transform domain. Performing spectral weighted averaging requires a transform to be performed on individual contributions, namely, the inter-layer prediction signal and the intra-enhancement layer prediction signal. However, avoid spectrally filtering either the inter-layer prediction signal or the intra-enhancement layer prediction signal in the spatial domain, including, for example, FIR or IIR filtering. However, performing the formation of spectral weighted averaging in the spatial domain can avoid detouring the individual contributions to the weighted average via the transform domain. The decision as to which region is actually selected to perform the formation of spectral weighted averaging can depend on whether the scalable video data stream includes the residual signal in the form of transform coefficients for the portion that is to be currently configured within the enhancement layer signal. If not, the detour via the transform domain can be stopped. On the other hand, if the residual signal is present, the detour via the transform domain is more advantageous because it allows directly adding the spectral weighted averaging in the transform domain to the transmitted residual signal in the transform domain.

[0015] One aspect of the present invention is that information available from base layer encoding / decoding, i.e., base layer hints, can be utilized to make motion compensation prediction in the enhancement layer more efficient by more efficiently encoding enhancement layer motion parameters. In particular, a set of motion parameter candidates collected from adjacent already reconstructed blocks of a frame of the enhancement layer signal can possibly be augmented by one or more sets of base layer motion parameters of blocks of the base layer signal (the base layer signal collocated with the blocks of the frame of the enhancement layer signal). As a result, the available quality of the set of motion parameter candidates is improved based on the fact that motion compensation prediction of a block of the enhancement layer signal can be performed by selecting one of the motion parameter candidates of the augmented set of motion parameter candidates and using the selected motion parameter candidate for prediction. Additionally, or alternatively, the list of motion parameter candidates of the enhancement layer signal can be ordered depending on the base layer motion parameters involved in base layer encoding / decoding. Thus, the probability distribution for selecting an enhancement layer motion parameter from the ordered list of motion parameter candidates can be compressed so that, for example, clearly signaled index syntax elements can be encoded using fewer bits, e.g., using entropy coding, etc. Further, additionally, or alternatively, the index used within base layer encoding / decoding can act as a basis for determining an index into the list of motion parameter candidates for the enhancement layer. Thus, any signaling of an index for the enhancement layer can be completely avoided. Or, simply the prediction deviation determined in this way for the index can be transmitted within the enhancement layer substream, and as a result, the encoding efficiency can be improved.

[0016] One aspect of the present invention is that scalable video coding is made more efficient by evaluating the spatial variation of base layer coding parameters on the base layer signal to derive / select the sub - division of sub - blocks of the enhancement layer to be used for enhancement layer prediction within a set of possible sub - divisions of sub - blocks of the enhancement layer block. For this reason, if so, less signaling overhead has to be spent in the enhancement layer data stream to signal this sub - division of sub - blocks. The sub - division of sub - blocks thus selected can be used when predictively encoding / decoding the enhancement layer signal.

[0017] One aspect of the present invention is that if the sub - division of sub - blocks of each transform coefficient block is controlled based on the base layer residual signal or the base layer signal, the coding based on the sub - blocks of the transform coefficient block of the enhancement layer can be made more efficient. In particular, by utilizing each base layer hint, the sub - blocks can be made longer along the horizontal spatial frequency axis with respect to the edge extension observable from the base layer residual signal or the base layer signal. For this reason, with an increasing probability, each sub - block is filled with either almost entirely significant transform coefficients (i.e., transform coefficients not quantized to zero) or insignificant transform coefficients (i.e., only transform coefficients quantized to zero), while with a decreasing probability, it is possible to adapt the shape of the sub - blocks to the estimated distribution of the energy of the transform coefficients of the enhancement layer transform coefficient block such that each sub - block has the same number of significant transform coefficients on one side and the same number of insignificant transform coefficients on the other side. However, due to the fact that sub - blocks having no significant transform coefficients can be efficiently signaled in the data stream, for example, by using a single flag, and due to the fact that sub - blocks almost entirely filled with significant transform coefficients do not require waste of signaling amount for encoding the insignificant transform coefficients that may be scattered therein, the coding efficiency for encoding the transform coefficient block of the enhancement layer increases.

[0018] One aspect of the present invention is that the coding efficiency of scalable video coding can be increased by substituting lost spatial intra prediction parameter candidates in the spatial neighborhood of the current block of the enhancement layer with the intra prediction parameters of the collocated blocks of the base layer signal. Therefore, the coding efficiency for coding the spatial intra prediction parameters is likely to increase due to the improved prediction quality of the set of intra prediction parameters of the enhancement layer, or, more precisely, is expected to increase. A suitable predictor for the intra prediction parameters for the intra predicted blocks of the enhancement layer is useful, and as a result, increases the likelihood that the signaling of the intra prediction parameters for each enhancement layer block can be performed with fewer bits on average.

[0019] Further advantageous implementations are described in the dependent claims.

[0020] Preferred embodiments are further described below with reference to the drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0021]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 15c

Figure 15d

Figure 16

Figure 17

Figure 18

Figure 19a

Figure 19b

Figure 19c

Figure 20

Figure 21

Figure 22

Figure 23

Figure 24a

Figure 24b

Figure 25

Figure 26

Figure 27

Figure 28

Figure 29

Figure 30

Figure 31

Figure 32

Figure 33

Figure 34

Figure 35

Figure 36

Figure 37

Figure 38

Figure 39

Figure 40

Figure 41

Figure 42

Figure 43

Figure 44

Figure 45

Figure 46

DETAILED DESCRIPTION OF THE INVENTION

[0022] FIG. 1 shows in a general way an embodiment for a scalable video encoder into which the embodiments further outlined below can be incorporated. The scalable video encoder of FIG. 1 is generally denoted using reference numeral 2 and receives and encodes video 4. The scalable video encoder 2 is configured to encode video 4 into data stream 6 in a scalable manner. That is, data stream 6 has a first portion 6a having video 4 encoded therein with a first information content amount and a further portion 6b having video 4 encoded therein with an information content amount greater than that of portion 6a. For example, the information content amounts of portions 6a and 6b can differ in terms of quality or fidelity, i.e., the amount of deviation per pixel from the original video 4 and / or in terms of spatial resolution. However, other forms of different information content amounts can also be applied, for example, in terms of color fidelity. Portion 6a can be called a base layer data stream or a base layer sub-stream. On the other hand, portion 6b can be called an enhancement layer data stream or an enhancement layer sub-stream.

[0023] The scalable video encoder 2 is configured to utilize the redundancy between versions 8a and 8b of the reconstructable video 4 from the base layer sub-stream 6a without the enhancement layer sub-stream 6b on the one hand and from both sub-streams 6a and 6b on the other hand. To do so, the scalable video encoder 2 can use inter-layer prediction.

[0024] As shown in FIG. 1, the scalable video encoder 2 can alternatively receive two versions 4a and 4b of the video 4. Both versions 4a and 4b have different information capacities such that the base layer substream 6a and the enhancement layer substream 6b are exactly different from each other. Thus, for example, the scalable video encoder 2 is configured to generate substreams 6a and 6b. As a result, the base layer substream 6a has the version 4a encoded therein. On the other hand, the enhancement layer data stream (substream) 6b has the version 4b encoded therein using inter-layer prediction based on the base layer substream 6b. The encoding of both substreams 6a and 6b can be lossy.

[0025] Even if the scalable video encoder 2 only receives the original version of the video 4, the scalable video encoder 2 can be configured to internally derive two versions 4a and 4b therefrom by obtaining the base layer version 4a, for example, by spatial downscaling and / or tone mapping (mapping) from a higher bit depth to a lower bit depth.

[0026] Figure 2 shows a scalable video decoder adapted to the scalable video encoder 2 of FIG. 1 in a similar manner suitable for incorporating the embodiments outlined below. The scalable video decoder of FIG. 2 is generally denoted using reference numeral 10, and the scalable video decoder is generally configured to decode (decipher) the encoded data stream 6 such that, if both parts 6a and 6b of the data stream 6 reach the scalable video decoder 10 in a complete manner, the enhanced layer version 8b of the video is reconstructed therefrom, or, if, for example, part 6b is unavailable due to a transmission loss or the like, the base layer version 8a of the video is reconstructed therefrom. That is, the scalable video decoder 10 is configured to be able to reconstruct version 8a from only the base layer sub-stream 6a and to be able to reconstruct version 8b using inter-layer prediction from both parts 6a and 6b.

[0027] Before the following more detailed embodiments of the present invention (i.e., the embodiments show how the embodiments of FIGS. 1 and 2 are implemented) are described in detail, a more detailed implementation of the scalable video encoder and decoder of FIGS. 1 and 2 will be described with respect to FIGS. 3 and 4. FIG. 3 shows a scalable video encoder 2 comprising a base layer encoder 12, an enhancement layer encoder 14 and a multiplexer 16. The base layer encoder 12 is configured to encode the base layer version 4a of the input video, while the enhancement layer encoder 14 is configured to encode the enhancement layer version 4b of the video. Accordingly, the multiplexer 16 receives the base layer sub-stream 6a from the base layer encoder 12 and the enhancement layer sub-stream 6b from the enhancement layer encoder 14 and multiplexes both of them into the encoded data stream 6 when outputting.

[0028] As shown in FIG. 3, both encoders 12 and 14 can be predictive encoders that use, for example, spatial prediction and / or temporal prediction to encode their respective input versions 4a and 4b into their respective sub-streams 6a and 6b. In particular, encoders 12 and 14 can each be hybrid video block encoders. That is, each of encoders 12 and 14 can be configured to encode each input version of the video on a block-by-block basis, for example, while different prediction modes are selected for each block of blocks into which the respective images or frames of input versions 4a and 4b of the video are each sub-divided. The different prediction modes of base layer encoder 12 can include spatial and / or temporal prediction modes. On the other hand, enhancement layer encoder 14 can additionally support an inter-layer prediction mode. The sub-division into blocks can be different between the base layer and the enhancement layer. The prediction mode, the prediction parameters for the prediction modes selected for the various blocks, the prediction residuals, and, optionally, the block sub-divisions of the respective video versions can be described by respective encoders 12, 14 using a respective syntax that includes syntax elements to be encoded in respective sub-streams 6a, 6b in turn using entropy coding. Inter-layer prediction is utilized, for example, one or more times to predict samples of the enhancement layer video, prediction mode, prediction parameters, and / or block sub-divisions as just mentioned in the examples of 2, 3. Thus, both base layer encoder 12 and enhancement layer encoder 14 can each include prediction encoders 18a, 18b followed by entropy encoders 19a, 19b. On the other hand, prediction encoders 18a, 18b form a syntax element stream from the respective inbound versions 4a and 4b using predictive coding. Entropy encoders 19a, 19b entropy code the syntax elements output by respective prediction encoders 18a, 18b. As just mentioned, the inter-layer prediction of encoder 2 can be relevant at different times within the encoding procedure of the enhancement layer.Accordingly, predictive encoder 18b is shown as being connected to one or more of predictive encoder 18a, its output, and entropy encoder 19a. Similarly, entropy encoder 19b may optionally utilize inter-layer prediction, for example, by predicting the context used for entropy encoding from the base layer. Accordingly, entropy encoder 19b is shown as being optionally connected to any of the elements of base layer encoder 12.

[0029] In the same manner as FIG. 2 with respect to FIG. 1, FIG. 4 shows a possible implementation of scalable video decoder 10 adapted to the scalable video encoder of FIG. 3. Accordingly, the scalable video decoder 10 of FIG. 4 includes a demultiplexer 40 that receives data stream 6 to obtain substreams 6a and 6b, a base layer decoder 80 configured to decode base layer substream 6a, and an enhancement layer decoder 60 configured to decode enhancement layer substream 6b. As shown, decoder 60 is connected to base layer decoder 80 to receive information therefrom for utilizing inter-layer prediction. Thereby, base layer decoder 80 can reconstruct base layer version 8a from base layer substream 6a. And enhancement layer decoder 60 is configured to reconstruct enhancement layer version 8b of the video using enhancement layer substream 6b. Similar to the scalable video encoder of FIG. 3, enhancement layer decoder 60 and base layer decoder 80 may each internally include entropy decoders 100, 320, followed by prediction decoders 102, 322.

[0030] To simplify the understanding of the following embodiments, FIG. 5 illustratively shows different versions of video 4, namely base layer versions 4a and 8a that are simply offset from each other by encoding losses. Similarly, enhancement layer versions 4b and 8b are simply offset from each other by encoding losses. As shown, the base layer signal and the enhancement layer signal can each be composed of a series of images 22a and 22b. They are shown in FIG. 5 to be registered with each other along the time axis 24, i.e., in addition to the temporally corresponding images 22b of the enhancement layer signal, also with the images 22a of the base layer version. As described above, the images 22b can represent video 4 with a higher spatial resolution and / or a higher fidelity, etc., for example, with a higher bit depth of the sample values of the images. Solid and dashed lines are used to show the order of encoding / decoding as defined between images 22a, 22b. According to the example shown in FIG. 5, the order of encoding / decoding is such that the base layer image 22a at a given time's time stamp / instance traverses in front of the enhancement layer image 22b of the same time stamp of the enhancement layer signal, crossing images 22a and 22b. With respect to the time axis 24, images 22a, 22b can be traversed by the order of encoding / decoding 26 in the order of provision time. However, an order that deviates from the order of provision time of images 22a, 22b is also possible. Neither encoder 2 nor decoder 10 needs to encode / decoder continuously along the order of encoding / decoding 26. Rather, encoding / decoding can be used in parallel. The order of encoding / decoding 26 can define the availability between portions of the base layer signal and the enhancement layer signal that are adjacent to each other in a spatial, temporal, and / or inter-layer sense. As a result, when encoding / decoding the current portion of the enhancement layer, the available portions of the current enhancement layer portion are defined through the order of encoding / decoding. Thus, since only adjacent portions that are available according to this order of encoding / decoding 26 are used by the encoder for prediction, the decoder accesses the same information source to correct the prediction.

[0031] For the following figures, how the scalable video encoder or decoder described above for FIGS. 1-4 is implemented to form an embodiment of the present invention according to one aspect of the present invention will be described. Possible implementations of the aspects described below are discussed using the designation "Aspect C".

[0032] In particular, FIG. 6 illustrates an image 22b of an enhancement layer signal indicated using reference numeral 360 and an image 22a of a base layer signal indicated using reference numeral 200. Temporally corresponding images of different layers are shown in a manner shown relative to the time axis 24. Using hatching, portions 200 and 36 within the base and enhancement layer signals that have already been encoded / decoded according to the encoding / decoding order are distinguished from portions that have not yet been encoded or decoded according to the encoding / decoding order shown in FIG. 5. Also, FIG. 6 shows a portion 28 of the enhancement layer signal 360 that is currently being encoded / decoded.

[0033] According to the presently described embodiment, the prediction of portion 28 uses both intra-layer prediction within the enhancement layer itself and inter-layer prediction from the base layer to predict portion 28. However, the predictions are combined so that they contribute to the final prediction of portion 28 in a spectrally varying manner. As a result, in particular, the ratio between both contributions varies spectrally.

[0034] In particular, portion 28 is predicted spatially or temporally from a portion of the enhanced layer signal 400 that has already been reconstructed (i.e., the portion indicated by the hatching in the enhanced layer signal 400 in FIG. 6). Spatial prediction is described using arrow 30. On the other hand, temporal prediction is described using arrow 32. Temporal prediction may include motion compensation prediction, for example, as motion vector information is transmitted within the enhanced layer substream for the current portion 28. The motion vector indicates the replacement of a portion of the reference picture of the enhanced layer signal 400 to be copied to obtain the temporal prediction of the current portion 28. Spatial prediction 30 may include a portion of the spatially adjacent portion to be estimated, the already encoded / decoded portion of picture 22b, and the spatially adjacent current portion 28 within the current portion 28. For this purpose, intra prediction information such as the estimation (or angular) direction may be signaled within the enhanced layer substream for the current portion 28. Also, a combination of spatial prediction 30 and temporal prediction 32 may be used as well. In any case, as a result, the enhanced layer internal prediction signal 34 is obtained as described in FIG. 7.

[0035] To obtain another prediction of the current portion 28, inter-layer prediction is used. For this purpose, the base layer signal 200 is at a portion 36 that spatially and temporally corresponds to the current portion 28 of the enhanced layer signal 400, and the inter-layer prediction signal for the current portion 28 undergoes an improvement in resolution or quality to obtain an increasing potential resolution. The improvement procedure is described using arrow 38 in FIG. 6 and results in an inter-layer prediction signal 39 as shown in FIG. 7.

[0036] Therefore, two prediction contributions 34 and 39 exist for the current portion 28. And the weighted average of both contributions is formed to obtain the enhancement layer prediction signal 42 for the current portion 28 in such a way that the weights by which the inter-layer prediction signal and the intra-enhancement layer prediction signal contribute to the enhancement layer prediction signal 42 vary differently for spatial frequency components, as schematically shown at 44 in FIG. 7. FIG. 7 exemplarily shows a case where, for every spatial frequency component, the weights by which the prediction signals 34 and 38 contribute to the final prediction signal, for all spectral components, however, add the same value 46 in a state where the ratio between the weight applied to the prediction signal 34 and the weight applied to the prediction signal 39 varies spectrally.

[0037] On the other hand, the prediction signal 42 can be directly used for the current portion 28 by the enhancement layer signal 400. Alternatively, the residual signal can be provided in the enhancement layer sub-stream 6b of the current portion 28 resulting from the combination 50 with the prediction signal 42, for example, as the addition shown in FIG. 7, within the reconstructed version 54 of the current portion 28. As an intermediate note, it should be noted that both the scalable video encoder and decoder are hybrid video decoders / encoders that use transform coding to encode / decode the prediction residuals and use predictive coding.

[0038] Summarizing the description of FIGS. 6 and 7, the enhancement layer substream 6b may include an intra prediction parameter 56 for controlling spatial and / or temporal prediction 30, 32 for the current portion 28, and optionally a weighting parameter 58 for controlling the formation of a spectrally weighted average 41, and residual information 59 for signaling the residual signal 48. On the other hand, the scalable video encoder accordingly determines all of these parameters 56, 58, 59 and inserts the parameters 56, 58, 59 into the enhancement layer substream 6b. The scalable video decoder uses the parameters 56, 58, 59 to reconstruct the current portion 28 as outlined above. All of these elements 56, 58, 59 may undergo some quantization. And accordingly, the scalable video encoder may determine these parameters / elements (i.e., using a rate / distortion cost function as quantization). Interestingly, the encoder 2 uses the parameters / elements 56, 58, 59 thus determined to serve as a basis for any prediction for the portion of the enhancement layer signal 400, for example, following the encoding / decoding order, to obtain a reconstructed version 54 of the current portion 28.

[0039] There are different possibilities for the weighting parameters 58 and the way they control the formation of the spectral weighted average at 41. For example, the weighting parameter 58 can signal for the current portion 28 only one of two states, namely, one state that activates the formation of the spectral weighted average as previously described, and the other state that deactivates the contribution of the interlayer prediction signal 38. As a result, the final enhancement layer prediction signal 42 is then created only by the enhancement layer internal prediction signal 34. The weighting parameter 58 for the current portion 28 can switch between the activation of one spectral weighting average formation and the interlayer prediction signal 39 that forms the enhancement layer prediction signal 42 alone on the other hand. Also, the weighting parameter 58 can be designed to signal one of the three states / alternatives mentioned. Alternatively / additionally, the weighting parameter 58 can further control the spectral weighted average formation 41 with respect to the spectral change of the ratio between the weightings by which the prediction signals 34 and 39 contribute to the final prediction signal 42 for the current portion 28. It will be explained later that the spectral weighted average formation 41 can involve filtering one or both of the prediction signals 34 and 39 before adding them, for example, using a high-pass filter and / or a low-pass filter. In that case, the weighting parameter 58 can signal the filter characteristics for the filter to be used for the prediction of the current portion 28. Or it can filter. Alternatively, the weighting parameter 58 can signal / set the individual weighting values of these spectral components, as it will be explained below that the spectral weighting in step 41 can be achieved by the individual weighting of the spectral components in the transform domain.

[0040] Additionally, or alternatively, the weighting parameter for the current portion 28 can signal whether the spectral weighting is performed within the transform domain or the spatial domain in step 41.

[0041] FIG. 9 illustrates an embodiment for performing spectral weighted average composition within a spatial region. The prediction signals 39 and 34 are illustrated as being obtained in the form of respective pixel arrays that match the pixel raster of the current portion 28. To perform spectral weighted average formation, the pixel arrays of both prediction signals 34 and 39 are shown to undergo filtering. FIG. 9 illustrates filtering as an example by showing filter kernels 62 and 64 that traverse the pixel arrays of prediction signals 34 and 39 to perform, for example, FIR filtering. However, IIR filtering is also possible. Further, only one of the prediction signals 34 and 39 may undergo filtering. Since the transfer functions of both filters 62 and 64 are different, the addition 66 of the results of filtering the pixel arrays of prediction signals 39 and 34 results in spectral weighted average formation, that is, the enhancement layer prediction signal 42. In other words, addition 66 easily adds the juxtaposed samples in the filtered prediction signals 39 and 34 using filters 62 and 64, respectively. As a result, 62 - 66 results in spectral weighted average formation 41. FIG. 9 shows that in the case of residual information 59 existing in the form of transform coefficients, the residual signal 48 within the transform region is signaled and the inverse transform 68 can be used to yield the spatial region in the form of pixel array 70, and as a result, the combination 52 that yields the reconstructed version 55 is realized by a simple pixel-wise addition of the residual signal array 70 and the enhancement layer prediction signal 42.

[0042] Again, it is to be recalled that prediction is performed by a scalable video encoder and decoder, using prediction for reconstruction within the decoder and encoder, respectively.

[0043] FIG. 10 illustratively shows how spectral weighted averaging is performed within the transform domain. Here, the pixel arrays of prediction signals 39 and 34 each undergo a transform 72 and 74, respectively, resulting in spectral decompositions 76 and 78, respectively. Each spectral decomposition 76 and 78 has one transform coefficient per spectral component, creating a transform coefficient array. Each transform coefficient block 76 and 78 is multiplied by a corresponding block of weights, i.e., blocks 82 and 84. As a result, the transform coefficients of blocks 76 and 78 are individually weighted for each spectral component. For each spectral component, the weighting values of blocks 82 and 84 may have a common value added to all spectral components, but this is not obligatory. In fact, multipliers 86 between blocks 76 and 82 and multiplier 88 between blocks 78 and block 84 each represent spectral filtering within the transform domain. And the transform coefficient / spectral component unit addition 90 concludes the spectral weighted averaging formation 41 to yield a transform domain version of the enhancement layer prediction signal 42 in the form of a block of transform coefficients. As shown in FIG. 10, in the case of the residual signal 59 that signals the residual signal 48 in the form of a transform coefficient block, the residual signal 59 is easily added to the transform coefficient block representing the enhancement layer prediction signal 42 (or another combination) 52 to yield a reconstructed version of the current portion 28 within the transform domain. Thus, the inverse transform 84 applied to the additional result of the combination 52 yields the pixel array that reconstructs the current portion 28, i.e., the reconstructed version 54.

[0044] As described above, the current parameters within the enhancement layer substream 6b for the current portion 28, such as the residual information 59 and the weighting parameter 58, can be signaled as to whether the averaging formation 41 is performed within the transform domain shown in FIG. 10 or within the spatial domain according to FIG. 9. For example, if the residual information 59 indicates the absence of any transform coefficient block for the current portion 28. Also, the spatial domain can be used. Alternatively, the weighting parameter 58 can switch between both domains regardless of whether the residual information 59 includes transform coefficients or not.

[0045] Thereafter, to obtain an intra-layer enhancement layer prediction signal, it is explained that a difference signal can be calculated and managed between the already reconstructed part of the enhancement layer signal and the inter-layer prediction signal. The spatial prediction of the difference signal at a first part juxtaposed to a part of the enhancement layer signal is currently reconstructed from a second part of the difference signal. Spatially adjacent to the first part of the enhancement layer signal and belonging to the already reconstructed part, it can then be used to spatially predict the difference signal. Alternatively, the temporal prediction of the difference signal at the first part is juxtaposed to a part of the enhancement layer signal and is currently reconstructed from the second part of the difference signal while belonging to a previously reconstructed frame of the enhancement layer signal, and can be used to obtain a temporally predicted difference signal. The combination of the inter-layer prediction signal and the predicted difference signal can be used to obtain an intra-layer enhancement layer prediction signal, which is then combined with the inter-layer prediction signal.

[0046] Regarding the following figures, it is described how a scalable video encoder or decoder as described above for FIGS. 1 to 4 is executed to form an embodiment of the present application according to another aspect of the application.

[0047] To explain this content, reference is made to FIG. 11. FIG. 11 shows the possibility of performing a spatial prediction 30 of the current part 28. As a result, the following description of FIG. 11 can be combined with the description regarding FIGS. 6 to 10. In particular, the content described below will be described later with respect to the illustrated example of implementation by referring to "aspects" X and Y.

[0048] The situation shown in FIG. 11 corresponds to that shown in FIG. 6. That is, the base layer signal 200 and the enhancement layer signal 400 are shown. The already encoded / decoded portions are shown using diagonal lines. Within the enhancement layer signal 400, the portion that is currently to be encoded / decoded has adjacent blocks 92 and 94. Here, by way of example, for both blocks 92 and 94 that exemplarily have the same size as the current block 28, block 92 is drawn above the current portion 28 and block 94 is drawn to the left. However, the size match is not obligatory. Rather, the portions of the blocks into which the image 22b of the enhancement layer signal 400 is sub-divided can have different sizes. They are not even restricted to being quadrilaterals. They may be rectangles or other shapes. Further, the current block 28 has adjacent blocks that are not explicitly shown in FIG. 11. However, the adjacent blocks have not yet been decoded / encoded. That is, the adjacent blocks follow in the order of encoding / decoding and as a result are not available for prediction. Beyond this, there are blocks other than blocks 92 and 94 that have already been encoded / decoded according to the order of encoding / decoding (blocks adjacent to the current block 28, such as block 96 that is diagonally adjacent at the upper left corner of the current block 28). However, blocks 92 and 94 are predetermined adjacent blocks that serve to predict the intra prediction parameters for the current block 28 that is the subject of intra prediction 30 in the example considered here. The number of such predetermined adjacent blocks is not limited to two. It may be more or simply one as well.

[0049] A scalable video encoder and a scalable video decoder may determine a set of predetermined adjacent blocks (here, blocks 92, 94) from a set of already encoded adjacent blocks. Here, blocks 92 to 96 depend on a predetermined sample position 98 within the current portion 28, such as its upper left sample. For example, only those already encoded adjacent blocks of the current portion 28 may form a set of "predetermined adjacent blocks" that includes sample positions immediately adjacent to the predetermined sample position 98. In any case, the adjacent already encoded / decoded blocks include samples 102 adjacent to the current block 28 based on the sample values of the samples for which the region of the current block 28 is to be spatially predicted. For this purpose, spatial prediction parameters such as 56 are signaled within the enhancement layer substream 6b. For example, the spatial prediction parameter for the current block 28 indicates the spatial direction in which the sample value of the sample 102 is to be copied within the region of the current block 28.

[0050] In any case, at least as long as the temporally corresponding image 22a is relevant to the spatially corresponding region, when spatially predicting the current block 28 using block-based prediction as described above, for example, using a block-based selection between a spatial prediction mode and a temporal prediction mode, the scalable video decoder / encoder has already reconstructed (encoded in the case of the encoder) the base layer 200 using the base layer substream 6a.

[0051] In FIG. 11, several blocks 104 into which the image 22a arranged in time of the base layer signal 200 is sub-divided exist within and around an area that locally corresponds to the current portion 28 shown illustratively. It is just the case of the spatially predicted blocks within the enhancement layer signal 400. Similar to the case of the spatially predicted blocks within the enhancement layer signal 400, the spatial prediction parameters are included in or signaled in the base layer sub-stream for those blocks 104 within the base layer signal 200 for which the selection of the spatial prediction mode is signaled.

[0052] Here, illustratively, in order to enable the reconstruction of the enhancement layer signal from the coded data stream for the block 28 for which the spatial intra-layer prediction 30 is selected, the intra prediction parameters are used and coded within the following bit stream.

[0053] The intra prediction parameters are often coded using the concept of the "most likely intra prediction parameters", which is a fairly small subset of all possible intra prediction parameters. For example, the set of the most likely intra prediction parameters can include one, two or three intra prediction parameters. On the other hand, for example, the set of all possible intra prediction parameters can include 35 intra prediction parameters. If the intra prediction parameter is included in the set of the most likely intra prediction parameters, it can be signaled in the bit stream with a small number of bits. If the intra prediction parameter is not included in the set of the most likely intra prediction parameters, its signaling in the bit stream requires more bits. Therefore, the amount of bits to be spent on the syntax elements for signaling the intra prediction parameters for the current intra-predicted block depends on the quality of the set of the most likely, or perhaps advantageous, intra prediction parameters. Assuming that this concept can be used to appropriately derive the set of the most likely intra prediction parameters, on average fewer bits are required to code the intra prediction parameters.

[0054] Typically, the set of most likely intra prediction parameters is selected in a way that it includes the intra prediction parameters of directly adjacent blocks and / or additionally often uses intra prediction parameters in the form of, for example, predefined parameters. For example, since the main gradient directions of adjacent blocks are the same, it is generally advantageous to include the intra prediction parameters of adjacent blocks within the set of most likely intra prediction parameters.

[0055] However, if the adjacent blocks are not coded in the spatial intra prediction mode, their parameters are not available at the decoder side.

[0056] In scalable coding, however, it is possible to use the intra prediction parameters of collocated base layer blocks. Thus, according to the embodiments outlined below, this situation is exploited using the intra prediction parameters of collocated base layer blocks in the case of non-coded adjacent blocks within the spatial intra prediction mode.

[0057] As a result, according to FIG. 11, a set of perhaps advantageous intra prediction parameters for the current enhancement layer block is constructed by examining the intra prediction parameters of predefined adjacent blocks and, for example, in any case where a predefined adjacent block does not have an appropriate intra prediction parameter associated therewith, since the predefined adjacent block is not coded in the intra prediction mode, by exceptionally reclassifying the block collocated in the base layer.

[0058] First, it is checked whether a predetermined adjacent block such as block 92 or 94 of the current block 28 has been predicted using the spatial intra prediction mode. That is, it is checked whether the spatial intra prediction mode has been selected for its adjacent block. Thereby, the intra prediction parameter of the adjacent block is included in a set of probably advantageous intra prediction parameters for the current block 28, or, if any, alternatively, in the intra prediction parameters of the collocated block 108 of the base layer. This process can be performed for each of the predetermined adjacent blocks 92 and 94.

[0059] For example, if each of the predetermined adjacent blocks is not a spatial intra prediction block, instead of using an initial prediction or the like, the intra prediction parameters of block 108 of the base layer signal 200 are included in a set of probably advantageous inter prediction parameters for the current block 28 that is collocated with the current block 28. For example, the collocated block 108 is determined using the predetermined sample position 98 of the current block 28. That is, block 108 covers position 106 that locally corresponds to the predetermined sample position 98 in the temporally aligned image 22a of the base layer signal 200. Naturally, a further check can be performed as to whether this collocated block 108 in the base layer signal 200 is actually a spatial intra prediction block. In the case of FIG. 11, this is illustratively explained to be the case. However, if the collocated block is also not encoded in the intra prediction mode, a set of probably advantageous intra prediction parameters may be left with no contribution for its predetermined adjacent block. Or, the initial intra prediction parameters may be used instead as an alternative. That is, the initial intra prediction parameters are inserted into a set of probably advantageous intra prediction parameters.

[0060] Thus, if the block 108 collocated with the current block 28 is spatially intra prediction, then, as an alternative, the intra prediction parameter signaled within the base layer substream 6a is used for the predefined adjacent block 92 or 94 of the current block 28 that has no intra prediction parameter because the intra prediction parameter is encoded using another prediction mode such as the temporal prediction mode.

[0061] According to another embodiment, in a given case, if each predefined adjacent block is in the intra prediction mode, the intra prediction parameter of the predefined adjacent block is replaced by the intra prediction parameter of the collocated base layer block. For example, a further check, such as whether the intra prediction parameter meets a given criterion, can be performed for any predefined adjacent block in the intra prediction mode. If the given criterion is not met by the intra prediction parameter of the adjacent block but is met by the intra prediction parameter of the collocated base layer block, then the replacement is performed regardless of the very adjacent block being intra-coded. For example, if the intra prediction parameter of the adjacent block does not represent the angular intra prediction mode (however, for example, the DC or planar intra prediction mode), but the intra prediction parameter of the collocated base layer block represents the angular intra prediction mode, then the intra prediction parameter of the adjacent block can be replaced by the intra prediction parameter of the base layer block.

[0062] The inter prediction parameters for the current block 28 are then determined based on the enhancement layer substream 6b for the current block 28 and syntax elements present in the coded data stream such as perhaps a set of advantageous intra prediction parameters. That is, the syntax elements may be coded using fewer bits in the case of the inter prediction parameters for the current block 28 which are perhaps members of a set of advantageous intra prediction parameters than in the case of the remaining members of the set of possible intra prediction parameters which do not lead to a set of perhaps advantageous intra prediction parameters.

[0063] The set of possible intra prediction parameters can include several angular direction modes according to which the current block is filled by copying from already coded / decoded adjacent samples along the angular direction of each mode / parameter, one DC mode according to which the samples of the current block are set to a constant value determined based on, for example, some averages of already coded / decoded adjacent samples, etc., and a plane mode according to which the samples of the current block are set to a value distribution following a linear function of the slopes and intercepts of x and y based on, for example, already coded / decoded adjacent samples.

[0064] FIG. 12 shows the possibility of how an alternative to the spatial prediction parameters obtained from the collocated blocks 108 of the base layer can be used together with syntax elements signaled in the enhancement layer substream. FIG. 12 shows an enlarged view of the current block 28 together with adjacent already coded / decoded samples 102 and predetermined adjacent blocks 92 and 94. Also, FIG. 12 exemplarily shows the angular direction 112 indicated by the spatial prediction parameters of the collocated blocks 108.

[0065] The syntax element 114 signaled within the enhancement layer substream 6b for the current block 28 can signal, for example as shown in FIG. 13, a conditionally encoded index 118 into a list 122, illustratively here as an angular direction 124, which is the result of possible favorable intra prediction parameters. Or, if the actual intra prediction parameter 116 is not within the most likely set 122 and is an index 123 within a list 125 of possible intra prediction modes that are potentially excluded as shown at 127, the candidates in list 122 then identify the actual intra prediction parameter 116. Encoding of the syntax element can consume fewer bits in the case of the actual intra prediction parameter belonging within list 122. For example, the syntax element can include a flag and an index field. The flag indicates whether the index points to either list 122 or list 125, i.e., whether to include or exclude it from the members of list 122. The syntax element includes a field that identifies one of the members 124 of list 122 or an escape code. And, in the case of an escape code, the syntax element includes a second field that identifies a member from list 125 that includes or excludes the member of list 122. The order within member 124 in list 122 can be determined, for example, based on an initial setting rule.

[0066] Accordingly, a scalable video decoder may obtain, or retrieve, a syntax element 114 from the enhancement layer sub-stream 6b. And a scalable video encoder may insert the syntax element 114 into the enhancement layer sub-stream 6b. And then, for example, the syntax element 114 is used to index one spatial prediction parameter from the list 122. When forming the list 122, the aforementioned alternative may be executed by checking whether the predetermined adjacent blocks 92 and 94 are of a spatial prediction coding mode type. Otherwise, as described above, it is checked whether the collocated block 108 is, for example, a spatially predicted block in order, and if so, the same spatial prediction parameter, such as the angular direction 112 used to spatially predict this collocated block 108, is included in the list 122. Also, if the base layer block 108 does not contain a suitable intra prediction parameter, the list 122 may be left without contribution from the respective predetermined adjacent blocks 92 or 94. Because, to avoid the list 122 being empty, for example, since both of the predetermined adjacent blocks 92, 98 are, for example, inter predicted, at least one of the members 124 may be unconditionally determined using the initial intra prediction parameter, similar to the collocated block 108 lacking a suitable intra prediction parameter. Alternatively, it may be allowed for the list 122 to be empty.

[0067] Naturally, the aspects described with respect to FIGS. 11 to 13 can be connected to the aspects outlined with respect to FIGS. 6 to 10. In particular, the intra prediction obtained using the spatial intra prediction parameters drawn out bypassing the base layer according to FIGS. 11 to 13 can represent, in particular, the enhancement layer internal prediction signal 34 of the aspects of FIGS. 6 to 10, because it is combined with the interlayer prediction signal 38 in a spectrally weighted manner as described above.

[0068] For the following drawings, as described above with respect to FIGS. 1-4, it is explained how a scalable video encoder or decoder can be implemented to form an embodiment of the present application in accordance with another aspect of the application. Later, additional implementation examples are presented with reference to Aspects T and U for the aspects described below.

[0069] Refer to FIG. 14, which shows the images 22b and 22a of the enhancement layer signal 400 and the base layer signal 200, respectively, in a temporal registration manner. Currently, the portion to be encoded / decoded is indicated by 28. In accordance with the current aspect, the base layer signal 200 is predictively encoded by a scalable video encoder using base layer encoding parameters that spatially vary the base layer signal, and is predictively reconstructed by a scalable video decoder. The spatial variation uses the hatched portion 132 where the base layer encoding parameters used to predictively encode / reconstruct the base layer encoding parameters in FIG. 14 are constant, and is surrounded by a non-hatched region where the base layer encoding parameters change when transitioning from the hatched portion 132 to the non-hatched region. In accordance with the aspect outlined above, the enhancement layer signal 400 is encoded / decoded in block units. The current portion 28 is such a block. In accordance with the aspect outlined above, the sub-division of sub-blocks for the current portion 28 is selected from a set of possible sub-block sub-divisions based on the spatial variation of the base layer encoding parameters within the juxtaposed portion 134 of the base layer signal 200, i.e., within the spatially juxtaposed portion of the temporally corresponding image 22a of the base layer signal 200.

[0070] In particular, instead of signaling within the sub-division information of the enhancement layer sub-stream 6b for the current portion 28, the above description presents selecting the sub-division of the sub-block within the set of possible sub-divisions of the current portion 28 such that the sub-division of the selected sub-block is the coarsest within the set of possible sub-divisions of the sub-block. Therein, when shifted onto the juxtaposed portion 134 of the base layer signal, the base layer coding parameters divide the base layer signal 200 such that they are sufficiently similar to each other within each sub-block of the sub-division of each sub-block. For ease of understanding, refer to FIG. 15a. FIG. 15a shows the portion 28 depicting the spatial variation of the base layer coding parameters within the juxtaposed portion 134 using hatching. In particular, the portion 28 shows three times the different sub-divisions of the sub-blocks applied to the block 28. In particular, the quadtree sub-division is illustratively used in the case of FIG. 15a. That is, the set of possible sub-divisions of the sub-block is the quadtree sub-division or is defined thereby. And the three specific examples of the sub-division of the sub-block of the portion 28 represented in FIG. 15a belong to different hierarchical levels of the quadtree sub-division of the block 28. From bottom to top, the level or coarseness of the sub-division of the block 28 within the sub-block increases. At the highest level, the portion 28 remains as it is. At the next lower level, the block 28 is divided into four sub-blocks. And at least one of the latter is further divided into four sub-blocks at the next lower level and so on. In FIG. 15a, at each level, the quadtree sub-division is selected at the place where the number of sub-blocks is the smallest and yet not at the sub-block overlapping with the base layer coding parameter change boundary. That is, in the case of FIG. 15a, the quadtree sub-division of the block 28 to be selected for dividing the block 28 is recognized as the lowest among those shown in FIG. 15a. Here, the base layer coding parameters of the base layer are constant within each portion juxtaposed to each sub-block of the sub-division of the sub-block.

[0071] Therefore, the sub-division information for block 28 need not be signaled within the enhancement layer sub-stream 6b. As a result, the coding efficiency is increased. Moreover, as outlined, the method of obtaining a sub-division is appropriate regardless of the current position of section 28 with respect to any grid (lattice), or any alignment of the sample array of the base layer signal 200. In particular, the sub-division derivation works in the case of a fragmented spatial resolution ratio between the base layer and the enhancement layer.

[0072] Based on the sub-division of the sub-blocks of section 28 determined in this way, section 28 can be pre-predictively reconstructed / encoded. It should be noted that for the above description, different possibilities exist to "measure" the coarseness of the sub-division of the different available sub-blocks of the current block 28. For example, the size of the coarseness can be determined based on the number of sub-blocks. The more sub-blocks each sub-division of a sub-block has, the lower the level. This definition is clearly not applicable in the case of FIG. 15a where the "size of the coarseness" is determined by the combination of the number of sub-blocks of each sub-division of a sub-block and the smallest size of all the sub-blocks of each sub-division of a sub-block.

[0073] To be complete, FIG. 15b illustratively shows the case of selecting one possible sub-division of a sub-block from a set of sub-divisions of sub-blocks available for the current block 28 when using the sub-division of FIG. 35 as an example of an available set. Different hatching (and non-hatching) indicates regions in the base layer signal where the regions juxtaposed thereto have the same base layer coding parameters associated therewith.

[0074] As described above, the selection outlined traverses the sub - division of the possible sub - blocks according to a certain continuous order, such as in the order of increasing or decreasing levels of coarseness, and within each sub - block of the sub - division of each sub - block, it can be implemented by selecting the sub - division of the possible sub - blocks from the sub - division of the possible sub - blocks in a situation where the base - layer coding parameters are sufficiently similar to each other. (When using traversal according to increasing coarseness) It is no longer applicable. Or, (when using traversal according to decreasing levels of coarseness) it is applied incidentally at first. Alternatively, all possible sub - divisions can be tested.

[0075] In the descriptions of FIGS. 14, 15a, and 15b, although the broad term "base - layer coding parameters" is used in the preferred embodiments, these base - layer coding parameters represent base - layer prediction parameters, that is, parameters related to the formation of the prediction of the base - layer signal but not related to the formation of the prediction residual. Thus, for example, the base - layer coding parameters include a prediction mode that distinguishes between spatial prediction and temporal prediction, prediction parameters for blocks / parts of the base - layer signal assigned to spatial prediction such as spatial prediction in the angular direction, and prediction parameters for blocks / parts of the base - layer signal assigned to temporal prediction such as motion parameters.

[0076] However, interestingly, within a given sub - block, the definition of "sufficient" similarity of the base - layer coding parameters only determines / defines a subset of the base - layer coding parameters. For example, the similarity can be determined based on the prediction mode only. Or also, prediction parameters that further adjust spatial prediction and / or temporal prediction can form parameters on which the similarity of the base - layer coding parameters within a given sub - block depends.

[0077] Furthermore, as already outlined, in order to be sufficiently similar to each other, within a given sub-block, the base layer coding parameters need to be exactly equal to each other within each respective sub-block. Alternatively, the degree of similarity used may need to be within a given range of intervals in order to meet the "similarity" criterion.

[0078] As outlined above, the sub-division of the selected sub-block is not only in terms of amounts that can be predicted from, or transferred from, the base layer signal. Rather, the base layer coding parameters themselves are transferred to the enhancement layer signal in order to derive, based thereon, enhancement layer coding parameters for the sub-blocks of the sub-division of the sub-block obtained by transferring the sub-division of the selected sub-block from the base layer signal to the enhancement layer signal. As far as motion parameters are concerned, for example, scaling can be used for transfer from the base layer to the enhancement layer. Preferably, only those parts or syntax elements of the prediction parameters of the base layer are used to set the sub-blocks of the sub-division of the current part of the base layer that affect the degree of similarity. By this degree, the fact that these syntax elements of the prediction parameters within each sub-block of the sub-division of the selected sub-block are somehow similar to each other ensures that the syntax elements of the base layer prediction parameters used to predict the corresponding prediction parameters of the sub-blocks of the current part 308 are similar or even equal to each other. As a result, in the first case allowing for some variations, some important "meanings" of the syntax elements of the base layer prediction parameters corresponding to the parts of the base layer signal covered by each sub-block can be used as predictors for the corresponding sub-blocks. However, also, only the part of the syntax elements contributing to the degree of similarity, while the mode-specific base layer prediction parameters participate in determining the degree of similarity, may be used to predict the prediction parameters of the sub-blocks of the sub-division of the enhancement layer by simply adding the transfer of the sub-division itself so as to merely infer or pre-set the mode of the sub-blocks of the current part 28.

[0079] One such possibility of not using only sub-partitioned inter-layer prediction from the base layer to the enhancement layer is now explained with respect to the following figure (Figure 16). Figure 16 shows an image 22b of the enhancement layer signal 400 and an image 22a of the base layer signal 200 in a registered manner along the presentation time axis 24.

[0080] According to the embodiment of FIG. 16, the base layer signal 200 is predictively reconstructed by a scalable video decoder by sub-dividing the image 22a of the base layer signal 200 into sub-blocks within an intra-block and an inter-block, and is predictively encoded by the use of a scalable video encoder. According to the example of FIG. 16, the latter sub-division is made by a two-stage method. First, the frame 22a is normally sub-divided along its periphery into the largest block or the largest coding unit, indicated by reference numeral 302 in FIG. 16, using a double line. Then, each of the largest blocks 302 is made subordinate to a hierarchical quadtree sub-division within the coding units forming the aforementioned intra-block and inter-block. As a result, they are the leaves of the quadtree sub-division of the largest block 302. In FIG. 16, reference numeral 304 is used to indicate these leaf blocks or coding units. Usually, a solid line is used to indicate the periphery of these coding units. On the other hand, spatial intra prediction is used for intra-blocks, and temporal inter prediction is used for inter-blocks. Prediction parameters related to spatial intra prediction and temporal inter prediction are respectively set within units of smaller blocks. However, the intra- and inter-blocks or coding units 304 are sub-divided. Such a sub-division is exemplarily shown in FIG. 16 for one of the coding units 304 using reference numeral 306 to indicate smaller blocks. The smaller blocks 304 are outlined using a dotted line. That is, in the case of the embodiment of FIG. 16, the spatial video encoder has an opportunity to select between one spatial prediction and the other temporal prediction for each coding unit 304 of the base layer. However, as far as the enhancement layer signal is concerned, the degree of freedom increases. Here, in particular, the frame 22b of the enhancement layer signal 400 is assigned to each one of a set of prediction modes including not only spatial intra prediction and temporal inter prediction but also inter-layer prediction as outlined in detail below within the coding unit in which the frame 22b of the enhancement layer signal 400 is sub-divided.The sub - division within these coding units can be done in the same way as described for the base - layer signal. First, frame 22b is normally sub - divided into rows and columns of the largest block whose contour is drawn using double lines, and then, using normal solid lines, it can be sub - divided within the coding unit whose outer shape is formed during the hierarchical quadtree sub - division process.

[0081] One coding unit 308 of the current image 22b of the enhancement - layer signal 400 is exemplarily inferred to be assigned to the inter - layer prediction mode and is shown using diagonal lines. In a similar way to FIGS. 14, 15a, and 15b, FIG. 16 shows at 312 how the sub - division of the coding unit 308 is predictively derived by local transfer from the base - layer signal. In particular, the local region superimposed by the coding unit 308 is shown at 312. Within this region, the dotted lines indicate the boundaries between adjacent blocks of the base - layer signal, or more generally the boundaries where the base - layer coding parameters of the base - layer probably change. As a result, these boundaries are the boundaries of the prediction blocks 306 of the base - layer signal 200 and can partially coincide with the boundaries between adjacent coding units 304 of the base - layer signal 200, or equally, between the largest adjacent coding units 302. The dotted lines in 312 indicate the sub - division of the current coding unit 308 into the prediction blocks derived / selected by local transfer from the base - layer signal 200. Details regarding local transfer have been described above.

[0082] According to the embodiment of FIG. 16, as already described, not only the sub - division into the prediction blocks but also the base - layer is adopted. Rather, the prediction parameters of the base - layer signal used within the region 312 are used to derive the prediction parameters to be used for performing prediction regarding the prediction blocks of the coding unit 308 of the enhancement - layer signal 400.

[0083] In particular, according to the embodiment of FIG. 16, not only is the sub-division into the prediction block derived from the base layer signal, but also the prediction mode is used within the base layer signal 200 to encode / reconstruct each region locally covered by each sub-block of the derived sub-division. One example is as follows. To derive the sub-division of the coding unit 308 as described above, the prediction mode is used in connection with the relevant base layer signal 200. Mode-specific prediction parameters can be used to determine the "similarity" discussed above. Thus, the different hatches shown in FIG. 16 may correspond to different prediction blocks 306 of the base layer. Each of the different prediction blocks 306 may have an intra or inter prediction mode, i.e., a spatial or temporal prediction mode associated therewith. As explained above, for it to be "sufficiently similar", the prediction mode used within the region juxtaposed to each sub-block of the sub-division of the coding unit 308 and the specific prediction parameters for each prediction mode within the sub-area must be exactly equal to each other. Alternatively, it may also tolerate some variation.

[0084] In particular, according to the embodiment of FIG. 16, all blocks indicated by the hatches extending from top left to bottom right can be set as intra prediction blocks of the coding unit 308 because the locally corresponding portions of the base layer signal are covered by the prediction blocks 306 having a spatial intra prediction mode associated therewith. On the other hand, the other blocks (i.e., the blocks indicated by the hatches extending from bottom left to top right) can be set as inter prediction blocks because the locally corresponding portions of the base layer signal are covered by the prediction blocks 306 having a temporal inter prediction mode associated therewith.

[0085] On the other hand, for an alternative embodiment, the derivation of the prediction stops here within the coding unit 308 where the details for performing the prediction are. That is, the derivation of the sub-division of the coding unit 308 within the prediction block and the assignment of these prediction blocks within the prediction blocks coded using non-temporal prediction or spatial prediction and within the prediction blocks coded using temporal prediction can be restricted, which does not follow the embodiment of FIG. 16.

[0086] According to the latter embodiment, all prediction blocks of the coding unit 308 having a non-temporal prediction mode assigned thereto receive non-temporal prediction such as spatial intra prediction while using prediction parameters derived from the prediction parameters of the locally matching intra-blocks of the base layer signal 200 as the enhancement layer prediction parameters of these non-temporal mode blocks. As a result, such a derivation may be related to the spatial prediction parameters of the intra-blocks locally juxtaposed of the base layer signal 200. For example, such spatial prediction parameters may be an indication of the angular direction in which the spatial prediction is performed. As outlined above, a definition of similarity by itself is required in either case where the spatial base layer prediction parameters overlap by being the same for each non-temporal prediction block of the coding unit 308, or for each non-temporal prediction block of the coding unit 308, some average of the spatial base layer prediction parameters overlaps by being used for each non-temporal prediction block to derive the prediction parameters of each non-temporal prediction block.

[0087] Alternatively, all prediction blocks of the coding unit 308 having an assigned non-temporal prediction mode may receive inter-layer prediction in the following manner. First, the base layer signal undergoes decomposition or quality improvement within those regions spatially juxtaposed at least to the non-temporal prediction mode prediction blocks of the coding unit 308 to obtain an inter-layer prediction signal. Then, next, these prediction blocks of the coding unit 308 are predicted using the inter-layer prediction signal.

[0088] Scalable video decoders and encoders may, by default, subject all of the coding units 308 to spatial prediction or inter-layer prediction. Alternatively, the scalable video encoder / decoder may support both alternatives and signal them within the coded video data stream signal. That version is used as long as it pertains to non-temporal prediction mode prediction blocks of the coding unit 308. In particular, the decision between both alternatives may be signaled within the data stream for any size of coding unit 308, for example, individually.

[0089] As long as it pertains to another prediction block of the coding unit 308, the coding unit 308 receives temporal inter prediction using prediction parameters derived from the prediction parameters of the inter-block that locally match as if it were a non-temporal prediction mode prediction block. As a result, the derivation is related, in turn, to the motion vectors assigned to the corresponding parts of the base layer signal.

[0090] For all other coding units that have both a spatial intra prediction mode and a temporal inter prediction mode assigned to them, another coding unit receives spatial prediction or temporal prediction in the following way. In particular, another coding unit is further sub-divided into prediction blocks that have the prediction mode assigned to it. The prediction mode is common to all of the prediction blocks within the coding unit and, in particular, is the same prediction mode assigned to each coding unit. That is, unlike coding units associated with an inter-layer prediction mode such as coding unit 308, coding units associated with a spatial intra prediction mode, or coding units associated with a temporal inter prediction mode, are sub-divided into prediction blocks that only have the prediction mode, that is, the prediction mode inherited from each coding unit derived by the sub-division of each coding unit.

[0091] The sub-division of all coding units including 308 may be a quadtree sub-division within the prediction block.

[0092] A further difference between an encoding unit in an inter-layer prediction mode such as the symbolization unit 308 and an encoding unit in a spatial intra prediction mode or a temporal inter prediction mode is when the prediction blocks of the spatial intra prediction mode encoding unit or the temporal inter prediction mode encoding unit are respectively subjected to spatial prediction and temporal prediction. The prediction parameters are set, for example, by a method of signaling within the enhancement layer substream 6b, without depending on the base layer signal 200 or the like. Even a sub-division of other encoding units having an inter-layer prediction mode related thereto such as the encoding unit 308 can be signaled within the enhancement layer signal 6b. That is, an inter-layer prediction mode encoding unit such as 308 has the advantage of the need for low-bitrate signaling. According to an embodiment, the mode indicator of the encoding unit 308 itself does not need to be signaled within the enhancement layer substream. Optionally, another parameter can be transmitted for the encoding unit 308, such as a prediction parameter residual, for an individual prediction block. Additionally, or alternatively, a prediction residual for the encoding unit 308 can be transmitted / signaled within the enhancement layer substream 6b. On the other hand, a scalable video decoder searches for this information from the enhancement layer substream, and a scalable video encoder according to the current embodiment determines these parameters and inserts these parameters into the enhancement layer substream 6b.

[0093] In other words, the prediction of the base layer signal 200 can be made using base layer encoding parameters in such a way that the base layer encoding parameters spatially vary the base layer signal 200 within the unit of the base layer block 304. The prediction modes available for the base layer may include, for example, spatial and temporal prediction. The base layer encoding parameters may further include individual prediction parameters of the prediction modes such as, as long as they are related to the spatially predicted block 304 in the angular direction, and as long as they are related to the motion vector and the temporally predicted block 304. The individual prediction parameters of the latter prediction mode can vary the base layer signal within a unit smaller than the base layer block 304, i.e., within the aforementioned prediction block 306. In order to satisfy the requirements outlined before sufficient similarity, it may be necessary that the prediction modes of all base layer blocks 304 whose sub-division regions of each possible sub-block overlap with each other are equal to each other. And only the sub-division of each sub-block is put into the selection candidate list to obtain the sub-division of the selected sub-block. However, the requirements are even stricter. It is that the individual prediction parameters of the prediction modes of the prediction blocks, whose common regions of the sub-division of each sub-block overlap with each other, must also be equal to each other. For each sub-division of each sub-block and the corresponding sub-block of the region within the base layer signal, only the sub-division of the sub-block that satisfies this requirement can be put into the selection candidate list to obtain the sub-division of the finally selected sub-block.

[0094] In particular, as briefly outlined above, there are various possibilities for how to perform the selection within a set of possible sub-block divisions. To outline this in more detail, reference is made to FIGS. 15c and 15d. Assume that set 352 encloses sub-divisions 354 of all possible sub-blocks of current block 28. Of course, FIG. 15c shows merely an example. The set 352 of possible or available sub-divisions of the current block 28 may be known to the scalable video decoder and the scalable video encoder by default, or may be signaled within an encoded data stream such as a series of images or the like. According to the example of FIG. 15c, each member of set 352, i.e., each available sub-division 354 of a sub-block, is subject to a check 356 to check whether regions within the juxtaposed portion 108 of the base layer signal to be sub-divided, by transferring each sub-division 354 of the sub-block from the enhancement layer to the base layer, are simply overlapped by prediction blocks 306 and coding units 304. And then, it is checked whether the base layer coding parameters meet the requirement of sufficient similarity. For example, refer to the exemplary sub-division with reference number 354. According to this exemplary available sub-division of the sub-block, the current block 28 is sub-divided into four quadrants / sub-blocks 358. And the upper left sub-block corresponds to region 362 within the base layer. Obviously, this region 362 does not correspond to another sub-division within the four blocks of the base layer, i.e., the prediction blocks, and as a result, represents the prediction blocks themselves, overlapping two prediction blocks 306 and two coding units 304. Thus, if all the base layer coding parameters of these prediction blocks overlapping region 362 meet the similarity criterion, and this is also the case for all sub-blocks / quadrants of the possible sub-division 354 of the sub-block and their corresponding regions in the base layer with overlapping coding parameters, then this possible sub-division 354 of the sub-block meets the sufficient requirements for all regions covered by the sub-blocks of the sub-division of the sub-block and belongs to the set 364 of sub-divisions of the sub-block.Then, within this set 364, the coarsest sub-division is selected as indicated by arrow 366, and as a result, from set 352, a sub-division 368 of the selected sub-block is obtained.

[0095] Obviously, it is preferable to avoid performing check 356 for all members of set 352. Thus, as shown in FIG. 15d and as described above, the possible sub-divisions 354 can be traversed to increase or decrease their coarseness. The traversal is indicated using the double-headed arrow 372. FIG. 15d shows that for at least some of the sub-divisions of the available sub-blocks, the levels or scales of coarseness are equal to each other. In other words, the ordering according to the increasing or decreasing levels of coarseness can be ambiguous. However, since only one of such equally coarse possible sub-divisions of the sub-block can belong to set 364, this does not prevent the search for the "coarsest sub-division of the sub-block" belonging to set 364. Thus, when traversing in the direction of the level where the coarseness increases, among the sub-divisions of the possible sub-blocks traversed from the second to the last, which are the sub-divisions 354 of the sub-blocks to be selected, as soon as the result of the reference check 356 changes from being filled to not being filled, the largest possible sub-division 368 of the sub-block can be found. Or, when traversing in the direction of the level where the coarseness decreases, among the sub-divisions of the sub-blocks where most of the sub-division 368 of the sub-block has been recently traversed, as soon as the result of the reference check 356 switches from not being filled to being filled, the coarsest possible sub-division 368 of the sub-block is found.

[0096] For the following figures, it is described how a scalable video encoder or decoder as described above with respect to FIGS. 1 to 4 is implemented to form an embodiment of the present application according to another embodiment of the present application. Possible implementation forms of the embodiments described below are presented below with reference to embodiments K, A, and M.

[0097] To describe an aspect, refer to FIG. 17. FIG. 17 shows the possibilities for the time prediction 32 of the current portion 28. As a result, the following description of FIG. 17 can be combined with the description regarding FIGS. 6 to 10 as long as it is related to the combination with the inter-layer prediction signal. Alternatively, it can be combined with the description regarding FIGS. 11 to 13 as long as it is related to the combination with the time inter-layer prediction mode.

[0098] The situation shown in FIG. 17 corresponds to the situation shown in FIG. 6. That is, the base layer signal 200 and the enhancement layer signal 400 are shown together with the already encoded / decoded portions indicated using diagonal lines. Within the enhancement layer signal 400, the portion to be currently encoded / decoded here has, by way of example, adjacent blocks 92 and 94 described as the block 92 above and the block 94 to the left of the current portion 28. Both blocks 92 and 94 have, by way of example, the same size as the current block 28. However, the size match is not essential. Rather, the portions of the blocks within the image 22b of the enhancement layer signal 400 that are sub-divided can have different sizes. They are not even restricted to being rectangular. They can be rectangular, or other shapes. The current block 28 having other adjacent blocks is not explicitly described in FIG. 17. However, the other adjacent blocks have not yet been encoded / decoded. That is, they follow in the order of encoding / decoding and as a result are not available for prediction. In addition to this, there are blocks other than the blocks 92 and 94 that have already been encoded / decoded in accordance with the encoding / decoding order, by way of example, the block 96 diagonally above and to the left of the current block 28, adjacent to the current block 28. However, in the example considered here, the blocks 92 and 94 serve to predict the inter-prediction parameters for the current block 28 that undergoes inter-prediction 30, and the adjacent blocks are predetermined. The number of such predetermined adjacent blocks is not restricted to two. It can be greater than 1 or simply 1. Discussion of possible embodiments is presented with respect to FIGS. 36 to 38.

[0099] A scalable video encoder and a scalable video decoder can determine a set of predetermined adjacent blocks (here, blocks 92 and 94) from a set of already encoded adjacent blocks (here, blocks 92 to 96) that depend on, for example, a predetermined sample position 98 within the current portion 28 such as the sample in the upper left. For example, only those already encoded adjacent blocks of the current portion 28 form a set of "predetermined adjacent blocks" that includes sample positions directly adjacent to the predetermined sample position 98. Further possibilities can be described with respect to FIGS. 36 to 38.

[0100] In any case, in accordance with the decoding / encoding order, a portion 502 of the previously encoded / decoded image 22b of the enhancement layer signal 400, which has been replaced from the collocated position of the current block 28 by the motion vector 504, includes sample values reconstructed based on the sample values of the portion 28 that can be predicted by mere copying or interpolation. For this purpose, the motion vector 504 is signaled within the enhancement layer substream 6b. For example, the temporal prediction parameter for the current block 28 indicates a displacement vector 506 that shows the replacement of the portion 502 from the collocated position of the portion 28 within the reference image 22b by interpolation, optionally, for being copied onto the samples of the portion 28.

[0101] In any case, when temporally predicting the current block 28, the scalable video decoder / encoder has already reconstructed (and in the case of the encoder, already encoded) the base layer 200 using the base layer substream 6a. At least as long as the relevant spatially corresponding region of the temporally corresponding image 22a is so related, block-based prediction is used as described above, and, for example, block-based selection between the spatial prediction mode and the temporal prediction mode is used.

[0102] In FIG. 17, the time when the image 22a of the base layer signal 200 is juxtaposed is sub-divided into several blocks 104. The block 104 is located around in a region that locally corresponds to the current portion 28, which is illustratively represented. Similar to the case of having spatially predicted blocks in the enhancement layer signal 400, the spatial prediction parameter is signaled for the selection of the spatial prediction mode. Whether they are included in the base layer sub-stream 6a or signaled for those blocks 104 in the base layer signal 200.

[0103] Here, illustratively, for the block 28 where the temporal intra-layer prediction 32 is selected, in order to enable the reconstruction of the enhancement layer signal from the coded data stream, the inter-prediction parameters such as the motion parameters are determined using any of the following methods.

[0104] The first possibility is described with respect to FIG. 18. In particular, first, a set 512 of motion parameter candidates 514 is collected or generated from adjacent already reconstructed blocks of the frame such as the pre-determined blocks 92 and 94. The motion parameter can be a motion vector. The motion vectors of the blocks 92 and 94 are represented using the arrows 516 and 518 marked with 1 and 2 respectively (therein). As shown in the figure, these motion parameters 516 and 518 can directly form the candidates 514. Some candidates can be formed by combining motion vectors such as 518 and 516 as shown in FIG. 18.

[0105] Furthermore, a set 522 of one or more base layer motion parameters 524 of the block 108 of the base layer signal 200 juxtaposed to the portion 28 is collected or generated from the base layer motion parameters. In other words, the motion parameters related to the block 108 juxtaposed in the base layer are used to derive one or more base layer motion parameters 524.

[0106] At that time, one or more base layer motion parameters 524, or a scaled version thereof, are added 526 to the set 512 of motion parameter candidates 514 to obtain an extended motion parameter candidate set 528 of motion parameter candidates. This can be done in various ways such as simply adding the base layer motion parameter 524 at the end of the list of candidates 514, or in a different way outlined for FIG. 19a.

[0107] Next, at least one of the motion parameter candidates 532 of the extended motion parameter candidate set 528 is selected. A temporal prediction 32 is then performed using the selected one of the motion parameter candidates of the extended motion parameter candidate set by motion compensation prediction of part 28. The selection 534 can be signaled in a data stream such as substream 6b for part 28 by way of an index 536 within the list / set 528, or can be performed in another way as described for FIG. 19a.

[0108] As described above, it can be checked whether the base layer motion parameter 523 has been encoded in an encoded data stream such as base layer substream 6a using merge. And if, hypothetically, the base layer motion parameter 523 is encoded in the encoded data stream using merge, the addition 526 can be suppressed.

[0109] The motion parameters described according to FIG. 18 can relate to only the motion vectors (motion vector prediction), or to a complete set of motion parameters including the number of motion hypotheses for each block, reference index, partitioning information (merge). Thus, the "scaled version" can result from the scaling of the motion parameters used within the base layer signal according to the spatial resolution ratio between the base layer signal and the enhancement layer signal in the case of spatial scalability. Depending on the method of the encoded data stream, the encoding / decoding of the base layer motion parameters of the base layer signal can be involved in, for example, spatial or temporal motion vector prediction, or merging.

[0110] The incorporation 526 of the motion parameters 523 used in the collocated portion 108 of the base layer signal into the set 528 of merge / motion vector candidates 532 enables a very effective indexing among the intra layer candidates 514 and one or more inter layer candidates 524. The selection 534 may involve the explicit signaling of an index into an extended set / list of motion parameter candidates within the enhancement layer signal 6b, for each prediction block, or for each coding unit or the like. Alternatively, the selected index 536 may also be inferred from other information in the enhancement layer signal 6b, or from inter layer information.

[0111] According to the possibility of FIG. 19a, the formation 542 of the final motion parameter candidate list for the enhancement layer signal for the portion 28 is only optionally performed, as outlined with respect to FIG. 18. That is, the formation 542 may also be 528 or 512. However, the lists 528 / 512 are ordered 544 depending on, for example, the base layer motion parameters such as the motion parameters represented by the motion vectors 523 of the collocated base layer blocks 108. For example, the rank of a member, i.e., a motion parameter candidate 532 or 514 of the list 528 / 512, is determined based on the deviation of each member with respect to a potentially scaled version of the motion parameter 523. The greater the deviation, the lower the rank of each member 532 / 512 within the ordered list 528 / 512'. As a result, the ordering 544 involves determining the magnitude of the deviation for each member 532 / 514 of the lists 528 / 512. The selection 534 of one candidate 532 / 512 within the ordered list 528 / 512' is performed and controlled via an explicitly signaled index syntax element 536 in the coded data stream in order to obtain the enhancement layer motion parameters from the ordered motion parameter candidate list 528 / 512' for the portion 28 of the enhancement layer signal. Then, in turn, the temporal prediction 32 is performed using the selected motion parameters, where the index 536 points to 534, by motion compensation prediction of the portion 28 of the enhancement layer signal.

[0112] For the motion parameters referred to in FIG. 19a, the motion parameters described above with respect to FIG. 18 are applied. Decoding of the base layer motion parameters 520 from the coded data stream may (optionally) involve spatial or temporal motion vector prediction, or merge. Ordering may be made according to the magnitude of the difference between each enhancement layer motion parameter candidate and the base layer motion parameter of the base layer signal, in relation to the block of the base layer signal juxtaposed to the current block of the enhancement layer signal. That is, for the current block of the enhancement layer signal, a list of enhancement layer motion parameter candidates may first be determined. Next, it is described that ordering is performed. Below, the selection is performed with explicit signaling.

[0113] Alternatively, the ordering 544 may be made according to a magnitude that measures a difference between a base layer motion parameter 523 of a base layer signal associated with a block 108 of the base layer signal juxtaposed to the current block 28 of the enhancement layer signal and a base layer motion parameter 546 of a spatially and / or temporally adjacent block 548 within the base layer. Next, the determined ordering within the base layer is transferred to the enhancement layer. As a result, the enhancement layer motion parameter candidates are ordered such that they have the same ordering as the determined ordering with respect to the corresponding base layer candidates. In this regard, when the associated base layer block 548 is spatially / temporally juxtaposed to adjacent enhancement layer blocks 92 and 94 associated with the considered enhancement layer motion parameter, the base layer motion parameter 546 may be said to correspond to the enhancement layer motion parameters of the adjacent enhancement layer blocks 92, 94. Alternatively, when the adjacency relationship (left adjacent, upper adjacent, A1, A2, B1, B2, B0, or refer to FIGS. 36 to 38 for further examples) between the associated base layer block 548 and the block 108 juxtaposed to the current enhancement layer block 28 is the same as the adjacency relationship between the current enhancement layer block 28 and the adjacent enhancement layer blocks 92, 94 respectively, the base layer motion parameter 546 may be said to correspond to the enhancement layer motion parameters of the adjacent enhancement layer blocks 92, 94. Based on the base layer ordering, the selection 534 is then performed by explicit signaling.

[0114] To explain this in more detail, refer to FIG. 19b. FIG. 19b shows the first of the alternative outlines for just obtaining an enhancement layer that is ordered for a list of motion parameter candidates using base layer hints. FIG. 19b shows the current block 28 and the positions of three different pre-determined samples, namely, by way of example, the upper left sample 581, the lower left sample 583, and the upper right sample 585. The examples are to be construed as illustrative only. Assume that there are four types of adjacencies including adjacent blocks 94a that cover, or include, sample position 587 that is adjacent to and above sample position 581 and adjacent blocks 94b that cover, or include, sample position 589 that is adjacent to and directly above sample position 585. Similarly, adjacent blocks 92a and 92b are those blocks that include sample positions 591 and 593 that are adjacent to and to the immediate left of sample positions 581 and 583. Also, note that, as described with respect to FIGS. 36-38, the number of pre-determined adjacent blocks can vary despite a pre-determined number of decision rules. Nevertheless, the pre-determined adjacent blocks 92a, 92b, 94a, 94b are distinguishable by their decision rules.

[0115] According to an alternative to FIG. 19b, juxtaposed blocks within the base layer are determined for each of the predefined adjacent blocks 92a, 92b, 94a, 94b. For example, for this purpose, the upper left samples 595 of each adjacent block are used. Just as it is the case with the upper left sample 581 formally mentioned in FIG. 19a, when having the current block 28. This is illustrated in FIG. 19b using dotted arrows. By this means, for each of the predefined adjacent blocks, a corresponding block 597 is found in addition to the juxtaposed block 108 and the juxtaposed current block 28. Using the motion parameters m1, m2, m3, m4 of the juxtaposed base layer block 597 and their respective differences with respect to the base layer motion parameter m of the juxtaposed base layer block 108, the enhancement layer motion parameters M1, M2, M3, M4 of the predefined adjacent blocks 92a, 92b, 94a, 94b can be ordered within list 528 or 512. For example, the larger any of the distances of m1 to m4, the higher the corresponding enhancement layer motion parameters M1 to M4. That is, a higher index may be required to index from list 528 / 512´ in the same state. The absolute difference is used with respect to the magnitude of the distance. Similarly, the motion parameter candidates 532 or 514 can be rearranged within the list with respect to their ranks which are a combination of the enhancement layer motion parameters M1 to M4.

[0116] FIG. 19c shows an alternative in which the corresponding blocks in the base layer are determined in a different way. In particular, FIG. 19c shows the predefined adjacent blocks 92a, 92b, 94a, 94b of the current block 28, and the juxtaposed block 108 of the current block 28. According to the embodiment of FIG. 19c, the corresponding base layer blocks of the current block 28, namely 92a, 92b, 94a, 94b, are determined in such a way that these base layer blocks are associated with the enhancement layer adjacent blocks 92a, 92b, 94a, 94b using the same adjacency determination rules for determining these base layer adjacent blocks. In particular, FIG. 19c shows the predefined sample positions of the juxtaposed block 108, namely the upper left, lower left, and upper right sample positions 601. Based on these sample positions, the four adjacent blocks of block 108 are determined in the same way as described for the enhancement layer adjacent blocks 92a, 92b, 94a, 94b with respect to the predefined sample positions 581, 583, 585 of the current block 28. The four base layer adjacent blocks 603a, 603b, 605a, 605b are found in this way. 603a clearly corresponds to the enhancement layer adjacent block 92a. The base layer block 603b corresponds to the enhancement layer adjacent block 92b. The base layer block 605a corresponds to the enhancement layer adjacent block 94a. The base layer block 605b corresponds to the enhancement layer adjacent block 94b. In the same way as previously described, the base layer motion parameters M1 to M4 of the base layer blocks 903a, 903b, 905a, 905b and their distances to the base layer motion parameter m of the juxtaposed base layer block 108 are used to order the motion parameter candidates within the list 528 / 512 formed from the motion parameters M1 to M4 of the enhancement layer blocks 92a, 92b, 94a, 94b.

[0117] According to the possibilities of FIG. 20, the formation 562 of the final motion parameter candidate list for the enhancement layer signal for part 28 is optionally performed simply as outlined for FIGS. 18 and / or 19. That is, the formation 562 is 528 or 512 or 528 / 512´. Reference numeral 564 may be used in FIG. 20. According to FIG. 20, the index 566 pointed out within the motion parameter candidate list 564 is determined, for example, for the juxtaposed block 108, depending on the index 567 within the motion parameter candidate list 568 used for encoding / decoding the base layer signal. For example, when reconstructing the base layer signal at block 108, the list 568 of motion parameter candidates is for a block 108 having the same adjacency relationship as the predetermined adjacency relationship between the adjacent enhancement layer blocks 92, 94 and the current block 28, and may be determined based on the motion parameters 548 of the adjacent block 548 of the block 108 having an adjacency relationship (adjacent to the left, adjacent to the top, A1, A2, B1, B2, B0, or for yet another example, see FIGS. 36 - 38). Here, the determination 572 of the list 567 potentially uses the same configuration rules as those used within the formation 562 such as the ordering within the list members of the lists 568 and 564. More generally, the index 566 for the enhancement layer can be determined in such a way that the index 566 points to the index 566 juxtaposed with the base layer block 548 related to the indexed base layer candidate, i.e., what the index 567 points to, and to which the adjacent enhancement layer blocks 92, 94 are related. As a result, the index 567 can function as an important prediction of the index 566. The enhancement layer motion parameters are then determined using the index 566 into the motion parameter candidate list 564, and the motion compensation prediction for block 28 is performed using the determined motion parameters.

[0118] Regarding the motion parameters mentioned in FIG. 20, the same applies as described above with respect to FIGS. 18 and 19.

[0119] Regarding the following figures, as described above for FIGS. 1-4, how a scalable video encoder or decoder can be implemented to form embodiments of the present application according to another aspect of the application is described. The detailed implementation of the aspects described below is described with reference to Example V below.

[0120] This aspect relates to residual coding within an enhancement layer. In particular, FIG. 21 illustratively shows an image 22b of an enhancement layer signal 400 and an image 22a of a base layer signal 200 in a temporally registered manner. FIG. 21 shows a method of reconstructing within a scalable video decoder or encoding within a scalable video encoder, and focuses on a predetermined transform coefficient block of transform coefficients 402 representing the enhancement layer signal 400 and a predetermined portion 404. In other words, the transform coefficient block 402 represents the spatial decomposition of the portion 404 of the enhancement layer signal 400. As already explained above according to the encoding / decoding ordering, the corresponding portion 406 of the base layer signal 200 can already be decoded / encoded when decoding / encoding the transform coefficient block 402. As far as the base layer signal 200 is concerned, predictive encoding / decoding can be used, including signaling of the base layer residual signal within an encoded data stream such as the base layer substream 6a.

[0121] According to the embodiment described with respect to FIG. 21, the scalable video decoder / encoder utilizes the fact that the evaluation 408 of the base layer signal or the base layer residual signal can result in an advantageous selection of the sub-division of the transform coefficient block 402 within the sub-block 412 at the portion 406 collocated with the portion 404. In particular, several possible sub-divisions of the transform coefficient block 402 into sub-blocks can be supported by the scalable video decoder / encoder. These possible sub-divisions of the sub-blocks can regularly sub-divide the transform coefficient block 402 into rectangular sub-blocks 412. That is, the transform coefficients 414 of the transform coefficient block 402 can be arranged in columns and rows, and according to the possible sub-divisions of the sub-blocks, these transform coefficients 414 are regularly packed within the sub-block 412, so that the sub-blocks 412 themselves are arranged in rows and columns. The evaluation 408 uses the sub-division of the sub-blocks thus selected to enable the encoding of the transform coefficient block 402 in such a way that the ratio between the number of rows and the number of columns of the sub-block 412, that is, the ratio between their width and height, is set. If, for example, the evaluation 408 determines that the reconstructed base layer signal 200 within the collocated portion 406, or at least the base layer residual signal within the corresponding portion 406, is mainly composed of horizontal edges in the spatial domain, then the transform coefficient block 402 has significance, that is, the transform coefficient level is non-zero, that is, the quantized transform coefficient is near the zero horizontal frequency side of the transform coefficient block 402, and probably exists. In the case of vertical edges, the transform coefficient block 402 has a non-zero transform coefficient level at a position near the zero vertical frequency side of the transform coefficient block 402 and probably exists. Therefore, first, the sub-block 412 is selected to be longer along the vertical direction and smaller along the horizontal direction. And second, the sub-block is made longer along the horizontal direction and smaller along the vertical direction. The latter case is schematically shown in FIG. 40.

[0122] That is, the scalable video decoder / encoder selects one sub-block sub-division within a set of possible sub-block sub-divisions based on the base layer residual signal or the base layer signal. Then, the encoding 414 or decoding of the transform coefficient block 402 is performed while applying the selected sub-block sub-division. In particular, since the positions of the transform coefficients 414 are traversed in units of sub-blocks 412, all positions in one sub-block are traversed in a manner that immediately follows the next sub-block in the sub-block ordering defined in the sub-block. For a currently visited sub-block, such as the sub-block 412 exemplarily shown at 22 in FIG. 40, a syntax element is signaled in a data stream, such as the enhancement layer sub-stream 6b, indicating whether the currently visited sub-block has any significant transform coefficients. In FIG. 21, the syntax element 416 is described for two exemplary sub-blocks. If the respective syntax element of the respective sub-block indicates a non-significant transform coefficient, nothing else needs to be transmitted in the data stream or enhancement layer sub-stream 6b. Rather, the scalable video decoder may set the transform coefficients in that sub-block to zero. However, if the syntax element 416 of the respective sub-block indicates that this sub-block has significant transform coefficients, another information related to the transform coefficients in that sub-block is signaled in the data stream or sub-stream 6b. On the decoding side, the scalable video decoder decodes from the data stream or sub-stream 6b a syntax element 418 indicating the level of the transform coefficients in the respective sub-block. The syntax element 418 may indicate the position of the significant transform coefficients in that sub-block according to the scanning order of these transform coefficients in the respective sub-block and, optionally, the scanning order of the transform coefficients in the respective sub-block.

[0123] Figure 22 shows different possibilities that each exist to perform a selection within a sub - division of possible sub - blocks within evaluation 408. Figure 22 depicts again part 404 of the enhancement layer signal, which is related to the spectral decomposition of part 404 by transform coefficient block 402. For example, transform coefficient block 402 represents the spectral decomposition of the enhancement layer residual signal, with a scalable video decoder / encoder that predictively encodes / decodes the enhancement layer signal. In particular, the encoding / decoding of the transform is used by the scalable video decoder / encoder to encode the enhancement layer residual signal. The encoding / decoding of the transform is performed in a block - by - block manner, i.e., within the blocks into which the image 22b of the enhancement layer signal is sub - divided. Figure 22 shows the corresponding or juxtaposed part 406 of the base layer signal. Here, the scalable video decoder / encoder applies predictive encoding / decoding to the base layer signal while using the encoding / decoding of the transform for the prediction residual of the base layer signal, i.e., for the base layer residual signal. In particular, the block - by - block transform is used for the base layer residual signal. That is, the base layer residual signal is transformed block - by - block in the individually transformed blocks depicted by the dotted lines in Figure 22. As depicted in Figure 22, the block boundaries of the base layer transform blocks do not have to match the outer shape of the juxtaposed part 406.

[0124] Nevertheless, to perform evaluation 408, one or a combination of the following options A - C may be used.

[0125] In particular, the scalable video decoder / encoder may perform transform 422 on the base layer residual signal or the reconstructed base layer signal within part 406 to obtain a transform coefficient block 424 of transform coefficients that match in size the transform coefficient block 402 to be encoded / decoded. Examination of the distribution of the values of the transform coefficients within transform coefficient blocks 424, 426 can be used to appropriately set the dimensions of sub - blocks 412 along the horizontal frequency direction 428 and to appropriately set the dimensions of sub - blocks 412 along the vertical frequency direction 432.

[0126] Additionally, or alternatively, the scalable video decoder / encoder may examine all of the transform coefficient blocks of the base layer transform block 434 depicted by different hatching in FIG. 22 that at least partially overlap the juxtaposed portion 406. In the exemplary case of FIG. 22, there are four base layer transform blocks, and then their transform coefficient blocks are examined. In particular, all of these base layer transform blocks can be of different sizes from each other and can be of further different sizes with respect to the transform coefficient block 412. Scaling 436 may be performed on these transform coefficient blocks that overlap the base layer transform block 434 to provide an approximation of the transform coefficient block 438 of the spectral decomposition of the base layer residual signal within the portion 406. The distribution of the values of the transform coefficients within that transform coefficient block 438, i.e., 442, may be used within the evaluation 408 to appropriately set the dimensions 428 and 432 of the sub-blocks. As a result, the sub-division of the sub-blocks of the transform coefficient block 402 is selected.

[0127] A further alternative that may be used, additionally or alternatively, to perform the evaluation 408 is to examine the base layer residual signal or the reconstructed base layer signal within the spatial region using edge detection 444 or determination of the main gradient direction and, for example, determine based on the expansion direction of the detected edges or the determined gradient within the co-located portion 406 to appropriately set the sub-block dimensions 428 and 432.

[0128] Although not explicitly described above, in traversing the position of the conversion factor and the unit of sub-block 412, it is preferable to start from the zero-frequency corner of the conversion factor block, i.e., the upper left corner of FIG. 21, and traverse sub-block 412 in order to reach the highest-frequency corner of block 402, i.e., the lower right corner of FIG. 21. Further, entropy coding can be used to signal syntax elements within data stream 6b. That is, syntax elements 416 and 418 can be entropy codings that are convenient for encoding, such as arithmetic, variable-length coding, or another form of entropy coding. The ordering for traversing sub-block 412 can also depend on the sub-block shape selected according to 408. For sub-blocks selected to be wider than their height, the traversal order can cross the sub-blocks in the first column direction and then proceed to the next column, etc. Beyond this, it should be noted again that the base layer information used to select the dimensions of the sub-block can be the base layer residual signal or the reconstructed base layer signal itself.

[0129] In the following, different embodiments that can be combined with the aspects described above are described. The embodiments described below relate to many different aspects or means for making scalable video coding more efficient. In part, the above aspects are described in detail below while maintaining the general concept and presenting another derived embodiment thereof. The following description presented can be used to obtain alternatives or extensions of the above embodiments / aspects. However, most of the embodiments described below relate to sub-aspects and can optionally be combined with the aspects already described above. That is, they can be implemented together with the above embodiments within a single scalable video decoder / encoder. However, this is not essential.

[0130] To more easily understand the foregoing description, more detailed embodiments for implementing a suitable scalable video encoder / decoder incorporating embodiments or combinations of embodiments are presented next. The different examples described below are enumerated by the use of alphanumeric symbols. Explanations of some of these aspects refer, according to one embodiment, to the elements in the figures being described now, in which these aspects can be implemented in common. However, it should be noted that, as far as individual embodiments are concerned, the provision of each element within the implementation of the scalable video decoder / encoder is not necessary in all aspects. Depending on the aspect in question, some elements and some interconnections may be omitted in the drawings described next. Only the elements referred to for each aspect are provided to perform the work or function mentioned in the description of each aspect. However, in particular, when several elements are listed for one function, there may be alternatives.

[0131] However, to provide an overview of the functions of the scalable video decoder / encoder, the aspects described next are implemented. The elements shown in the following figures are now briefly described.

[0132] FIG. 23 shows a scalable video decoder for decoding an encoded data stream 6 in which the main subpart (i.e., 6a) of the encoded data stream 6 represents video at a first resolution or quality level. An additional part 6b of the encoded data stream corresponds to the representation of the video at increasing resolution or quality levels. To keep the data volume of the encoded data stream 6 low, the inter-layer redundancy between sub-streams 6a and 6b is utilized in forming sub-stream 6b. Some of the aspects described below are directed to inter-layer prediction from the associated base layer for sub-stream 6a and to the associated enhancement layer for sub-stream 6b.

[0133] The scalable video decoder includes two block-based prediction decoders 80, 60 operating in parallel, and receives sub-streams 6a and 6b respectively. As shown in the figure, the demultiplexer 40 can provide the decoding stages 80 and 60 separately, together with the corresponding sub-streams 6a and 6b.

[0134] The internal structures of the block-based prediction coding stages 80 and 60 can be the same as shown in the figure. From the inputs of the respective decoding stages 80, 60, the entropy coding modules 100; 320, the inverse transformers 560; 580, the adders 180; 340, the optional filters 120; 300 and 140; 280 are connected in series in this order of description. As a result, at the end of this series connection, the reconstructed base layer signal 600 and the reconstructed enhancement layer signal 360 can be derived respectively. On the other hand, the outputs of the adders 180, 340 and the filters 120, 140, 300, 280 provide different versions of the reconstruction of the base layer signal and the enhancement layer signal respectively. Then, each prediction provider 160; 260 receives a subset or all of these versions and provides, based on them, prediction signals to the residual inputs of the adders 180; 340 respectively. The entropy decoding stages 100; 320 decode respectively from the respective input signals 6a and 6b, and the transform coefficient blocks enter the inverse transformers 560; 580 and encode the parameters including the prediction parameters for the prediction providers 160; 260.

[0135] Therefore, the prediction providers 160 and 260 predict the blocks of the video frame at their respective resolution / quality levels. And for this purpose, the prediction providers 160 and 260 can be selected within a predetermined prediction mode such as a spatial intra prediction mode or a temporal inter prediction mode. Both modes are intra-layer prediction modes (i.e., prediction modes that depend only on the data within the sub-stream in which each level is contained).

[0136] However, in order to utilize the aforementioned redundancy between layers, the enhancement layer decoding stage 60 additionally includes a coding parameter inter-layer predictor 240, a resolution / quality improver 220, and / or a predictor 260 to be compared with the predictor 160. Further, or alternatively, the enhancement layer decoding stage 60 supports an inter-layer prediction mode that can provide an enhancement layer prediction signal 420 based on data derived from internal stages of the base layer decoding stage 80. The resolution / quality improver 220 subjects either the reconstructed base layer signals 200a, 200b, 200c or the base layer residual signal 480 to resolution or quality improvement to obtain an inter-layer prediction signal 380. The coding parameter inter-layer predictor 240 is to predict coding parameters, such as prediction parameters and motion parameters, in some form. The predictor 260 can support the inter-layer prediction mode, for example, further in accordance with the reconstructed portions of the base layer signals such as 200a, 200b, 200c. Alternatively, the reconstructed portions of the base layer residual signal 640 potentially improved at increasing resolution / quality levels are used as a reference / basis.

[0137] As described above, the decoding stages 60 and 80 can be operated in a block-based manner. That is, the frames of the video can be sub-divided into parts such as blocks. Different coarseness levels can be used to assign the prediction mode executed by the prediction providers 160, 260, the local transformation by the inverse transformers 560, 580, the filter coefficient selection by the filters 120, 140, and the prediction parameter setting for the prediction mode by the prediction providers 160, 260. That is, sub-dividing the frame into prediction blocks is in turn a continuation of the sub-division of the frame into blocks for which the prediction mode is selected, for example, so-called coding units or prediction units. The sub-division of the frame into blocks for transform coding, so-called transform units, can be different from the partition into prediction units. Some of the inter-layer prediction modes used by the prediction provider 260 are described by way of examples below. The prediction provider 260 has several intra-layer prediction modes, that is, prediction modes that internally derive the respective prediction signals input to the respective adders 180, 340, that is, each applied based solely on the state related to the coding stages 60, 80 of the current level.

[0138] Some further details of the illustrated blocks will become apparent from the description of the individual aspects below. Note that these descriptions are equally generally transferable to the description of other aspects and figures, unless they are explicitly related to the aspects for which such descriptions are provided.

[0139] In particular, the embodiment for the scalable video decoder of FIG. 23 represents a possible implementation of the scalable video decoder according to FIGS. 2 and 4. Although the scalable video decoder according to FIG. 23 has been described above, FIG. 23 shows the corresponding scalable video encoder, and the same reference numerals are used for the internal elements of the prediction coding / decoding method in FIGS. 23 and 24. The reason is as described above. Also, for the purpose of maintaining a common prediction basis between the encoder and the decoder, a reconstructable version of the base and enhancement layer signals is used in the encoder, and the already encoded part up to this point is reconstructed to obtain a reconstructable version of the scalable video. Therefore, the only difference from the description of FIG. 23 is that, similar to the coding parameter interlayer predictor 240, the prediction providers 160 and 260 determine the prediction parameters within some ratio / distortion optimization process rather than receiving the prediction parameters from the data stream. Rather, the providers transmit the prediction parameters thus determined to the entropy decoders 19a and 19b. The entropy decoders 19a and 19b transmit the respective base layer substream 6a and enhancement layer substream 6b in turn via the multiplexer 16 for inclusion in the data stream 6. In a similar manner, instead of these entropy encoders 19a and 19b outputting such entropy decoding results of the residuals, the reconstructed base layer signal 200 and the reconstructed enhancement layer signal 400 and the prediction residuals between the original base layer and enhancement layer versions 4a, 4b are received so as to be obtained via the subsequent subtractors 720 and 722 by the conversion modules 724, 726. However, otherwise, the structure of the scalable video encoder of FIG. 24 is consistent with the structure of the scalable video decoder of FIG. 23. Therefore, with regard to these issues, reference is made to the above description of FIG. 23. Here, as just outlined, any derivation from any data stream has to be changed to the determination of each element with its subsequent insertion into the respective data stream.

[0140] The techniques for intra - coding of enhancement layer signals used in the embodiments described below include multiple methods for generating an intra - prediction signal for an enhancement layer block (using base layer data). These methods are provided in addition to methods for generating an intra - prediction signal based only on samples of the reconstructed enhancement layer.

[0141] Intra - prediction is part of the process of reconstructing an intra - coded block. The final reconstructed block is obtained by adding the transformed and coded residual signal (which may be zero) to the intra - prediction signal. The residual signal is generated by inverse quantization (scaling) of the transform coefficient levels transmitted in the bitstream followed by inverse transform.

[0142] The following description applies to scalable coding with a quality enhancement layer (where the enhancement layer represents an input video having the same resolution as the base layer but higher quality or fidelity) and to scalable coding with a spatial enhancement layer (where the enhancement layer has a higher resolution than the base layer, i.e., more samples). In the case of a quality enhancement layer, up - sampling of the base layer signal is not required within blocks such as block 220, but can be applied to filtering 500 of the samples of the reconstructed base layer. In the case of a spatial enhancement layer, generally, up - sampling of the base layer signal is required, for example, within block 220.

[0143] The next-described aspects support different ways to use the reconstructed base layer samples (compared to 200) or base layer residual samples (compared to 640) for intra prediction of the enhancement layer blocks. One or more of the methods described below can be supported in addition to intra layer intra coding (where only the reconstructed enhancement layer samples (compared to 400) are used for intra prediction). The use of a particular method can be signaled at the level of the largest supported block size (such as the size of a block in H.264 / AVC within HEVC or a macroblock within an encoding tree / largest coding unit), or it can be signaled at all supported block sizes, or alternatively, it can be signaled for a subset of the supported block sizes.

[0144] For all of the methods described below, the prediction signal can be used directly as the reconstruction signal for the block. That is, no residual is transmitted at all. Or, the selected method for inter-layer intra prediction can be combined with residual coding. In a particular embodiment, the residual signal is transmitted via transform coding. That is, the quantized transform coefficients (transform coefficient levels) are transmitted using an entropy coding technique (e.g., variable length coding or arithmetic coding (compared to 19b)). And the residual is obtained by inverse quantizing (scaling) the transmitted transform coefficient levels and applying an inverse transform (compared to 580). In a particular version, the complete residual block corresponding to the block where the inter-layer intra prediction signal is generated is transformed using one transform (compared to 726). That is, the entire block is transformed using one transform of the same size as the prediction block. In another embodiment, the prediction block can be further sub-divided into smaller blocks (e.g., using hierarchical decomposition). And for each of the smaller blocks (which can have different block sizes), a separate transform can be applied. In another embodiment, the coding unit can be divided into smaller prediction blocks. And for zero or more of the prediction blocks, the prediction signal is generated using one of the methods for inter-layer intra prediction. And then, the residual of the entire coding unit is transformed using one transform (compared to 726). Or, the coding unit is sub-divided into different transform units. Here, the sub-division for forming the transform unit (the block to which one transform is applied) is different from the sub-division for decomposing the coding unit into prediction blocks.

[0145] In certain embodiments, the (upsampled / filtered) reconstructed base layer signal (compared to 380) is used directly as a prediction signal. Multiple methods for using the base layer for intra prediction of the enhancement layer include the following. The (upsampled / filtered) reconstructed base layer signal (compared to 380) is used directly as an enhancement layer prediction signal. This method is similar to the well-known H.264 / SVC inter-layer intra prediction mode. In this method, a prediction block for the enhancement layer is formed by juxtaposed samples of the base layer reconstruction signal that are upsampled (compared to 220) to match the corresponding sample positions in the enhancement layer and are optionally filtered before or after upsampling. In contrast to the SVC inter-layer intra prediction mode, this mode can be supported not only at the macroblock level (or the largest supported block size), but also at any block size. That is, the mode can be signaled not only for the largest supported block size, but also the blocks of the largest supported block size (macroblocks in MPEG4, H.264 and coding tree blocks / largest coding units in HEVC) are hierarchically sub-divided into smaller blocks / coding units, and the use of the inter-layer intra prediction mode can be signaled at any supported block size (for the corresponding blocks). In certain embodiments, this mode is only supported for the selected block size. Next, the syntax element signaling the use of this mode can be sent only for the corresponding block size. Or the value of the syntax element signaling the use of this mode (within another coding parameter) can be correspondingly restricted for another block size.Another difference from the inter-layer intra prediction mode within the SVC extension of H.264 / AVC is that the inter-layer intra prediction mode is supported not only when the collocated regions within the base layer are intra-coded, but also when the collocated base layer regions are inter-coded or partially inter-coded.

[0146] In certain embodiments, spatial intra prediction of the differential signal (see Aspect A) is performed. The multiple methods include the following methods. The reconstructed base layer signal (compared with 380) (potentially upsampled / filtered) is combined with the spatial intra prediction signal. Therein, the spatial intra prediction (compared with 420) is derived based on the differential samples for adjacent blocks (compared with 260). The differential samples represent the difference between the reconstructed enhancement layer signal (compared with 400) and the reconstructed base layer signal (compared with 380) (potentially upsampled / filtered).

[0147] FIG. 25 shows the generation of an inter-layer intra prediction signal by a sum 732 of a (upsampled / filtered) base layer reconstruction signal 380 (BLReco) and a spatial intra prediction using a difference signal 734 (EHDiff) of already encoded adjacent blocks 736. Therein, a difference signal (EHDiff) for already encoded block 736 is generated by subtracting 738 the (upsampled / filtered) base layer reconstruction signal 380 (BLReco) from a reconstructed enhancement layer signal (EHReco) (compared to 400), where the already encoded / decoded portion is hatched, and the currently encoded / decoded block / region / portion is 28. That is, the inter-layer intra prediction method described in FIG. 25 uses two overlapping input signals to generate a prediction block. In particular, for scalable spatial coding, usually, the difference signal 734 mainly contains high-frequency components. The difference signal 734 is available for all already reconstructed blocks (i.e., all enhancement layer blocks already encoded / decoded). The difference signal 734 for adjacent samples 742 of already encoded / decoded block 736 is used as an input to a spatial intra prediction technique (such as a spatial intra prediction mode specified within H.264 / AVC or HEVC). By the spatial intra prediction indicated by arrow 744, a prediction signal 746 for different components of block 28 to be predicted is generated. In certain embodiments, any clip function of the spatial intra prediction process (as known from H.264 / AVC or HEVC) is modified or disabled to match the dynamic range of the difference signal 734. The actually used intra prediction method (one of the provided methods, which can include planar intra prediction, DC intra prediction, or directional intra prediction 744 having any particular angle) is signaled within the bitstream 6b. It is possible to use a spatial intra prediction technique different from the methods provided in H.264 / AVC and HEVC (a method for generating a prediction signal using samples of already encoded adjacent blocks).The obtained predicted block 746, using the differential samples of adjacent blocks, is the first part of the final predicted block 420.

[0148] The second part of the prediction signal is generated using the collocated region 28 within the reconstructed signal 200 of the base layer. For the quality enhancement layer, the collocated base layer samples are either used directly or optionally filtered, for example, by a low-pass filter or a filter 500 that attenuates high-frequency components. For the spatial enhancement layer, the collocated base layer samples are upsampled. For upsampling 220, an FIR filter or a set of FIR filters can be used. An IIR filter can also be used. Optionally, the reconstructed base layer samples 200 can be filtered before upsampling. Alternatively, the base layer prediction signal (the signal obtained after upsampling the base layer) can be filtered after the upsampling stage. The processing of the base layer reconstruction can include one or more additional filters such as a non-blocking filter (compared to 120) or an adaptive loop filter (compared to 140). The base layer reconstruction 200 used for upsampling can be the reconstructed signal before any loop filter (compared to 200c). Alternatively, it can be the reconstructed signal after the non-blocking filter but before any other filter (compared to 200b). Alternatively, it can be the reconstructed signal after a specific filter or the reconstructed signal after applying all the filters used in the base layer decoding process (compared to 200a).

[0149] The two generated parts of the prediction signal (the spatially predicted differential signal 746 and the potentially filtered / upsampled base layer reconstruction 380) are added sample by sample 732 to form the final prediction signal 420.

[0150] Transferring exactly the aspects outlined above to the embodiments of FIGS. 6-10 may mean that the exactly outlined possibility of predicting the current block 28 of the enhancement layer signal is supported by each scalable video decoder / encoder as an alternative to the prediction scheme outlined with respect to FIGS. 6-10. The modes used are signaled within the enhancement layer substream 6b via respective prediction mode identifiers not shown in FIG. 8.

[0151] In certain embodiments, intra prediction follows inter-layer residual prediction (see Example B). Multiple methods for generating an intra prediction signal using base layer data include the following. A conventional spatial intra prediction signal (derived using samples of an adjacent reconstructed enhancement layer) is combined with an (upsampled / filtered) base layer residual signal (the inverse transform of the base layer transform coefficients, or the difference between base layer reconstruction and base layer prediction).

[0152] FIG. 26 shows the generation of an inter-layer intra prediction signal 420, etc., by a sum 752 of a spatial intra prediction 756 using an (upsampled / filtered) base layer residual signal 754 (BLResi) and samples 758 (EHReco) of a reconstructed enhancement layer of an already encoded adjacent block described by a dotted line 762.

[0153] The concept shown in FIG. 26 forms prediction block 420 by overlapping two prediction signals. Therein, one prediction signal 764 is generated from the already reconstructed enhancement layer samples 758, and the other prediction signal 754 is generated from the base layer residual samples 480. The first portion 764 of prediction signal 420 is derived by applying spatial intra prediction 756 using the reconstructed enhancement layer samples 758. The spatial intra prediction 756 can be one of the methods specified within H.264 / AVC. Or, it can be one of the methods specified within HEVC. Alternatively, it can be another spatial intra prediction technique that generates prediction signal 764 for the current block 18 that forms samples 758 of adjacent block 762. The actually used intra prediction method 756 (which can be one of the provided methods, including planar intra prediction, DC intra prediction, or directional intra prediction having any specific angle) is signaled within bitstream 6b. It is possible to use a spatial intra prediction technique (a method for generating a prediction signal using samples of already encoded adjacent blocks) different from the methods provided for H.264 / AVC and HEVC. The second portion 754 of prediction signal 420 is generated using the collocated residual signal 480 of the base layer. For the quality enhancement layer, the residual signal can be used such that it is reconstructed within the base layer. Or, the residual signal can be additionally filtered. For the spatial enhancement layer 480, the residual signal is upsampled 220 (for mapping the base layer sample positions to the enhancement layer sample positions) before it is used as the second portion of the prediction signal. Also, the base layer residual signal 480 can be filtered before or after the upsampling stage. An FIR filter can be applied to upsample 220 the residual signal. The upsampling process can be configured in a way that it is not filtered across the transform block boundaries within the base layer that are applied for the purpose of upsampling.

[0154] The base layer residual signal 480 used for inter-layer prediction can be the residual signal from which the transform coefficient levels of the base layer are obtained by scaling and inverse transformation 560. Or, the base layer residual signal 480 can be the difference between the reconstructed base layer signal 200 (before or after non-blocking and additional filtering, or during any filtering operation) and the prediction signal 660 used within the base layer.

[0155] The two generated signal components (spatial intra prediction signal 764 and inter-layer residual prediction signal 754) are added 752 to form the final enhancement layer intra prediction signal.

[0156] This means that the prediction mode just outlined with respect to FIG. 26 can be used or supported by any scalable video decoder / encoder according to FIGS. 6 - 10 to form the alternative prediction modes described above with respect to FIGS. 6 - 10 for the currently encoded / decoded portion 28.

[0157] In certain embodiments, a weighted prediction of spatial intra prediction and base layer reconstruction (see Mode C) is used. This represents the description of such weighted prediction, which was interpreted not only as an alternative to the embodiments outlined above with respect to FIGS. 6 - 10, but also as a description of the possibility of implementing the embodiments outlined above with respect to FIGS. 6 - 10 in a manner different from the given mode, in the specification published above of a particular implementation of the embodiments outlined above with respect to FIGS. 6 - 10.

[0158] Multiple methods for generating an intra prediction signal using base layer data include the following. The (upsampled / filtered) reconstructed base layer signal is combined with the spatial intra prediction signal. There, the spatial intra prediction is derived based on samples of the reconstructed enhancement layer of adjacent blocks. The final prediction signal is obtained by weighting (compared to 41) the spatial prediction signal and the base layer prediction signal in a way that different frequency components use different weightings. This can be achieved, for example, by filtering (compared to 62) the base layer prediction signal (compared to 38) with a low-pass filter, and filtering (compared to 64) the spatial intra prediction signal (compared to 34) with a high-pass filter, and then adding (compared to 66) the obtained filtered signals. Or, frequency-based weighting can be achieved by transforming (compared to 72, 74) the base layer prediction signal (compared to 38) and the enhancement layer prediction signal (compared to 34), and adding the obtained transformed blocks (compared to 76, 78), where different weighting coefficients (compared to 82, 84) are used for different frequency positions. Next, the obtained transformed block (compared to 42 in FIG. 10) can be inverse-transformed (compared to 84) and used as the enhancement layer prediction signal (compared to 54). Alternatively, the obtained transformation coefficients are added (compared to 52) to the scaled transmitted transformation coefficient levels (compared to 59), and then inverse-transformed (compared to 84) to obtain the reconstructed block (compared to 54) before deblocking and in-loop processing.

[0159] FIG. 27 shows the generation of an inter-layer intra prediction signal by the frequency-weighted sum of the (upsampled / filtered) base layer reconstruction signal (BLReco) and the spatial intra prediction using samples of the reconstructed enhancement layer of already encoded adjacent blocks (EHReco).

[0160] The concept of FIG. 27 forms prediction block 420 using two overlapping signals 772, 774. The first portion 774 of signal 420 is derived by applying a spatial intra prediction 776 corresponding to 30 in FIG. 6 using the reconstructed samples 778 of adjacent blocks already configured within the enhancement layer. The second portion 772 of the prediction signal 420 is generated using the collocated reconstructed signal 200 of the base layer. For the quality enhancement layer, the collocated base layer samples 200 are used directly. Or they may optionally be filtered, for example, by a low-pass filter or a filter that attenuates high-frequency components. For the spatial enhancement layer, the collocated base layer samples are upsampled 220. For upsampling, a FIR filter or a set of FIR filters may be used. It is also possible to use an IIR filter. Optionally, the reconstructed base layer samples are filtered before upsampling. Alternatively, the base layer prediction signal (the signal obtained after upsampling the base layer) may be filtered after the upsampling stage. The process of reconstructing the base layer can include one or more additional filters such as a non-blocking filter 120 or an adaptive loop filter 140. The base layer reconstruction 200 used for upsampling can be the reconstructed signal 200c before any of the loop filters 120, 140. Alternatively, it can be the reconstructed signal 200b after the non-blocking filter 120 but before another filter. Alternatively, it can be the reconstructed signal 200a after a particular filter, or the reconstructed signal after applying all the filters 120, 140 used in the base layer decoding process.

[0161] When the reference signs used in FIGS. 23 and 24 are compared with the reference signs used in connection with FIGS. 6-10, block 220 corresponds to reference sign 38 used in FIG. 6. 39 corresponds to a portion of 380. At least as far as the portion collocated with the current portion 28 is concerned, 420 collocated with the current portion 28 corresponds to 42. The spatial prediction 776 corresponds to 32.

[0162] Two prediction signals (potentially upsampled / filtered base layer reconstruction 386 and enhancement layer intra prediction 782) are combined to form the final prediction signal 420. A method for combining these signals can have the property that different weighting factors are used for different frequency components. In a particular embodiment, the upsampled base layer reconstruction is filtered with a low-pass filter (compared to 62) (it is also possible to filter the base layer reconstruction before upsampling 220). And the intra prediction signal (compared to 34 obtained by 30) is filtered with a high-pass filter (compared to 64). Both filtered signals are added (compared to 66) to form the final prediction signal 420. A pair of a low-pass filter and a high-pass filter can represent an orthogonal mirror filter pair, but this is not necessarily required.

[0163] In another specific embodiment (compared to FIG. 10), the combining process of the two prediction signals 380 and 782 is realized via a spatial transform. Both the (potentially upsampled / filtered) base layer reconstruction 380 and the intra prediction signal 782 are transformed (compared to 72, 74) using a spatial transform. Next, the transform coefficients of both signals (compared to 76, 78) are scaled with appropriate weighting coefficients (compared to 82, 84), and then added (compared to 90) to form the transform coefficient block of the final prediction signal (compared to 42). In one version, the weighting coefficients (compared to 82, 84) are selected such that for each transform coefficient position, the sum of the weighting coefficients for the components of both signals equals 1. In another version, for some or all of the transform coefficient positions, the sum of the weighting coefficients does not equal 1. In a particular version, the weighting coefficients are selected such that for the transform coefficients representing low frequency components, the weighting coefficient for the base layer reconstruction is larger than the weighting coefficient for the enhancement layer intra prediction signal, and for the transform coefficients representing high frequency components, the weighting coefficient for the base layer reconstruction is smaller than the weighting coefficient for the enhancement layer intra prediction signal.

[0164] In one embodiment, the obtained transform coefficient block (obtained by combining the weighted transformed signals for both components) is inverse-transformed (compared to 84) to form the final prediction signal 420 (compared to 54) compared to 42. In another embodiment, the prediction is made directly in the transform domain. That is, the encoded transform coefficient levels (compared to 59) are scaled (i.e., inverse quantized) and added to the transform coefficients of the prediction signal (compared to 42) obtained by adding the weighted transform signals for both components, and then added (compared to 52) to the block resulting from the inverse transform (compared to 84) of the transform coefficients (not shown in FIG. 10 but before the potential non-blocking 120 and further in-loop filtering step 140) to obtain the reconstructed signal 420 for the current block. In other words, in the first embodiment, the transform block obtained by adding the weighted transform signals for both components can be inverse-transformed and used as the enhancement layer prediction signal. Alternatively, in the second embodiment, the obtained transform coefficients can be added to the scaled transmitted transform coefficient levels and inverse-transformed to obtain a reconstructed block before non-blocking and in-loop processing.

[0165] Selection of the base layer reconstruction and the residual signal (see Mode D) can also be used. The following versions are used for the method of using the reconstructed base layer signal (as described above).

[0166] · The reconstructed base layer samples 200c before non-blocking 120 and further in-loop processing 140 (such as a sample adaptive offset filter or an adaptive loop filter). · The reconstructed base layer samples 200b after non-blocking 120 but before another in-loop processing 140 (such as a sample adaptive offset filter or an adaptive loop filter). ·The reconstructed base layer sample 200a after non-blocking 120 and further in-loop processing 140 (such as sample adaptive offset filter or adaptive loop filter), or between multiple in-loop processing steps.

[0167] The selection of the corresponding base layer signals 200a, b, c can be fixed for a particular decoder (and encoder) implementation. Or it can be signaled within the bitstream 6. In the latter case, different versions can be used. The use of a particular version of the base layer signal can be signaled at the sequence level, or at the picture level, or at the slice level, or at the largest coding unit level, or at the coding unit level, or at the prediction block level, or at the transform block level, or at any other block level. In another version, the selection can be made to depend on other coding parameters (such as the coding mode) or on the characteristics of the base layer signal.

[0168] In another embodiment, multiple versions of a method using an (upsampled / filtered) base layer signal 200 may be used. For example, two different modes that directly use the upsampled base layer signal (i.e., 200a) may be provided. Therein, the two modes use different interpolation filters. Alternatively, one mode uses additional filtering 500 of the (upsampled) base layer reconstruction signal. Similarly, multiple different versions for other modes as described above may be provided. The upsampled / filtered base layer signals 380 employed for different versions of the mode may differ within the interpolation filter used (including the interpolation filter that also applies to integer sample positions). Alternatively, the upsampled / filtered base layer signal 380 for a second version may be obtained by filtering 500 the upsampled / filtered base layer signal for a first version. One selection of different versions may be signaled at the sequence, picture, slice, maximum coding unit, coding unit level, prediction block level, or transform block level. Alternatively, it may be inferred from the characteristics of the corresponding reconstructed base layer signal or transmitted coding parameters.

[0169] The same applies to modes that use the reconstructed base layer residual signal via 480. Here, different versions of the interpolation filter or additional filtering steps used may also be used.

[0170] Different filters may be used to upsample / filter the reconstructed base layer signal and the base layer residual signal. This means that different approaches are used for upsampling the base layer residual signal than for upsampling the base layer reconstruction signal.

[0171] For the base layer block, the residual signal is zero (i.e., the transform coefficient levels are not sent to the block at all). The corresponding base layer residual signal can be replaced with another signal derived from the base layer. For example, this can be a high-pass filter version of the reconstructed base layer block, or another differential signal derived from the reconstructed base layer samples of adjacent blocks or the reconstructed base layer residual samples.

[0172] As far as the samples (see Mode H) used for spatial intra prediction within the enhancement layer are concerned, the following special processing is provided. For the modes using spatial intra prediction, the non-available adjacent samples within the enhancement layer (since the adjacent blocks are encoded after the current block, the adjacent samples are not available) can be replaced with the corresponding samples of the upsampled / filtered base layer signal.

[0173] As far as the coding of the intra prediction mode (refer to Mode X) is concerned, the following special modes and functionality may be provided. For the mode using spatial intra prediction such as 30a, the coding of the intra prediction mode may be changed (if available) in a way that the information about the intra prediction mode in the base layer is used to more efficiently code the intra prediction mode in the enhancement layer. This may be used, for example, for parameter 56. If the collocated area in the base layer (compared with 36) is intra-coded using a specific spatial intra prediction mode, a similar intra prediction mode is also used within the enhancement layer block (compared with 28). The intra prediction mode is usually signaled in a way that within the set of possible intra prediction modes, one or more modes are classified as the most likely modes. There, it can be signaled with a shorter codeword. Alternatively, with the following arithmetic coding decision, it results in fewer bits. Within the intra prediction of HEVC, (if available) the intra prediction mode of the upper block and (if available) the intra prediction mode of the left block are included within the set of the most likely modes. In addition to these modes, one or more additional modes (frequently used) are included in the list of the most likely modes. There, the actual additional modes depend on the usefulness of the intra prediction modes of the block above the current block and the block to the left of the current block. Within HEVC, three modes are accurately classified as the most likely modes. Within H.264 / AVC, one mode is classified as the most likely mode. This mode is derived based on the intra prediction modes used for the block above the current block and the block to the left of the current block. Any other concept (different between H.264 / AVC and HEVC) for classifying the intra prediction mode is possible and may be used for the following extensions.

[0174] To use base layer data for efficient coding of intra prediction modes within an enhancement layer, the concept of using one or more most likely modes is modified in a way that (if the corresponding base layer block is intra coded) includes the intra prediction mode used within the base layer block where the most likely modes are collocated. In certain embodiments, the following approach may be used. Given a current enhancement layer block, the collocated base layer block is determined. In a particular version, the collocated base layer block is the base layer block covering the collocated position of the top-left sample of the enhancement block. In another version, the collocated base layer block is the base layer block covering the collocated position of the sample at the center of the enhancement block. In other versions, another sample within the enhancement layer block may be used to determine the collocated base layer block. If the determined collocated base layer block is intra coded and the base layer intra prediction mode specifies an angular intra prediction mode and the intra prediction mode derived from the enhancement layer block to the left of the current enhancement layer block does not use the angular intra prediction mode, then the intra prediction mode derived from the left enhancement layer block is replaced with the corresponding base layer intra prediction mode. Otherwise, if the determined collocated base layer block is intra coded, the base layer intra prediction mode specifies an angular intra prediction mode, and the intra prediction mode derived from the enhancement layer block above the current enhancement layer block does not use the angular intra prediction mode, then the intra prediction mode derived from the above enhancement layer block is replaced with the corresponding base layer intra prediction mode. In other versions, different approaches are used to modify the list of most likely modes (consisting of a single element) using the base layer intra prediction mode.

[0175] Intercoding techniques for the spatial and quality enhancement layer are then provided.

[0176] In the state-of-the-art hybrid video coding standards (such as H.264 / AVC or the upcoming HEVC), the images of an image sequence are partitioned into blocks of samples. The block size can be fixed, or the coding technique can provide a hierarchical structure that allows the blocks to be further sub-divided into blocks with smaller block sizes. The reconstruction of a block is typically obtained by generating a prediction signal for the block and adding the transmitted residual signal. The residual signal is typically transmitted using transform coding, which means that the quantization indices for the transform coefficients (also called transform coefficient levels) are transmitted using entropy coding techniques. Then, on the decoder side, these transmitted transform coefficient levels are scaled and then inverse-transformed to obtain the residual signal that is added to the prediction signal. The residual signal is generated either by intra prediction (using only the data already transmitted for the current time instant) or by inter prediction (using the data already transmitted for different time instants).

[0177] In inter prediction, the prediction block is derived by motion-compensated prediction using the samples of a previously reconstructed frame. This is done by uni-directional prediction (using one reference image and one set of motion parameters). Alternatively, the prediction signal can be generated by multi-hypothesis prediction. In the latter case, two or more prediction signals are superimposed. That is, for each sample, a weighted average is constructed to form the final prediction signal. The multi-prediction signals (that are superimposed) can be generated using different motion parameters (e.g., different reference images or motion vectors) for different hypotheses. Also, for uni-directional prediction, it is possible to multiply the samples of the motion-compensated prediction signal with a constant coefficient and add a constant offset to form the final prediction signal. Also, such scaling and offset correction can be used for all or selected hypotheses within the multi-hypothesis prediction.

[0178] Even within scalable video coding, base layer information can be utilized to support inter-prediction processing for the enhancement layer. The SVC extension of the state-of-the-art video coding standard for scalable coding, H.264 / AVC, has one additional mode for improving the coding efficiency of inter-prediction processing within the enhancement layer. This mode is signaled at the macroblock level (a block of 16×16 luma samples). Within this mode, the reconstructed residual samples in the lower layer are used to improve the motion-compensated prediction signal in the enhancement layer. This mode is also called inter-layer residual prediction. If this mode is selected for a macroblock in the quality enhancement layer, the inter-layer prediction signal is assembled by the collocated samples of the reconstructed lower layer residual signal. If the inter-layer residual prediction mode is selected within the spatial enhancement layer, the prediction signal is generated by upsampling the collocated reconstructed base layer residual signal. For upsampling, an FIR filter is used. However, filtering is not applied across the transform block boundary. The prediction signal generated from the samples of the reconstructed base layer residual is added to the conventional motion-compensated prediction signal to form the final prediction signal for the enhancement layer block. Generally, for the inter-layer residual prediction mode, an additional residual signal is sent by transform coding. The transmission of the residual signal can be omitted (inferred to be equal to zero) if it is correspondingly signaled within the bitstream. The final reconstructed signal is obtained by adding the (reconstructed residual signal obtained by scaling the transmitted transform coefficient levels and applying the inverse spatial transform) to the prediction signal (which is obtained by adding the inter-layer residual prediction signal to the motion-compensated prediction signal).

[0179] Next, techniques for inter-coding of enhancement layer signals are described. This section describes methods for using the base layer signal in addition to the already reconstructed enhancement layer signal to intra-predict the enhancement layer signal to be encoded within a scalable video coding scenario. By using the base layer signal to inter-predict the enhancement layer signal to be encoded, the prediction error can be sufficiently suppressed. This results in an overall bitrate saving for the encoding of the enhancement layer. The main focus of this section is to increase the block-based motion compensation of samples in the enhancement layer using samples of the already encoded enhancement layer with additional signals from the base layer. The following description provides possibilities for using various signals from the encoded base layer. Quad-tree block partitioning is generally adopted as a preferred embodiment, but the examples presented are applicable to a general block-based hybrid coding approach without assuming any particular block partitioning. Even the use of the base layer reconstruction at the current time index, the base layer residual at the current time index, or the base layer reconstruction of the already encoded picture for inter-prediction of the enhancement layer blocks to be encoded is described. Also, ways are described in which the base layer signal can be combined with the already encoded enhancement layer signal to obtain a better prediction for the current enhancement layer. One of the state-of-the-art main techniques is the inter-layer residual prediction within H.264 / SVC. The inter-layer residual prediction within H.264 / SVC can be adopted for all inter-coded macroblocks, regardless of whether they are coded or not, by using the SVC macroblock type signaled by whether they use either the base mode flag or the conventional macroblock type. The flag is added to signal the macroblock syntax for the spatial and quality enhancement layers and the usage of the inter-layer residual prediction. When this residual prediction flag is equal to 1, the residual signal of the corresponding region within the reference layer is block-wise upsampled using a bilinear filter and used as the prediction for the residual signal of the enhancement layer macroblock. As a result, only the corresponding difference signal needs to be coded within the enhancement layer. For the description in this section, the following notations are used. t0 := the temporal index of the current picture t1 := the temporal index of the already reconstructed picture EL := Enhancement Layer BL := Base Layer EL(t0) := the current enhancement layer picture to be coded EL_reco := enhancement layer reconstruction BL_reco := base layer reconstruction BL_resi := base layer residual signal (inverse transform of the base layer transform coefficients, or the difference between the base layer reconstruction and the base layer prediction) EL_diff := the difference between the enhancement layer reconstruction and the upsampled / filtered base layer reconstruction The different base layer signals and enhancement layer signals are used within the description explained in FIG. 28.

[0180] For the description, the following characteristics of the filter are used. · Linearity: Although many of the filters mentioned in the specification are linear, non-linear filters can also be used. · Number of output samples: In the upsampling operation, the number of output samples is larger than the number of input samples. Here, filtering of the input data creates more samples than the input values. In conventional filtering, the number of output samples is equal to the number of input samples. Such a filtering operation can be used, for example, within high-quality scalable coding. · Phase delay: For filtering of samples at integer positions, the phase delay is usually zero (or a delay of an integer value within the sample). For the generation of samples at fractional positions (e.g., half-pel position or quarter-pel position), a filter having a fractional delay (within the unit of the sample) is usually applied to the samples of the integer lattice.

[0181] The conventional motion compensation prediction used in all hybrid video coding standards (e.g., MPEG-2, H.264 / AVC, or the upcoming HEVC standard) is illustrated in FIG. 29. To predict the signal of the current block, an area of the already reconstructed image is replaced and used as the prediction signal. For the signaling of the replacement, motion vectors are typically encoded within the bitstream. For integer sample precision motion vectors, a reference area within the reference image can be directly copied to form the prediction signal. However, it is also possible to transmit motion vectors with fractional sample precision. In this case, the prediction signal is obtained by filtering the reference signal with a filter having a fractional sample delay. The reference image used can typically be specified by including a reference image index in the bitstream syntax. In general, it is also possible to superimpose two or more prediction signals to form the final prediction signal. The concept is supported, for example, within a B slice with two motion hypotheses. In this case, the multiple prediction signals are generated using different motion parameters (e.g., different reference images or motion vectors) for different hypotheses. For single-direction prediction, it is also possible to multiply the samples of the motion compensation prediction signal having a certain factor and add a certain offset to form the final prediction signal. Such scaling and offset correction can also be used for all or selected hypotheses within multi-hypothesis prediction.

[0182] The following description applies to scalable coding with a quality enhancement layer (the enhancement layer represents the input video with the same resolution as the base layer but with higher quality or fidelity) and to scalable coding with a spatial enhancement layer (the enhancement layer has a higher resolution than the base layer, i.e., a larger number of samples). For the quality enhancement layer, upsampling of the base layer signal is not necessary, but filtering of the samples of the reconstructed base layer can be applied. In the case of the spatial enhancement layer, upsampling of the base layer signal is generally required.

[0183] Embodiments support different ways for using the reconstructed base layer samples or base layer residual samples for inter prediction of enhancement layer blocks. It is possible to support one or more of the methods described below in addition to conventional inter prediction and intra prediction. The usage of a particular method can be signaled at the level of the most supported block size (such as macroblocks in H.264 / AVC, or coding tree blocks / largest coding units in HEVC). It can be signaled for all supported block sizes. Or, it can be signaled for a subset of the supported block sizes.

[0184] For all the methods described below, the prediction signal can be used directly as the reconstruction signal for a block. Alternatively, a selected method for inter-layer prediction can be combined with residual coding. In certain embodiments, the residual signal is transmitted via transform coding. That is, the quantized transform coefficients (transform coefficient levels) are transmitted using an entropy coding technique (e.g., variable length coding or arithmetic coding), and the residual is obtained by inverse quantizing (scaling) the transmitted transform coefficient levels and applying an inverse transform. In a particular version, the complete residual block corresponding to the block where the inter-layer prediction signal is generated is transformed using a single transform. (That is, the entire block is transformed using a single transform of the same size as the prediction block.) In another embodiment, the prediction block can be further sub-divided into smaller blocks, for example, using hierarchical decomposition. And for the smaller blocks having different block sizes, separate transforms are applied. In another embodiment, the coding unit can be divided into smaller prediction blocks. And for zero or more of the prediction blocks, the prediction signal is generated using one of the methods for inter-layer prediction. And then, the residual of the entire coding unit is transformed using a single transform. Or, the coding unit is sub-divided into different transform units. Wherein, the sub-division for forming the transform unit (the block to which a single transform is applied) is different from the sub-division for decomposing the coding unit into prediction blocks.

[0185] Below, the possibility of performing prediction using the base layer residual and enhancement layer reconstruction is explained. The multiple methods include the following methods. A conventional inter-prediction signal (obtained by motion compensated interpolation of the already reconstructed enhancement layer image) is combined with the (upsampled / filtered) base layer residual signal (the inverse transform of the base layer transform coefficients, or the difference between the base layer reconstruction and the base layer prediction). This method is also referred to as the "BL_resi" mode (illustrated in FIG. 30).

[0186] In short, predictions for samples of the enhancement layer can be described below. EL prediction=filter(BL_resi(t0))+MCP_filter(EL_reco(t1)) It is also possible that two or more hypotheses of the enhancement layer reconstruction signal are used. For example, EL prediction=filter(BL_resi(t0))+MCP_filter1(EL_reco(t1))+MCP_filter2(EL_reco(t2)) The motion compensation prediction (MCP) filter used on the enhancement layer (EL) reference image is of integer or fractional sample accuracy. The MCP filter used on the EL reference image is the same as, or different from, the MCP filter used on the BL reference image during the BL decoding process. The motion vector MV(x, y, t) is defined to indicate a specific position within the EL reference image. The parameters x and y indicate the spatial position within the image. The parameter t is used to describe the time index of the reference image and is also called the reference index. Often, the term motion vector is used to refer to only the two spatial components (x, y). The integer part of the MV is used to take a set of samples from the reference image. And the fractional part of the MV is used to select the MCP filter from a set of filters. The obtained reference samples are filtered to create filtered reference samples. Motion vectors are generally encoded using different predictions. It means that the motion vector predictor is derived based on the already encoded motion vectors (and the syntax element potentially indicates one of the used set of potential motion vector predictors), and different vectors are included within the bitstream. The final motion vector is obtained by adding the transmitted motion vector difference to the motion vector predictor. Usually, it is also possible to fully derive the motion parameters for a block. Thus, usually, a list of potential motion parameter candidates is constructed based on the already encoded data. This list can include the motion parameters of spatially adjacent blocks as well as the motion parameters derived based on the motion parameters within the collocated blocks within the reference frame. The base layer (BL) residual signal can be defined as one of the following. The inverse transform of the BL transform coefficients, or, · The difference between the BL reconstruction and the BL prediction, or, · For a BL block where the inverse transform of the BL transform coefficients is zero, it can be replaced by another signal derived from the BL, e.g., the high-pass filtered version of the reconstructed BL block, or, · A combination of the above methods. To calculate the EL prediction elements from the current BL residuals, the regions in the BL image juxtaposed with the regions considered in the EL image are identified, and the residual signal is taken from the identified BL regions. The definition of the juxtaposed regions can be made such that it accounts for an integer scaling factor of the BL resolution (e.g., 2× scalability), or a fractional scaling factor of the BL resolution (e.g., 1.5× scalability). Or it can even produce the same EL resolution as the BL resolution (e.g., quality scalability). In the case of quality scalability, the juxtaposed blocks in the BL image have the same coordinates as the EL blocks to be predicted. The juxtaposed BL residuals can be upsampled / filtered to generate filtered BL residual samples. The final EL prediction is obtained by adding the filtered EL reconstruction samples and the filtered BL residual samples.

[0187] The multiple methods for prediction using the base layer reconstruction and the enhancement layer difference signal (see Example J) include the following method. The (upsampled / filtered) reconstructed base layer signal is combined with the motion compensated prediction signal. There, the motion compensated prediction signal is obtained by the motion compensated difference image. The difference image represents the difference between the reference image and the sum of the reconstructed enhancement layer signal and the (upsampled / filtered) reconstructed base layer signal. This method is also called the BL_reco mode.

[0188] This concept is illustrated in Figure 31. Briefly, the prediction for the EL samples is described below. EL prediction = filter(BL_reco(t0)) + MCP_filter(EL_diff(t1))

[0189] It is also possible to use two or more hypotheses of the EL difference signal. For example, EL prediction = filter(BL_resi(t0)) + MCP_filter1(EL_diff(t1)) + MCP_filter2(EL_diff(t2))

[0190] For the EL difference signal, the following versions may be used. · The difference between the EL reconstruction and the upsampled / filtered BL reconstruction, or, · The difference between the EL reconstruction before or during the loop filtering stage (such as non-blocking, SAO, ALF) and the upsampled / filtered BL reconstruction.

[0191] The usage of a specific version can be fixed within the decoder, or it can be signaled at the sequence level, picture level, slice level, largest coding unit level, coding unit level, or another partition level. Alternatively, it can depend on other coding parameters.

[0192] When the EL difference signal is defined using the difference between the EL reconstruction and the upsampled / filtered BL reconstruction, this allows the EL reconstruction and BL reconstruction to be saved and the EL difference signal of the blocks using the prediction (sprediction) mode to be calculated in-place. As a result, the memory required to store the EL difference signal can be saved. However, it incurs a slight computational overhead.

[0193] The MCP filter used on the EL difference image can be of integer or fractional sample accuracy. · For the MCP of the difference image, a different interpolation filter than the MCP of the reconstructed image can be used. · For the MCP of the difference image, the interpolation filter can be selected based on the characteristics of the corresponding region within the difference image (or based on the coding parameters within the bitstream, or based on the transmitted information).

[0194] The motion vector MV(x, y, t) is defined to point to a specific position within the EL difference image. The parameters x and y refer to the spatial position within the image, and the parameter t is used to specify the time index of the difference image.

[0195] The integer part of the MV is used to obtain a set of samples from the difference image, and the fractional part of the MV is used to select an MCP filter from a set of filters. The obtained difference samples are filtered to create filtered difference samples.

[0196] The dynamic (active) range of the difference image can theoretically exceed the dynamic range of the original image. Assuming an 8-bit representation of the image within the range [0 255], the difference image can have a range [-255 255]. However, in practice, most of the amplitudes are distributed near plus or minus 0. In a preferred embodiment for storing the difference image, a constant offset of 128 is added, and the result is clipped to the range [0 255] and stored as a normal 8-bit image. Then, within the encoding and decoding processes, the offset of 128 is subtracted from the difference amplitude read from the difference image.

[0197] For the method of using the reconstructed BL signal, the following versions can be used. This can be signaled as fixed, or it can be signaled at the sequence level, image level, slice level, maximum coding unit level, coding unit level, or another partitioning level. Alternatively, it can depend on another coding parameter. · Reconstructed base layer samples (such as sample adaptive offset filters or adaptive loop filters, etc.) before non-blocking and further in-loop processing. · Reconstructed base layer samples (such as sample adaptive offset filters or adaptive loop filters, etc.) after non-blocking, but before further in-loop processing. · After non-blocking and further in-loop processing (such as sample adaptive offset filter or adaptive loop filter), or reconstructed base layer samples between multiple in-loop processing steps.

[0198] To calculate the EL prediction component from the current BL reconstruction, the region in the BL image juxtaposed to the region considered in the EL image is identified. Then, the reconstruction signal is taken from the identified BL region. The definition of the juxtaposed region can be made such that it even accounts for producing the same EL resolution as the integer scaling factor of the BL resolution (e.g., 2× scalability), or the fractional scaling factor of the BL resolution (e.g., 1.5× scalability), or the BL resolution (e.g., SNR scalability). In the case of SNR scalability, the juxtaposed block in the BL image has the same coordinates as the EL block to be predicted.

[0199] The final EL prediction is obtained by adding the filtered EL difference samples and the filtered BL reconstruction samples.

[0200] Some possible variations of the mode of combining the (upsampled / filtered) base layer reconstruction signal and the motion-compensated enhancement layer difference signal are described below. ·Multiple versions of a method of using an (upsampled / filtered) BL signal may be used. The upsampled / filtered BL signals employed for these versions may differ in the interpolation filter used (including the interpolation filter that also filters the integer sample positions used), or the upsampled / filtered BL signal for a second version may be obtained by filtering the upsampled / filtered BL signal for the first version. One selection of the different versions may be signaled in a sequence, in an image, in a slice, in a maximum coding unit, at a coding unit level, or at another level of an image partition. Alternatively, it may be inferred from the characteristics of the corresponding reconstructed BL signal or the transmitted coded parameters. ·Different filters may be used to upsample / filter the BL reconstructed signal in the case of the BL_reco mode and the BL residual signal in the case of the BL_resi mode. ·The upsampled / filtered BL signal may also be combined with two or more hypotheses of the motion compensated difference signal. This is illustrated in Figure 32.

[0201] Considering the above, the prediction can be performed using a combination of base layer reconstruction and enhancement layer reconstruction (see Example C). One major difference from the above description regarding FIGS. 11, 12, and 13 is the coding mode for obtaining the intra-layer prediction 34 that is performed temporally rather than spatially. That is, instead of the spatial prediction 30, the temporal prediction 32 is used to form the intra-layer prediction signal 34. Therefore, some of the embodiments described below can be easily applied to the respective embodiments above FIGS. 6 to 10 and FIGS. 11 to 13. The multiple methods include the following methods. The (upsampled / filtered) reconstructed base layer signal is combined with the inter prediction signal, where the inter prediction is derived by motion compensation prediction using the image of the reconstructed enhancement layer. The final prediction signal is obtained by weighting the inter prediction signal and the base layer prediction signal in a way that different frequency components use different weightings. For example, this can be achieved by any of the following. · Filtering the base layer prediction signal with a low-pass filter, filtering the inter prediction signal with a high-pass filter, and summing the obtained filtered signals. · Transforming the base layer prediction signal and the inter prediction signal and superimposing the obtained transformed blocks. Therein, different weighting coefficients are used for different frequency positions. Next, the obtained transformed blocks are inverse-transformed and used as the enhancement layer prediction signal. Alternatively, the obtained transformation coefficients are added to the scaled transmitted transformation coefficient levels and inverse-transformed to obtain a reconstructed block before non-blocking and in-loop processing.

[0202] This mode may also be referred to as the "BL_comb" mode described in FIG. 33.

[0203] In short, the EL prediction can be shown as follows. EL prediction=BL_weighting(BL_reco(t0))+EL_weighting(MCP_filter(EL_reco(t1)))

[0204] In a preferred embodiment, the weighting is made depending on the ratio of the EL resolution to the BL resolution. For example, when the BL is to be scaled up by a factor within the range [1, 1.25), a predetermined set of weightings for EL and BL reconstruction can be used. When the BL is to be scaled up by a factor within the range [1.25 1.75), a different set of weightings can be used. When the BL is to be scaled up by a factor of 1.75 or more, another different set of weightings and the like can be used.

[0205] It is also possible in another embodiment regarding spatial intra-layer prediction to perform a specific weighting depending on the scaling factor separation base layer and the enhancement layer.

[0206] In another preferred embodiment, the weighting is made depending on the EL block size to be predicted. For example, for a 4×4 block within the EL, a weighting matrix specifying the weighting for the EL reconstruction transform coefficients can be defined. And another weighting matrix specifying the weighting for the BL reconstruction transform coefficients can be defined. The weighting matrix for the BL reconstruction transform coefficients can be, for example, as follows. 64, 63, 61, 49, 63, 62, 57, 40, 61, 56, 44, 28, 49, 46, 32, 15, And the weighting matrix for the EL reconstruction transform coefficients is, for example, as follows. 0, 2, 8, 24, 3, 7, 16, 32, 9, 18, 20, 26, 22, 31, 30, 23,

[0207] Similarly, for block sizes such as 8×8, 16×16, 32×32, etc., a separate weighting matrix can be defined.

[0208] The actual transform used for frequency domain weighting can be the same as, or different from, the transform used to encode the prediction residual. For example, the integer approximation for DCT can be used for both frequency domain weighting and calculating the transform coefficients of the prediction residual to be encoded in the frequency domain.

[0209] In another preferred embodiment, the maximum transform size is defined for frequency domain weighting in order to limit the computational complexity. If the EL block size under consideration is larger than the maximum transform size, then the EL reconstruction and the BL reconstruction are spatially separated into a series of adjacent sub-blocks, the frequency domain weighting is performed over the sub-blocks, and the final prediction signal is formed by assembling the weighted results.

[0210] Moreover, the weighting can be performed on the luminance and chrominance components, or a selected subset of the color components.

[0211] Various possibilities for deriving enhancement layer coding parameters are described below. The coding (or prediction) parameters to be used to reconstruct the enhancement layer blocks can be derived by multiple methods from the collocated coding parameters in the base layer. The base layer and the enhancement layer can have different spatial resolutions, or they can have the same spatial resolution.

[0212] In the scalable video extension of H.264 / AVC, inter-layer motion prediction is performed for macroblock types signaled by a syntax element base mode flag. If the base mode flag is equal to 1 and the corresponding reference macroblock in the base layer is inter-coded, then the enhancement layer macroblock is also inter-coded, and all motion parameters are inferred from the collocated base layer block. Otherwise (the base mode flag is equal to 0), for each motion vector, so-called "motion prediction flag", syntax elements are sent and it is specified whether the base layer motion vector is used as a motion vector predictor. If the "motion prediction flag" is equal to 1, the motion vector predictor of the collocated reference block in the base layer is scaled according to the resolution ratio and used as a motion vector predictor. If the "motion prediction flag" is equal to 0, the motion vector predictor is calculated as specified within H.264 / AVC.

[0213] A method for deriving enhancement layer coding parameters is described below. The sample array associated with the base layer image is decomposed into blocks, and each block has coding (or prediction) parameters associated with it. In other words, all sample positions within a particular block have particular associated coding (or prediction) parameters. The coding parameters may include parameters for motion compensation prediction including the number of motion hypotheses, reference indices, motion vectors, motion vector prediction identifiers, and merge identifiers. The coding parameters may also include intra prediction parameters such as the intra prediction direction.

[0214] It is signaled in the bitstream that blocks in the enhancement layer are coded using collocated information from the base layer.

[0215] For example, the derivation of enhancement layer coding parameters (see Example T) can be made as follows. For the N×M blocks of the enhancement layer signaled using the collocated base layer information, the coding parameters related to the sample positions within the block can be derived based on the coding parameters related to the collocated sample positions within the base layer sample array.

[0216] In a particular embodiment, this process is done by the following steps. 1. Derivation of the coding parameters for each sample position within the N×M enhancement layer block based on the base layer coding parameters. 2. Derivation of the partitioning of the N×M enhancement layer blocks within the sub-block such that all sample positions within a particular sub-block have the same related coding parameters.

[0217] Also, the second step can be omitted.

[0218] Step 1 is to use the function f el for the enhancement layer sample position p c to give the coding parameter c, that is, c = f c (p el ) and execute it.

[0219] TIFF2025102790000001.tif82170

[0220] The distance between two horizontally or vertically adjacent base layer sample positions of the base layer is thus equal to 1. Both the top left most base layer sample and the top left most enhancement layer sample have the position of p=(0,0).

[0221] As another example, the function f c (p el ) is the base layer sample position p el closest to the blIt can return to the encoding parameter c related thereto. Also, the function f c (p el ) can interpolate the encoding parameter when a specific enhancement layer sample position has a fractional component within the unit of the distance between the base layer sample positions.

[0222] Before returning the motion parameter, the function f c rounds the spatial replacement component of the motion parameter to the nearest available value within the enhancement layer sampling grid.

[0223] Since each sample position is related to the prediction parameter after step 1, after step 1, the samples of each enhancement layer can be predicted. Nevertheless, in step 2, a block partition can be derived to perform a prediction operation on the samples of a larger block or to transform and encode the prediction residual within the blocks of the derived partition.

[0224] Step 2 can be performed by classifying the enhancement layer sample positions into square or rectangular blocks. Each is decomposed into one of a set of possible decompositions into sub-blocks. The square or rectangular blocks correspond to the leaves within a quadtree structure where they can exist at different levels represented in FIG. 34.

[0225] The level and decomposition of each square or rectangular block can be determined by performing the following ordered steps. a) Set the highest level to the level corresponding to a block of size N×M. Set the current level to the lowest level, i.e., the level where the square or rectangular block contains a single block of the smallest block size. Proceed to step b). b) For each square or rectangular block at the current level, if there is a possible decomposition of the square or rectangular block, then all sample positions within each sub-block are related to the same coding parameter, or (according to some magnitude of difference) are related to the coding parameter with a small difference. Such a decomposition is a candidate decomposition. Among all candidate decompositions, select the one that decomposes the square or rectangular block into the minimum number of sub-blocks. If the current level is the highest level, go to step c). Otherwise, set the current level to the next higher level and go to step b). c) End

[0226] function f c can be selected in such a way that at a certain level within step b), there is always at least one candidate decomposition.

[0227] Grouping of blocks having the same coding parameter is not limited to square blocks, but blocks can be grouped into rectangular blocks. Furthermore, the grouping is not limited to a quadtree structure. It is also possible to use a structure in which a block is decomposed into two rectangular blocks of the same size or two rectangular blocks of different sizes. It is also possible to use a decomposition structure that uses quadtree decomposition up to a certain level and then uses decomposition into two rectangular blocks. Also, any other block decomposition is possible.

[0228] As opposed to the motion parameter prediction mode between SVC layers, the described mode is supported not only at the macroblock level (or the most widely supported block size), but also at any block size. It means that the mode can be signaled not only for the most widely supported block size, but also that blocks of the most widely supported block size (macroblocks in MPEG4, H.264, and coding tree blocks / largest coding units in HEVC) can be hierarchically subdivided into smaller blocks / coding units, and the use of the inter-layer motion mode can be signaled for any supported block size (for the corresponding blocks). In certain embodiments, this mode supports only the selected block size. Next, the syntax element signaling the usage of this mode can be sent only for the corresponding block size. Or, the value of the syntax element signaling the usage of this mode (within another coding parameter) can be restricted to correspond to another block size. Also, the difference from the inter-layer motion parameter prediction mode in the SVC extension of H.264 / AVC is that the blocks coded in this mode are not completely inter-coded. The block can include intra-coded sub-blocks depending on the collocated base layer signal.

[0229] One of several methods for reconstructing samples of an M×M enhancement layer block using the coding parameters derived by the method described above can be signaled in the bitstream. Such a method for predicting an enhancement layer block using the derived coding parameters can include the following. · For motion compensation, using the derived motion parameters and the reconstructed enhancement layer reference image to derive a prediction signal for the enhancement layer block. ·A combination of (a) the (upsampled / filtered) base layer reconstruction for the current image and (b) a motion compensation signal, using the derived motion parameters and the enhanced layer reference image, generated by subtracting the (upsampled / filtered) base layer reconstruction from the reconstructed enhanced layer image. ·A combination of (a) the (upsampled / filtered) base layer residual for the current image (the difference between the reconstructed signal and the prediction, or the inverse transform of the encoded transform coefficient values) and (b) a motion compensation signal using the derived motion parameters and the reconstructed enhanced layer reference image.

[0230] While another sub-block is classified as being inter-coded, a process for obtaining a partition into smaller blocks within the current block and deriving coding parameters for the sub-block can classify several sub-blocks as being intra-coded. For an inter-coded sub-block, motion parameters are obtained from the neighboring base layer blocks. However, if the neighboring base layer blocks are intra-coded, the corresponding sub-blocks within the enhanced layer can be classified as being intra-coded. For samples of such intra-coded sub-blocks, the enhanced layer signal can be predicted by using information from the base layer. For example, ·The (upsampled / filtered) version of the corresponding base layer reconstruction is used as an intra-prediction signal. ·The derived intra-prediction parameters are used for spatial intra-prediction within the enhanced layer.

[0231] The following embodiments relate to a method for predicting an enhancement layer block using a weighted combination of prediction signals, including generating a prediction signal for an enhancement layer block by combining (a) an enhancement layer intra prediction signal obtained by spatial or temporal (i.e., motion compensated) prediction using samples of a reconstructed enhancement layer, and (b) a base layer prediction signal that is an (upsampled / filtered) base layer reconstruction for the current image. The final prediction signal is obtained by weighting the enhancement layer intra prediction signal and the base layer prediction signal such that for each sample, a weighting according to a weighting function is used.

[0232] For example, the weighting function can be implemented in the following way. A low-pass filtered version of the original enhancement layer intra prediction signal v is compared to a low-pass filtered version of the base layer reconstruction u. From this comparison, a weighting for each sample position to be used for combining the original inter prediction signal and the (upsampled / filtered) base layer reconstruction is derived. For example, the weighting is derived by mapping the difference u - v to the weighting w using a transfer function t. That is, t(u - v)=w

[0233] Different weighting functions can be used for different block sizes of the current block to be predicted. Also, the weighting function can be varied according to the temporal distance of the reference image from which the inter prediction hypothesis is obtained.

[0234] In the case of an enhancement layer intra prediction signal that is an intra prediction signal, the weighting function can be implemented using different weightings that depend, for example, on the position within the current block to be predicted.

[0235] In a preferred embodiment, a method for deriving enhancement layer coding parameters is used. And step 2 of the method uses a set of possible decompositions of square blocks, as described in FIG. 35.

[0236] In a preferred embodiment, the function f c (p el ) returns coding parameters related to the base layer sample positions given by the function f p,m × n (p el ) described above with m = 4 and n = 4.

[0237] In an embodiment, the function f c (p el ) returns the following coding parameter c. · First, the base layer sample position is derived as p bl =f p,4 ×4(p el ). · If p bl has related inter-prediction parameters obtained by merging with a previously coded base layer block (or has the same motion parameters), then c is equal to the motion parameters of the enhancement layer block corresponding to the base layer block used for merging within the base layer (i.e., the motion parameters are copied from the corresponding enhancement layer block). · Otherwise, c is equal to the coding parameters related to p bl .

[0238] Also, combinations of the above embodiments are possible.

[0239] In another embodiment, using the juxtaposed base layer information, for the enhancement layer blocks to be signaled, since the intra prediction parameters obtained from the initial set of motion parameters are related to those enhancement layer sample positions, the blocks can be merged with the blocks containing these samples (i.e., copies of the initial set of motion parameters). The initial set of motion parameters consists of an indicator for using one or two hypotheses, a reference index referring to the first image in the list of reference images, and motion vectors with zero padding.

[0240] In another embodiment, using the juxtaposed base layer information, for the enhancement layer blocks to be signaled, the enhancement layer samples having the derived motion parameters are first predicted and reconstructed in a certain ordering. Thereafter, the samples having the derived intra prediction parameters are predicted in the intra reconstruction order. As a result, the intra prediction can use the already reconstructed sample values from (a) the adjacent inter prediction blocks and (b) the previous adjacent intra prediction blocks within the intra reconstruction order.

[0241] In another embodiment, for the enhancement layer blocks that are merged (i.e., taking the motion parameters derived from another inter prediction block), the list of merge candidates additionally includes candidates from the corresponding base layer block. And if the enhancement layer has a higher spatial sampling ratio than the base layer, additionally, up to four candidates derived from the base layer candidates are included by improving only the available adjacent values within the enhancement layer with the spatial replacement component.

[0242] In another embodiment, the magnitude of the difference used in step 2b) claims that there is a small difference within the sub - block only if the difference disappears completely. That is, the sub - block can be formed only when all the included sample positions have the same derived coding parameters.

[0243] In another embodiment, if the magnitude of the difference used in step 2b) is such that either (a) all included sample positions have derived motion parameters and pairs of sample positions within a block do not have derived motion parameters that differ by more than a certain value according to the vector norm applied to the corresponding motion vectors, or (b) all included sample positions have derived intra prediction parameters and pairs of sample positions within a block do not have derived intra prediction parameters that differ by more than a certain angle of directional intra prediction, then it is asserted that there is a small difference within the sub-block. The parameters provided for the sub-block are calculated by an average or median operation. In another embodiment, the partition obtained by inferring coding parameters from the base layer can be further refined based on side information signaled within the bitstream. In another embodiment, the residual coding for a block for which coding parameters are inferred from the base layer is independent of the partition within the block inferred from the base layer. For example, it means that although the inference of coding parameters from the base layer divides the block into several sub-blocks each having a separate set of coding parameters, a single transform can be applied to the block. Or, the block for which the partition and coding parameters for the sub-blocks are inferred from the base layer can be divided into smaller blocks for the purpose of transform coding the residual. There, the division into transform blocks is independent of the inferred partition within the block having different coding parameters.

[0244] In another embodiment, the residual encoding for a block whose encoding parameters are inferred from the base layer depends on the partitions within the block inferred from the base layer. For example, for transform encoding, it means that the splitting of blocks within a transform block depends on the partitions inferred from the base layer. In one version, a single transform can be applied to each sub-block having different encoding parameters. In another version, the partitions can be refined based on side information included in the bitstream. In another version, some sub-blocks can be grouped into larger blocks so that the residual signals are signaled in the bitstream for the purpose of transform encoding.

[0245] Also, embodiments obtained by combinations of the above-described embodiments are also possible.

[0246] Regarding enhancement layer motion vector encoding, the following section describes a method for reducing motion information in a scalable video encoding application by providing multiple enhancement layer predictors and using the motion information encoded in the base layer to efficiently encode the motion information of the enhancement layer. This idea is suitable for scalable video encoding including spatial, temporal, and quality scalability.

[0247] In the scalable video extension between H.264 / AVC layers, motion prediction is performed for the macroblock type signaled for the syntax element "base mode flag". If the "base mode flag" is equal to 1 and the corresponding reference macroblock in the base layer is inter-coded, then the enhancement layer macroblock is also inter-coded. And all motion parameters are inferred from the collocated base layer blocks. Otherwise (if the "base mode flag" is equal to 0), each motion vector (a syntax element of the so-called "motion prediction flag") is sent and specified regardless of whether the base layer motion vector is used as a motion vector predictor. If the "motion prediction flag" is equal to 1, the motion vector predictor of the collocated reference blocks in the base layer is scaled according to the resolution ratio and used as a motion vector predictor. If the "motion prediction flag" is equal to 0, the motion vector predictor is calculated as defined in H.264 / AVC. In HEVC, motion parameters are predicted by applying adaptive motion vector prediction (AMVP). AMVP features two competing spatial motion vector predictors and one temporal motion vector predictor. The spatial candidates are selected from the positions of adjacent prediction blocks located to the left or above the current prediction block. The temporal candidate is selected among the collocated positions of the previously encoded pictures. The positions of all spatial and temporal candidates are shown in Figure 36.

[0248] After the spatial and temporal candidates are inferred, a redundancy check that may introduce the zero motion vector as a candidate into the list is performed. The candidate list that describes the indexes is sent to identify the motion vector predictor used with the motion vector difference for motion compensation prediction. HEVC also uses a block merge algorithm that aims to reduce the coding redundancy of motion parameters resulting from the quad-tree based on the coding configuration. This is achieved by creating regions consisting of multiple prediction blocks sharing specific motion parameters. These motion parameters only need to be coded once for the first prediction block of each region where new kinds of motion information are sown. Similar to AMVP, the block merge algorithm constructs a list containing possible merge candidates for each prediction block. The number of candidates is defined by "NumMergeCands" which is signaled in the slice header and ranges from 1 to 5. The candidates are inferred from the prediction blocks in the temporal picture juxtaposed with the spatially neighboring prediction blocks. The possible sample positions for the prediction blocks considered as candidates are equal to the positions shown in Figure 36. An example of the block merge algorithm with the partition of possible prediction blocks in HEVC is illustrated in Figure 37. The thick lines in Figure 37(a) define all the prediction blocks merged into one region to hold specific motion data. This motion data is sent only to block S. The current prediction block to be coded is indicated by "X". The blocks in the striped region are successors of the prediction block X in the block scanning order and thus do not yet have the associated prediction data. The dots indicate the sample positions of the adjacent blocks which are possible spatial merge candidates. Before the possible candidates are inserted into the predictor list, a redundancy check for the spatial candidates is performed as shown in Figure 37(b).

[0249] If the number of spatial and temporal candidates is less than "NumMergeCands", additional candidates are provided by combining with existing candidates or by inserting zero motion vector candidates. If a candidate is added to the list, it has an index used to identify the candidate. When a new candidate is added to the list, the merge index (starting from 0) increases until the list is complete at the last candidate identified by index "NumMergeCands" - 1. Fixed - length codewords are used to encode the merge candidate index to ensure independent operations of candidate list derivation and bitstream syntax analysis.

[0250] The following section describes a method for using a multiple enhancement layer predictor that includes a predictor obtained from the base layer to encode the motion parameters of the enhancement layer. The motion information already encoded for the base layer can be used to significantly reduce the motion data rate while encoding the enhancement layer. This method includes the possibility of directly deriving all the motion data of the prediction block from the base layer. In this case, additional motion data need not be encoded. In the following description, the term prediction block refers to a prediction unit within HEVC (an M×N block within H.264 / AVC) and can be understood as a general set of samples within an image.

[0251] The first part of the present section relates to extending the list of motion vector prediction candidates by a base layer motion vector predictor (see aspect K). The base layer motion vectors are added to the motion vector predictor list during enhancement layer encoding. This is achieved by inferring one or multiple motion vector predictors of the collocated prediction blocks from the base layer and using them as candidates in the list of predictors for motion compensation prediction. The collocated prediction blocks of the base layer are located at the center, left, top, right or bottom of the current block. If the predicted block of the base layer at the selected position does not contain motion related data or exists outside the current range and is therefore not currently accessible, alternative positions can be used to infer the motion vector predictor. These alternative positions are represented in FIG. 38.

[0252] The inferred motion vectors of the base layer can be scaled according to the resolution ratio before they can be used as predictor candidates. Similar to the motion vector difference, an index describing the candidate list of motion vector predictors is sent to the prediction block that specifies the final motion vector used for motion compensation prediction. Contrasted with the scalable extension of the H.264 / AVC standard, the embodiments presented here do not constitute the usage of motion vector predictors of collocated blocks within the reference picture - rather it is available within a list in another predictor and can be described by the index sent. In an embodiment, the motion vector is obtained from the center position C1 of the adjacent prediction blocks in the base layer and added to the head of the candidate list as the first entry. The candidate list of the motion vector predictor is extended by one entry. If there is no motion data at all in the base layer available for the sample position C1, the list structure is not touched. In another embodiment, any sequence of sample positions in the base layer can be checked for motion data. When motion data is found, the motion vector predictor at the corresponding position is inserted into the candidate list and is available for motion compensation prediction in the enhancement layer. Moreover, the motion vector predictor derived from the base layer is inserted into the candidate list at any other position of the list. In another embodiment, if certain regulations are recognized, the base layer motion predictor may only be inserted into the candidate list. These constraints include the value of the merge flag of the adjacent reference blocks which must not be equal to zero. Another constraint may be the width of the prediction block in the enhancement layer equal to the width of the base adjacent prediction block with respect to the resolution ratio. For example, in the application of K× spatial scalability, if the width of the adjacent block in the base layer is equal to N and the width of the prediction block to be encoded in the enhancement layer is K * The motion vector predictor can only be inferred if it is equal to N. In another embodiment, one or more motion vector predictors from several sample positions of the base layer can be added to the candidate list of the enhancement layer. In another embodiment, a candidate having a motion vector predictor inferred from the adjacent blocks can replace a spatial or temporal candidate in the list rather than extending the list. It is also possible to include multiple motion vector predictors derived from the base layer data in the motion vector predictor candidate list.

[0253] The second part relates to extending the list of merge candidates by base layer candidates (see Aspect K). Motion data of one or more juxtaposed blocks of the base layer is added to the merge candidate list. This method enables the possibility of creating a merge area that shares specific motion parameters across the base layer and the enhancement layer. Similar to the previous section, as represented in FIG. 38, the base layer blocks covering the samples juxtaposed at the central position can be obtained from any position in the immediate vicinity, rather than being restricted to this central position. If any motion data is not available or accessible for a given position, an alternative position can be selected to infer possible merge candidates. Before the derived motion data is inserted into the merge candidate list, it can be scaled according to the resolution ratio. The index describing the merge candidate list is transmitted and defines the motion vector. It is used for motion compensation prediction. However, the method can also suppress possible motion predictor candidates that depend on the motion data of the prediction blocks within the base layer. In an embodiment, the motion vector predictor of the collocated blocks within the base layer covering the sample position C1 in FIG. 38 is considered a possible merge candidate for encoding the current prediction block in the enhancement layer. However, if the "merge_flag" (merge flag) of the reference block is equal to 1, or if the collocated reference block contains no motion data at all, the motion vector predictor is not inserted into the list. In any other case, the obtained motion vector predictor is added to the merge candidate list as the second entry. Note that in this embodiment, the length of the merge candidate list is maintained and not extended. In another embodiment, as represented in FIG. 38, one or more motion vector predictors are derived from the prediction blocks covering any of the sample positions such that they are added to merge the candidate lists. In another embodiment, one or several motion vector predictors of the base layer can be added to the merge candidate list at any position. In another embodiment, if certain constraints are met, one or more motion vector predictors are only added to the merge candidate list. Such constraints include the size of the prediction block in the enhancement layer that matches the size of the collocated blocks in the base layer (regarding the resolution ratio described in the section of the previous embodiment for motion vector prediction). Another constraint within another embodiment is a "merge_flag" value equal to 1. In another embodiment, the length of the merge candidate list can be extended by the number of motion vector predictors inferred from the collocated reference blocks in the base layer.

[0254] The third part of this specification is for reordering the motion parameter (or merge) candidate list using the base layer data (see Example L), and describes the process of reordering the merge candidate list according to the already encoded information in the base layer. If the collocated base layer block covering the samples of the current block is a motion compensated prediction having candidates derived from a particular original, the corresponding enhancement layer candidate from the equivalent original (if it exists) is placed at the head of the merge candidate list as the first entry. This step is equivalent to describing this candidate with the lowest index. The lowest index assigns the simplest codeword to this candidate. In an embodiment, collocated base layer blocks are motion compensated predicted along with candidates generated from a prediction block covering sample position A1, as represented in FIG. 38. If the merge candidate list of prediction blocks in the enhancement layer includes a candidate generated from the corresponding sample position A1 in the enhancement layer by the motion vector predictor, this candidate is placed in the list as the first entry. As a result, this candidate is indexed by index 0 and thus is assigned the shortest fixed length codeword. In this embodiment, this step is performed after the derivation of the motion vector predictor of the collocated base layer blocks with respect to the merge candidate list in the enhancement layer. Thus, the reordering process assigns the lowest index to the candidate generated from the corresponding block as the motion vector predictor of the collocated base layer blocks. The second lowest index is assigned to the candidate derived from the collocated blocks in the base layer, as explained in the second part of this section. Moreover, the reordering process is performed only when the "merge_flag" of the collocated blocks in the base layer is equal to 1. In another embodiment, the reordering process may be performed regardless of the value of the "merge_flag" of the collocated prediction blocks in the base layer. In another embodiment, a candidate with the corresponding original motion vector predictor may be placed at any position in the merge candidate list. In another embodiment, the reordering process may remove all other candidates in the merge candidate list. Here, only the candidate having the same original as the motion vector predictor used for motion compensated prediction of the collocated blocks in the base layer remains in the list. In this case, a single candidate is utilized and no index is transmitted at all.

[0255] The fourth part of this specification is for reordering the motion vector predictor candidate list using base layer data (refer to Aspect L), and the process of reordering the candidate list for motion vector prediction using the motion parameters of the base layer block is taken as the embodiment. If the collocated base layer block covering the samples of the current prediction block uses a motion vector from a specific origin, then the motion vector predictor from the corresponding origin in the enhancement layer is used as the first entry in the motion vector predictor list of the current prediction block. This results in assigning the least expensive codeword to this candidate. In an embodiment, the collocated base layer block is motion compensated predicted along with the candidates generated from the prediction block covering the sample position A1, as shown in FIG. 38. If the motion vector predictor candidate list of the block in the enhancement layer includes a candidate where the motion vector predictor is generated from the corresponding sample position A1 in the enhancement layer, then this candidate is placed as the first entry in the list. As a result, this candidate is indexed by index 0 and thus is assigned the shortest fixed-length codeword. In this embodiment, this step is performed after the derivation of the motion vector predictor of the collocated base layer block for the motion vector predictor list in the enhancement layer. Thus, the reordering process assigns the lowest index to the candidate generated from the corresponding block as the motion vector predictor of the collocated base layer block. The second lowest index is assigned to the candidate derived from the collocated blocks in the base layer as described in the first part of this section. Moreover, the reordering process is performed only when the "merge_flag" of the collocated blocks in the base layer is equal to 0. In another embodiment, the reordering process can be performed regardless of the value of the "merge_flag" of the collocated prediction blocks in the base layer. In another embodiment, the candidate with the motion vector predictor of the corresponding origin can be placed at any position in the motion vector predictor candidate list.

[0256] The following relates to the enhancement layer coding of the conversion coefficient.

[0257] In the state-of-the-art video and image coding, the residual of the prediction signal is previously converted and the resulting quantized conversion coefficients are signaled in the bitstream. This coefficient coding follows a fixed scheme.

[0258] Depending on the conversion size (for luma residuals: 4×4, 8×8, 16×16, 32×32), different scanning directions are defined. The first and last positions in the scanning order are given, and these scans uniquely determine which coefficient positions may be significant and, as a result, need to be coded. In all scans, the first coefficient is set to be the DC coefficient at position (0,0), and the last position has to be signaled in the bitstream, which is done by coding the (horizontal) x and (vertical) y positions within the conversion block. Starting from the last position, the signaling of the significant coefficients is done in reverse scanning order until the DC position is reached.

[0259] For conversion sizes 16×16 and 32×32, only one scan, namely the "diagonal scan", is defined. However, conversion blocks of sizes 2×2, 4×4, and 8×8 can additionally utilize "vertical" and "horizontal" scans. However, the use of vertical and horizontal scans is restricted to the residuals of the intra prediction coding unit. And the scan actually used is derived from the indication mode of that intra prediction. Indication modes with indices in the range of 6 to 14 result in a vertical scan. Indication modes with indices in the range of 22 to 30 result in a horizontal scan. All residual indication modes result in a diagonal scan.

[0260] Figure 39 shows the diagonal scan, vertical scan, and horizontal scan defined for a 4×4 transform block. Larger transform coefficients are divided into subgroups of 16 coefficients. These subgroups enable hierarchical coding of significant coefficient positions. Subgroups signaled as non-significant do not contain significant coefficients. The scans for 8×8 and 16×16 are represented respectively in Figures 40 and 41 along with their associated subgroup partitions. The large arrows represent the scan order of the coefficient subgroups.

[0261] In zigzag scanning, for blocks of size larger than 4×4, the subgroups consist of 4×4 pixel blocks scanned in zigzag. The subgroups are scanned in a zigzag manner. Figure 42 shows the vertical scan for 16×16 transform as proposed within JCTVC-G703.

[0262] The following paragraphs describe extensions to transform coefficient coding. These include the introduction of new scan modes (ways of assigning scans to transform blocks and modified coding of significant coefficient positions). These extensions allow for better adaptation to different coefficient distributions within the transform block, and as a result, achieve coding gain in terms of rate distortion.

[0263] New realizations for vertical-horizontal scan patterns are introduced for 16×16 and 32×32 transform blocks. In contrast to previously proposed scan patterns, the size of the scan subgroups is 16×1 for horizontal scans and 1×16 for vertical scans respectively. Also, subgroups with sizes of 8×2 and 2×8 can be selected respectively. The subgroups themselves are scanned in the same way.

[0264] Vertical scanning is efficient for transform coefficients located within column-wise spreads. This can be found in images containing horizontal edges.

[0265] Horizontal scanning is efficient for transform coefficients found within row-wise spreads. This is observed in images containing vertical edges.

[0266] Figure 43 shows the implementation of vertical-horizontal scanning for a 16×16 transform block. The coefficient subgroups are defined as one row or one column each. The vertical-horizontal scanning is the introduced scanning pattern. The scanning pattern enables the encoding of coefficients within a row by column-direction scanning. For a 4×4 block, the first row is scanned following the rest of the first column, then the rest of the second row is scanned, then the rest of the coefficients of the second column are scanned. Next, the rest of the third row is scanned, and finally the rest of the fourth column and row are scanned.

[0267] For larger blocks, the block is divided into 4×4 subgroups. These 4×4 blocks are scanned by vertical-horizontal scanning, and the subgroups are scanned by the vertical-horizontal scanning itself.

[0268] The vertical-horizontal scanning can be used when the coefficients are located in the first row and column within the block. In this way, the coefficients are scanned earlier than when using another scanning, such as diagonal scanning. This is found for images that include both horizontal edges and vertical edges.

[0269] Figure 44 shows vertical and horizontal scanning for a 16×16 transform block.

[0270] Other scans are similarly possible. For example, all combinations between scans and subgroups can be used. For example, horizontal scanning is used for 4×4 blocks and diagonal scanning is used for subgroups. An appropriate selection of scans can be applied by selecting different scans for each subgroup.

[0271] It should be stated that different scans can be realized in the way that the transform coefficients are rearranged after quantization on the encoder side and conventional encoding is used. On the decoder side, the transform coefficients are rearranged after conventional decoding and before scaling and inverse transform (or after scaling and before inverse transform).

[0272] Different parts of the base layer signal are utilized to derive encoding parameters from the base layer signal. The following are within the base layer signal. · Collocated reconstructed base layer signals · Collocated residual base layer signals · Estimated residual signal of the enhancement layer obtained by subtracting the enhancement layer prediction signal from the reconstructed base layer signal · Image partition of the base layer frame

[0273] [Gradient parameter] The gradient parameter can be derived as follows: For each pixel of the investigated block, a gradient is calculated. From these gradients, a magnitude and an angle are calculated. The most occurring angle within the block is the one associated with the block (block angle). The angle is rounded to use only three directions: horizontal (0°), vertical (90°), diagonal (45°).

[0274] [Edge detection] The edge detector can be applied to the investigated blocks as follows: First, the block is smoothed by an n×n smoothing filter (e.g., Gaussian). A gradient matrix of size m×m is used to calculate the gradient of each pixel. The magnitude and angle of every pixel are calculated. The angle is rounded to use only three directions: horizontal (0°), vertical (90°), diagonal (45°). For every pixel having a magnitude greater than a predetermined threshold 1, the adjacent pixels are checked. If an adjacent pixel has a magnitude greater than threshold 2 and has the same angle as the current pixel, the counter for this angle is incremented. For the entire block, the counter with the highest value is selected as the angle of the block.

[0275] [Obtaining base layer coefficients by previous transformation] For a particular TU, the examined and collocated signals (reconstructed base layer signal / residual base layer signal / estimated enhancement layer signal) are converted into the frequency domain in order to derive encoding parameters from the frequency domain of the base layer signal. Preferably, this is performed using the same transform as used by that particular enhancement layer TU. The resulting base layer transform coefficients may or may not be quantized. Rate-distortion quantization with modified lambda is used to obtain a coefficient distribution comparable to that of the enhancement layer block.

[0276] [Scanning validity score for specific distribution and scanning] The scanning validity score for a specific significant coefficient distribution is defined as follows: Let each position of the examined block be represented by an index in the order of the examined scan. Then, the sum of the index values of the significant coefficient positions is defined as the validity score of this scan. As a result, the smaller the score of the scan, the better the efficiency represented by the specific distribution.

[0277] [Selection of appropriate scanning pattern for transform coefficient coding] If several scans are available for a particular TU, a rule for uniquely selecting one of the scans needs to be defined.

[0278] [Method for scanning pattern selection] The selected scan can be directly derived from the already decoded signal (without transmitting any additional data). This is possible either based on the characteristics of the collocated base layer signal or by using only the enhancement layer signal. The scanning pattern can be derived from the EL signal as follows. · The aforementioned state-of-the-art derivation rules. · Using the scanning pattern for the chrominance residual selected for the collocated luminance residual. ·Define a fixed mapping between the symbolization mode and the scanning pattern used. ·Obtain the scanning pattern from the last significant coefficient position (proportional to the estimated fixed scanning pattern). ·In a preferred embodiment, the scanning pattern is selected depending on the last position already decoded as follows:

[0279] The last position is represented as x and y coordinates within the transformation block and is already decoded (for a scan depending on the last encoded, a fixed scanning pattern is estimated for the decoding process of the last position. It can be the leading scanning pattern of that TU). Let T be a defined threshold depending on a specific transformation size. If not only the x coordinate but also the y coordinate of the last significant position does not exceed T, a diagonal scan is selected.

[0280] Otherwise, x is compared with y. If x exceeds y, a horizontal scan is selected and a vertical scan is not selected. The preferred value of T for a 4×4 TU is 1. The preferred value of T for a TU larger than 4×4 is 4.

[0281] In another preferred embodiment, the derivation of the scanning pattern, as described within the previous embodiment, is restricted to only for TUs of sizes 16×16 and 32×32. It can be further restricted to only the luminance signal.

[0282] Also, the scanning pattern is derived from the BL signal. Any of the encoding parameters described above can be used to derive the scanning pattern selected from the base layer signal. In particular, the gradient of the juxtaposed base layer signals can be calculated and compared with a predefined threshold and / or potentially discovered edges can be utilized.

[0283] In a preferred embodiment, the scanning direction is derived depending on the block gradient angle as follows. For a gradient quantized in the horizontal direction, vertical scanning is used. For a gradient quantized in the vertical direction, horizontal scanning is used. Otherwise, diagonal scanning is selected.

[0284] In another preferred embodiment, the scanning pattern is obtained as described in the previous embodiment. However, it is obtained only for those transform blocks where the number of occurrences of the block angle exceeds a threshold. The remaining transform units are decoded using the leading scanning pattern of the TU.

[0285] If the base layer coefficients of juxtaposed blocks are valid and are clearly signaled within the base layer data stream or are calculated by a previous transformation, the base layer coefficients can be utilized in the following ways. · For each available scan, the cost for encoding the base layer coefficients can be evaluated. The scan with the lowest cost is used to decode the enhancement layer coefficients. · The effective score of each available scan is calculated for the base layer coefficient distribution. The scan with the smallest score is used to decode the enhancement layer coefficients. · The distribution of the base layer coefficients within a transform block is classified into one of a predefined set of distributions associated with a particular scan pattern. · The scan pattern is selected depending on the last significant base layer coefficient.

[0286] If juxtaposed base layer blocks are predicted using intra prediction, the intra direction of that prediction can be used to derive the enhancement layer scan pattern.

[0287] Moreover, the transform size of juxtaposed base layer blocks is utilized to derive the scan pattern.

[0288] In a preferred embodiment, the scanning pattern is derived from the BL signal only for the TUs representing the residual of the INTRA_COPY mode prediction block. And those collocated base layer blocks are intra-predicted. For those blocks, a modified state-of-the-art scan selection is used. In contrast to the state-of-the-art scan selection, the intra-prediction direction of the collocated base layer blocks is used to select the scanning pattern.

[0289] Signaling the scan pattern index within the bitstream (see aspect R). The scan pattern of the transform block can be selected by the encoder in terms of rate distortion and then signaled within the bitstream.

[0290] A particular scan pattern can be encoded by signaling an index to a list of available scan pattern candidates. This list can be a fixed list of scan patterns defined for a particular transform size or can be filled dynamically during the decoding process. Filling the list dynamically allows for a suitable selection of those scan patterns, and the scan pattern can encode a particular coefficient distribution, perhaps most efficiently. By doing so, the number of available scan patterns for a particular TU can be reduced. And as a result, signaling the index to that list is less costly. If the number of scan patterns in a particular list is reduced to one, signaling is not necessary. For a particular TU, the process of selecting scan pattern candidates may utilize any of the encoding parameters described above and / or follow a predetermined rule that utilizes particular characteristics of that particular TU. Among them are the following. · The TU represents the residual of the luminance / color signal. · The TU has a particular size. · The TU represents the residual of a particular prediction mode. · The last significant position within the TU is known to the decoder and belongs to a particular sub-division of the TU. ·TU is a part of one I / B / P - slice. ·The coefficients of TU are quantized using specific quantization parameters.

[0291] In a preferred embodiment, the list of scan pattern candidates includes three scans for all TUs: "diagonal scan", "vertical scan" and "horizontal scan".

[0292] Another embodiment can be obtained by including any combination of scan patterns in the candidate list.

[0293] In a particular preferred embodiment, the list of scan pattern candidates may include any of "diagonal scan", "vertical scan" and "horizontal scan".

[0294] However, the scan pattern selected by the state - of - the - art scan derivation (described above) is initially set to be within the list. Another candidate is added to the list only if a particular TU has a size of 16×16 or 32×32. The order of the remaining scan patterns depends on the last significant coefficient position.

[0295] (Note: The diagonal scan is always the first pattern in the list for estimating 16×16 and 32×32 conversions)

[0296] If the magnitude of the x - coordinate exceeds the magnitude of the y - coordinate, the horizontal scan is then selected. And the vertical scan is placed in the last position. Otherwise, the vertical scan is placed in the second position followed by the horizontal scan.

[0297] Another preferred embodiment is obtained by further restricting the conditions because there is one or more candidates in the list.

[0298] In another embodiment, if the coefficients of the transform block represent the residual of the luminance signal, the vertical and horizontal scans are only added to the candidate list for 16×16 and 32×32 transform blocks.

[0299] In another embodiment, if both the x and y coordinates of the last significant position are greater than a specific threshold, the vertical and horizontal scans are added to the candidate list of transform blocks. This threshold can be size-dependent mode and / or TU. The preferred threshold is 3 for all sizes larger than 4×4, and 1 for 4×4 TU.

[0300] In another embodiment, if either the x or y coordinate of the last significant position is greater than a specific threshold, the vertical and horizontal scans are only added to the candidate list of transform blocks. This threshold can be size-dependent mode and / or TU. The preferred threshold is 3 for all sizes larger than 4×4, and 1 for 4×4 TU.

[0301] In another embodiment, if both the x and y coordinates of the last significant position are greater than a specific threshold, the vertical and horizontal scans are only added to the candidate lists of 16×16 and 32×32 transform blocks. This threshold can be size-dependent mode and / or TU. The preferred threshold is 3 for all sizes larger than 4×4, and 1 for 4×4 TU.

[0302] In another embodiment, if either the x or y coordinate of the last significant position is greater than a specific threshold, the vertical and horizontal scans are only added to the candidate lists of 16×16 and 32×32 transform blocks. This threshold can depend on the mode and / or TU size. The preferred threshold is 3 for all sizes larger than 4×4, and 1 for 4×4 TU.

[0303] For any of the described embodiments, a specific scan pattern is signaled within the bitstream. The signaling itself can be done at different signaling levels. In particular, the signaling can be done at the CU / LCU level, or at the slice level, for any node of the residual quad-tree (all sub-TUs of that node that use the signaled scan and the same candidate list index), or for each TU that is reduced within a subgroup of TUs having the signaled scan pattern.

[0304] The index into the candidate list can be sent using fixed-length coding, variable-length coding, arithmetic coding (including context-adaptive binary arithmetic coding), or PIPE coding. If context-adaptive coding is used, the context can be derived based on the parameters of adjacent blocks, the coding modes described above, and / or the characteristics of the specific TU itself.

[0305] In a preferred embodiment, context-adaptive coding is used to signal the index into the scan pattern candidate list for a TU. However, the context model is derived based on the transform size and / or position of the last significant position within the TU. Any of the methods described above for deriving the scan pattern can be used to derive a context model for signaling an explicit scan pattern for a specific TU.

[0306] The following changes can be used within the enhancement layer to code the last significant scan position. · Separate context models are used for all or a subset of the coding modes using base layer information. It is also possible to use different context models for different modes having base layer information. · The context model can depend on the data within the collocated base layer blocks (e.g., the transform coefficient distribution in the base layer, the gradient information of the base layer, the last scan position within the collocated base layer blocks). · The last scan position can be coded as the difference from the last base layer scan position. · If the last scan position is coded by signaling its x and y positions within the TU, the context model for the second signaled coordinate can depend on the value of the first signaling. · Any of the above methods can be used to derive a context model for signaling the last significant position in order to derive a scan pattern independent of the last significant position.

[0307] In a particular version, scan pattern derivation depends on the last significant position: · If the last scan position is coded by signaling its x and y positions within the TU, the context model for the second coordinate can depend on those scan patterns that are still possible candidates when the first coordinate is already known. · If the last scan position is coded by signaling its x and y positions within the TU, the context model for the second coordinate can depend on whether the scan pattern has already been uniquely selected when the first coordinate is already known.

[0308] In another version, scan pattern derivation is independent of the last significant position. · The context model can depend on the scan pattern used within a particular TU. · Any of the methods described above for deriving the scan pattern can be used to derive a context model for signaling the last significant position.

[0309] For coding the significant positions and significant flags (subgroup flags and / or significant flags for one transform coefficient) within the TU, respectively, the following changes can be used within the enhancement layer: ·Separate context models are used for all or a subset of the coding modes that use the base layer information. It is also possible to use different context models for different modes having the base layer information. ·The context model can depend on data within juxtaposed base layer blocks (e.g., the number of significant transform coefficients for a particular frequency position). ·Any of the methods described above for deriving the scan pattern can be used to derive the context model for signaling significant positions and / or their levels. ·A generalized template is used that evaluates both the significant number of already - coded transform coefficient levels within the spatial neighborhood of the coefficient to be coded and the number of significant transform coefficients within the juxtaposed base layer signals at similar frequency positions. ·A generalized template can be used that evaluates both the significant number of already - coded transform coefficient levels within the spatial neighborhood of the coefficient to be coded and the levels of the significant transform coefficients within the juxtaposed base layer signals at similar frequency positions. ·The context modeled for the subgroup flag can depend on the scan pattern used and / or the particular transform size.

[0310] The usage of different context initialization tables for the base layer and the enhancement layer can be used. The context model initialization for the enhancement layer can be changed in the following ways. ·The enhancement layer uses separate sets of initialization values. ·The enhancement layer uses separate sets of initialization values for different operation modes (spatial / temporal, or quality scalability). ·An enhancement layer context model having corresponding parts within the base layer can use the states of those corresponding parts as the initialization state. ·An algorithm for deriving the initial state of the context model can depend on the base layer QP and / or the delta QP.

[0311] Next, using the base layer data, the possibility of subsequent adaptive enhancement layer encoding is described. The following part describes a method for generating an enhancement layer prediction signal within a scalable video coding system. The method uses the base layer decoded for sample information of an image to infer the value of a prediction parameter, which is not transmitted in the encoded video bitstream but is used to form a prediction signal for the enhancement layer. Accordingly, the overall bitrate required to encode the enhancement layer signal is reduced.

[0312] State-of-the-art hybrid video encoders typically decompose the original (source) image into blocks of different sizes according to a hierarchical structure. For each block, the video signal is predicted from spatially adjacent blocks (intra prediction) or from previously temporally encoded images (inter prediction). The difference between the prediction and the actual image is transformation and quantization. The resulting prediction parameters and transform coefficients are entropy encoded to form the encoded video bitstream. A compliant decoder follows the steps in reverse order... A scalable video encoding the bitstream consists of various layers: a base layer providing a fully decodable video and an enhancement layer that can be additionally used for decoding. The enhancement layer can provide higher spatial resolution (spatial scalability), temporal resolution (temporal scalability) or quality (SNR scalability). In previous standards such as H.264 / AVC SVC, syntax elements such as motion vectors, reference picture indices or intra prediction modes are directly predicted from corresponding syntax elements within the encoded base layer. Within the enhancement layer, a mechanism exists to switch between them using a prediction signal derived from the base layer syntax elements, or predicted from another enhancement layer syntax element or decoded enhancement layer samples, at the block level.

[0313] In the following part, the base layer data is used to derive the enhancement layer parameters on the decoder side.

[0314] [Method 1: Derivation of Motion Parameter Candidates] For a block (a) of the image in the spatial or quality enhancement layer, the corresponding block (b) of the base layer image is determined, which covers the same image area. The inter-prediction signal for the block (a) of the enhancement layer is formed using the following method: 1. A set of motion compensation parameter candidates is determined from, for example, temporally or spatially adjacent enhancement layer blocks or derivatives thereof. 2. Motion compensation is performed for each candidate set of motion compensation parameters to form an inter-prediction signal within the enhancement layer. 3. The best set of motion compensation parameters is selected by minimizing the magnitude of the error between the prediction signal for the enhancement layer block (a) and the reconstruction signal of the base layer block (b). For spatial scalability, the base layer block (b) can be spatially upsampled using an interpolation filter.

[0315] A set of motion compensation parameters includes a specific combination of motion compensation parameters.

[0316] Motion compensation parameters can be a motion vector, a reference image index, a selection between one and two predictions and another parameter.

[0317] In an alternative embodiment, candidates for the motion compensation parameter set are used from the base layer block. Also, inter prediction is performed within the base layer (using the reference picture of the base layer). To apply the error magnitude, the base layer block (b) reconstruction signal can be used directly without being upsampled. The selected optimal motion compensation parameter set is applied to the reference picture of the enhancement layer to form the prediction signal for block (a). When the motion vector is applied within the spatial enhancement layer, the motion vector is scaled according to the resolution change. Both the encoder and the decoder can perform the same prediction steps to select an optimal motion compensation parameter set from the available candidates to create the same prediction signal. These parameters are not signaled within the coded video bitstream.

[0318] The selection of the prediction method can be signaled within the bitstream and coded using entropy coding. Within a hierarchical block sub-division structure, this coding method can be selected for any sub-level or alternatively only for a subset of the coding hierarchy. In an alternative embodiment, the encoder can send a refined motion parameter set prediction signal to the decoder. The refinement signal differentially includes the coded values of the motion parameters. The refinement signal can be entropy coded.

[0319] In an alternative embodiment, the decoder generates a list of best candidates. The index of the motion parameter set used is signaled within the coded video bitstream. The index is entropy coded. In an example, the list can be ordered by increasing error magnitude.

[0320] Examples generate motion compensation parameter set candidates using the adaptive motion vector prediction (AMVP) candidate list of HEVC. Another example generates motion compensation parameter set candidates using the merge mode candidate list of HEVC.

[0321] [Method 2: Motion Vector Derivation] For a block (a) of an image in a spatial or quality enhancement layer, a corresponding block (b) of the base layer image covering the same image region is determined.

[0322] The inter-prediction signal for block (a) of the enhancement layer is formed using the following method: 1. A motion vector predictor is selected. 2. An estimation of motion for a defined set of search positions is performed on the enhancement layer reference image. 3. For each search position, the magnitude of the error is determined and the motion vector having the smallest error is selected. 4. The prediction signal for block (a) is formed using the selected motion vector.

[0323] In an alternative embodiment, the search is performed on the reconstructed base layer signal. For spatial scalability, the selected motion vector is scaled according to the spatial resolution change before generating the prediction signal in step 4.

[0324] The search positions may be at full resolution or sub-pel resolution. Also, the search can be performed in multiple steps, for example, first determining the best full-pel position followed by another set of candidates based on the selected full-pel position. For example, the search can be terminated early when the magnitude of the error is below a defined threshold.

[0325] Both the encoder and the decoder can perform the same prediction steps to select the optimal motion vector within the candidates and generate the same prediction signal. These vectors are not signaled in the coded video bitstream.

[0326] The selection of the prediction method can be signaled within the bitstream and encoded using entropy coding. Within a hierarchical block sub-division structure, this coding method can be selected within any sub-level or only a subset of alternative coding levels. In an alternative embodiment, the encoder can send an improved motion vector prediction signal to the decoder. The refinement signal can be entropy coded.

[0327] The example selects a motion vector predictor using the algorithm described in Method 1.

[0328] Another example selects a motion vector predictor from temporally or spatially adjacent blocks in the enhancement layer using the Adaptive Motion Vector Prediction (AMVP) method of HEVC.

[0329] [Method 3: Intra Prediction Mode Derivation] For each block (a) in the enhancement layer (n) image, a corresponding block (b) is determined that covers the same region in the reconstructed base layer (n-1) image.

[0330] In a scalable video decoder, for each base layer block (b), an intra prediction signal is formed using an intra prediction mode (p) inferred by the following algorithm. 1) The intra prediction signal is generated for each available intra prediction mode according to the rules for intra prediction in the enhancement layer, but using sample values from the base layer. 2) The best prediction mode (p best ) is determined by minimizing the magnitude of the error (e.g., the sum of absolute differences) between the intra prediction signal and the decoded base layer block (b). 3) The prediction (p best ) mode selected in step 2) is used to generate a prediction signal for the enhancement layer block (a) according to the intra prediction rules for the enhancement layer.

[0331] Both the encoder and the decoder can execute the same steps to select the best prediction mode (p best ) and form a consistent prediction signal. Therefore, the actual intra prediction mode (p best ) is not signaled in the encoded video bitstream.

[0332] The selection of the prediction method can be signaled in the bitstream and encoded using entropy coding. Within a hierarchical block sub-division structure, this coding mode can be selected within any sub-level or, alternatively, only for a subset of the coding hierarchy. An alternative embodiment uses samples from the enhancement layer in step 2) to generate an intra prediction signal. For a spatially scalable enhancement layer, the base layer can be upsampled using an interpolation filter to apply the magnitude of the error.

[0333] An alternative embodiment divides the enhancement layer block into a plurality of blocks of a smaller block size (a i ) (e.g., a 16×16 block (a) can be divided into 16 4×4 blocks (a i )). The algorithms described above are applied to each sub-block (a i ) and the corresponding base layer block (b i ). After prediction of the block (a i ), residual coding is applied, and the result is used to predict the block (a i+1 ).

[0334] An alternative embodiment determines the predicted intra prediction mode (p i ) using the sample values around (b) or (b best ). For example, when a 4×4 block (a i ) of the spatial enhancement layer (n) has a corresponding 2×2 base layer block (b i ), the samples around (b i ) are the predicted intra prediction mode (pbest ) for determining 4×4 blocks (c i ) are used to form.

[0335] In an alternative embodiment, the encoder can send a refinement intra prediction direction signal to the decoder. For example, in a video codec such as HEVC, most intra prediction modes correspond to the angles at which the boundary pixels are used to form the prediction signal. The offset to the optimal mode can be sent as the difference to the predicted prediction mode (p best ) determined as described above. The improved mode can be entropy encoded.

[0336] Intra prediction modes are usually encoded depending on their probabilities. In H.264 / AVC, the most likely mode is determined based on the mode used in the (spatial) neighborhood of the block. In the HEVC list, the most likely mode is created. These most likely modes can be selected using fewer symbols in the bitstream than the total number of modes requires. An alternative embodiment uses the predicted intra prediction mode (p best ) for block (a) (determined as described in the above algorithm) as the most likely mode, or as a member of a list of most likely modes.

[0337] [Method 4: Intra Prediction Using a Boundary Region] In a scalable video decoder for forming an intra prediction signal for a block (a) (see FIG. 45) of a scalable or quality enhancement layer, the lines of samples (b) from the surrounding region of the same layer are used to be filled within the block region. These samples are obtained from the already encoded region (usually, however, not necessary on the upper and left boundaries).

[0338] The following alternative variations for selecting these pixels can be used. a) If the pixels in the surrounding area are not yet encoded, the pixel values are not used to predict the current block. b) If the pixels in the surrounding area are not yet encoded, the pixel values are derived from the already encoded adjacent pixels (e.g., by iteration). c) If the pixels in the surrounding area are not yet encoded, the pixel values are derived from the pixels in the corresponding area of the decoded base layer image.

[0339] To form the intra prediction of block (a), the adjacent lines of the pixels (b) derived as described above are used as a template, and each line (a j ) in block (a) is filled.

[0340] The lines (a j ) of block (a) are filled one by one along the x-axis. To achieve the best possible prediction signal, the columns of the template samples (b) are shifted along the y-axis to form the prediction signal (b´ j ) for the relevant lines (a j ).

[0341] To find the optimal prediction within each line, the shift offset (o j ) is determined by minimizing the magnitude of the error between the resulting prediction signal (a j ) and the sample values of the corresponding line in the base layer.

[0342] If (o j ) is a non-integer value, an interpolation filter can be used to map the values of (b) to the integer sample positions of (a j ), as shown in (b´7).

[0343] If spatial scalability is used, an interpolation filter can be used to create a matching number of sample values of the corresponding line in the base layer.

[0344] The filling direction (x-axis) can be horizontal (left - right), vertical (up - down), diagonal, or any other angle. The samples used for the template line (b) are the samples directly adjacent to the block along the x-axis. The template line (b) migrates along the y-axis which forms an angle of 90° with respect to the x-axis.

[0345] To find the optimal direction of the x-axis, a complete intra prediction signal is generated for the block (a). The angle with the minimum error magnitude between the prediction signal and the corresponding base layer block is selected. The number of possible angles can be limited.

[0346] Both the encoder and the decoder execute the same algorithm to determine the best prediction angle and offset. No explicit angle information or offset information needs to be signaled within the bitstream. In an alternative embodiment, only the samples of the base layer image are used to determine the offset (o i ).

[0347] In an alternative embodiment, an improvement of the predicted offset (o i ), e.g., the difference value, is signaled within the bitstream. Entropy coding can be used to code the improved offset value.

[0348] In an alternative embodiment, an improvement of the predicted direction, e.g., the difference value, is signaled within the bitstream. Entropy coding can be used to code the improved direction value.

[0349] If a line (b´ j ) is used for prediction, an alternative embodiment uses a threshold to select. If the error magnitude for the optimal offset (o j ) is less than the threshold, the line (c i ) is used to determine the value of the block line (a j ). If the optimal offset (oj ) If the magnitude of the error for () is equal to or greater than the threshold value, the (upsampled) base layer signal is used to determine the value of block line a j ) is used.

[0350] [Method 5: Another Prediction Parameter] Another prediction information is inferred in the same manner as Methods 1 to 3, for example, for partitioning blocks within a sub-block.

[0351] For block (a) of the image in the spatial or quality enhancement layer, the corresponding block (b) of the base layer image is determined. It covers the same image area.

[0352] The prediction signal for block (a) of the enhancement layer is formed using the following method. 1) The prediction signal is generated for each possible value of the tested parameter. 2) The best prediction mode (p best ) is determined by minimizing the magnitude of the error (e.g., the sum of absolute differences) between the prediction signal and the decoded base layer block (b). 3) The prediction (p best ) mode selected in step 2) is used to generate the prediction signal for the enhancement layer block (a).

[0353] Both the encoder and the decoder can execute the same prediction steps to select the optimal prediction mode within the possible candidates and generate the same prediction signal. The actual prediction mode is not signaled in the coded video bitstream.

[0354] The selection of the prediction method can be signaled in the bitstream and encoded using entropy coding. In a hierarchical block sub-division structure, this coding method can be alternatively selected within every sub-level or only for a subset of the coding hierarchy.

[0355] The following description briefly summarizes some of the above embodiments.

[0356] [Enhanced layer encoding having multiple methods for generating an intra prediction signal using samples of a reconstructed base layer] Main Example: For encoding blocks within the enhancement layer, multiple methods for generating an intra prediction signal using samples of the reconstructed base layer are provided in addition to a method for generating a prediction signal based only on samples of the reconstructed enhancement layer.

[0357] Sub - Example: · The multiple methods include the following. The (upsampled / filtered) reconstructed base layer signal is used directly as the enhancement layer prediction signal. · The multiple methods include the following. The (upsampled / filtered) reconstructed base layer signal is combined with a spatial intra prediction signal. Therein, the spatial intra prediction is derived based on differential samples for adjacent blocks. The differential samples represent the difference between the reconstructed enhancement layer signal and the (upsampled / filtered) reconstructed base layer signal (see Mode A). · The multiple methods include the following. A conventional spatial intra prediction signal (obtained using samples of an adjacent reconstructed enhancement layer) is combined with the (upsampled / filtered) base layer residual signal (the inverse transform of the base layer transform coefficients, or the difference between the base layer reconstruction and the base layer prediction) (see Mode B). · The multiple methods include the following. The (upsampled / filtered) reconstructed base layer signal is combined with a spatial intra prediction signal. Therein, the spatial intra prediction is derived based on samples of the reconstructed enhancement layer of adjacent blocks. The final prediction signal is obtained by weighting the spatial prediction signal and the base layer prediction signal in a way that different frequency components use different weightings (see Mode C1). This can be achieved, for example, by any of the following. ○ Filter the base layer prediction signal by a low-pass filter, filter the spatial intra prediction signal by a high-pass filter, and add the resulting filtered signals (see Mode C2). ○ Convert the base layer prediction signal and the enhancement layer prediction signal, and superimpose the resulting conversion blocks. Different weighting coefficients are used for different frequency positions there (see Example C3). The resulting conversion blocks can be inverse-transformed and used as the enhancement layer prediction signal. Alternatively, the resulting conversion coefficients are added to the scaled transmitted conversion coefficient levels and then inverse-transformed to obtain a block reconstructed before deblocking and in-loop processing (see Mode C4). · For the method of using the reconstructed base layer signal, the following versions can be used. This can be fixed, or it can be signaled at the sequence level, picture level, slice level, maximum coding unit level, coding unit level. Alternatively, it can be created depending on other coding parameters. ○ Samples of the reconstructed base layer before deblocking and in-loop processing (such as sample adaptive offset filter or adaptive loop filter). ○ Samples of the reconstructed base layer after deblocking and before in-loop processing (such as sample adaptive offset filter or adaptive loop filter). ○ Samples of the reconstructed base layer after deblocking and further in-loop processing (such as sample adaptive offset filter or adaptive loop filter), or after multiple in-loop processing steps (see Mode D). · Multiple versions of a method using an (upsampled / filtered) base layer signal may be used. The upsampled / filtered base layer signals employed for these versions may differ within the interpolation filter used (including an interpolation filter that filters integer sample positions). Alternatively, the upsampled / filtered base layer signal for a second version may be obtained by filtering the upsampled / filtered base layer signal for the first version. One selection of the different versions may be signaled at the slice level, picture level, slice level, maximum coding unit level, coding unit level. It may be inferred from the characteristics of the corresponding reconstructed base layer signal, or the transmitted coding parameters (see Mode E). · Different filters may be used to upsample / filter the reconstructed base layer signal (see Mode E) and the base layer residual signal (see Mode F). · For a base layer block for which the residual signal is zero, it may be replaced by another signal derived from the base layer (e.g., a high-pass filtered version of the reconstructed base layer block) (see Mode G). · For a mode using spatial intra prediction, adjacent samples not available within the enhancement layer (in a particular coding order) may be replaced by the corresponding samples of the upsampled / filtered base layer signal (see Mode H). · For a mode using spatial intra prediction, the coding of the intra prediction mode may be changed. The list of most likely modes includes the intra prediction modes of the collocated base layer signal. · In a specific version, the image of the enhancement layer is decoded within a two-stage process. In the first stage, only the blocks that use only the base layer signal for prediction (without using adjacent blocks), or the inter-prediction signal, are decoded and reconstructed. In the second stage, the remaining blocks that use adjacent samples for prediction can be reconstructed. For the blocks reconstructed in the second stage, the spatial intra-prediction concept is extended (see Aspect I). Based on the usefulness of the already reconstructed blocks, not only the samples adjacent to the upper or left side of the current block but also the samples adjacent to the lower or right side can be used for spatial intra-prediction.

[0358] [Enhancement layer coding with multiple methods for generating an inter-prediction signal using samples of the reconstructed base layer] Main embodiment: For encoding blocks within the enhancement layer, a multiple method for generating an inter-prediction signal using samples of the reconstructed base layer is provided in addition to the method of generating a prediction signal based only on samples of the reconstructed enhancement layer.

[0359] Sub-embodiment: · The multiple method includes the following methods. The conventional inter-prediction signal (derived by motion-compensated interpolation of the already reconstructed enhancement layer image) is combined with the (upsampled / filtered) base layer residual signal (the inverse transform of the base layer transform coefficients, or the difference between the base layer reconstruction and the base layer prediction). · The multiple method includes the following methods. The (upsampled / filtered) reconstructed base layer signal is combined with the motion-compensated prediction signal. Therein, the motion-compensated prediction signal is obtained by the motion-compensated difference image. The difference image represents the difference between the reconstructed enhancement layer signal and the (upsampled / filtered) reconstructed base layer signal with respect to the reference image (see Example J). · The multiple method includes the following methods. The (upsampled / filtered) reconstructed base layer signal is combined with the inter prediction signal. Therein, the inter prediction is derived by motion compensation prediction using the reconstructed enhancement layer image. The final prediction signal is obtained by weighting the inter prediction signal and the base layer prediction signal in a way that different frequency components use different weightings (see Example C). This can be achieved, for example, by any of the following. ○ Filtering the base layer prediction signal with a low-pass filter, filtering the inter prediction signal with a high-pass filter, and adding the obtained filtered signals. ○ Transforming the base layer prediction signal and the inter prediction signal, and overlapping the obtained transform blocks. Therein, different weighting coefficients are used for different frequency positions. The obtained transform blocks are inverse-transformed to obtain a reconstructed block before deblocking and in-loop processing and used as the enhancement layer prediction signal, or the obtained transform coefficients are added to the scaled transmitted transform coefficient levels and then inverse-transformed. · For the method using the reconstructed base layer signal, the following versions can be used. This can be fixed, or it can be signaled at the sequence level, picture level, slice level, maximum coding unit level, coding unit level. Or it can be created depending on another coding parameter. ○ Samples of the reconstructed base layer before deblocking and in-loop processing (such as sample adaptive offset filter or adaptive loop filter). ○ Samples of the reconstructed base layer after deblocking and before performing in-loop processing (such as sample adaptive offset filter or adaptive loop filter). ○ Samples of the reconstructed base layer after deblocking and in-loop processing (such as sample adaptive offset filter or adaptive loop filter), or during multiple in-loop processing steps (see Aspect D). ·For a base layer block where the residual signal is zero, it can be replaced by another signal derived from the base layer (e.g., a high-pass filtered version of the reconstructed base layer block) (see Example G). ·Multiple versions of a method using the (upsampled / filtered) base layer signal can be used. The upsampled / filtered base layer signals employed for these versions differ within the interpolation filter used (including an interpolation filter that filters integer sample positions). Or, the upsampled / filtered base layer signal for a second version can be obtained by filtering the upsampled / filtered base layer signal for a first version. One selection of the different versions can be signaled at the sequence level, picture level, slice level, maximum coding unit level, coding unit level. It can also be inferred from the corresponding reconstructed base layer signal or the characteristics of the transmitted coded parameters (see Example E). ·Different filters can be used to upsample / filter the reconstructed base layer signal (see Example E) and the base layer residual signal (see Example F). ·For motion compensated prediction of the difference image (the difference between the enhancement layer reconstruction and the upsampled / filtered base layer signal) (see Example J), different interpolation filters are used more for motion compensated prediction of the reconstructed image. ·For motion compensated prediction of the difference image (the difference between the enhancement layer reconstruction and the upsampled / filtered base layer signal) (see Example J), the interpolation filter is selected based on the characteristics of the corresponding region within the difference image (or based on coded parameters or based on information transmitted within the bitstream).

[0360] [Enhancement Layer Motion Parameter Coding] Main aspect: Use of a plurality of enhancement layer predictors and at least one predictor derived from the base layer for encoding enhancement layer motion parameters.

[0361] Sub - example: · Add the (scaled) base layer motion vector to the motion vector predictor list (see Example K). ○ Use of the base layer block covering the collocated samples at the center position of the current block (possible another derivation). ○ Scaled motion vector according to the resolution ratio. · Add the motion data of the collocated base layer block to the merge candidate list (see Example K). ○ Use of the base layer block covering the collocated samples at the center position of the current block (possible another derivation). ○ Scaled motion vector according to the resolution ratio. ○ If, within the base layer, the "merge_flag" is equal to 1, do not add. · Re - ordering of the merge candidate list based on the base layer merge information (see Example L) ○ If the collocated base layer block is merged into a specific candidate, the corresponding enhancement layer candidate is used as the first entry in the enhancement layer merge candidate list. · Re - ordering of the motion predictor candidate list based on the base layer motion predictor information (see Example L) ○ If the collocated base layer block uses a specific motion vector predictor, the corresponding enhancement layer motion vector predictor is used as the first entry in the enhancement layer motion vector predictor candidate list. · Derivation of the merge index (i.e., candidates for which the current block is to be merged) is based on the base layer information within the collocated blocks (see Example M). As an example, if, hypothetically, a base layer block is merged into a particular adjacent block and it is signaled within a bitstream that it also merges an enhancement layer block, then the merge index is not transmitted at all. Instead, the enhancement layer block is merged into the same adjacent block (but within the enhancement layer) as the collocated base layer block.

[0362] [Enhancement Layer Partitioning and Motion Parameter Inference] Main aspect: Inference of enhancement layer partitioning and motion parameters based on base layer partitioning and motion parameters (presumably, it may be required to combine this aspect with one of the sub - aspects).

[0363] Sub - aspects: · Derive motion parameters for N×M sub - blocks of the enhancement layer based on collocated base layer motion data. Group blocks having the same derived parameters (or parameters with a small difference) into larger blocks. Determine prediction and coding units. (See Example T) · Motion parameters may include the number of motion hypotheses, reference indices, motion vectors, motion vector predictor identifiers, merge identifiers. · Signal one of multiple methods for generating an enhancement layer prediction signal. Such methods may include the following. ○ Motion compensation using the derived motion parameters and the reconstructed reference image of the enhancement layer. ○ Combine (a) the (upsampled / filtered) base layer reconstruction for the current image, (b) the motion compensation signal using the obtained motion parameters, and the reference image of the enhancement layer generated by subtracting the (upsampled / filtered) base layer reconstruction from the reconstructed enhancement layer image. ○ Combine (a) the (upsampled / filtered) base layer residual (the difference between the reconstructed signal and the prediction, or the inverse transform of the encoded transform coefficient values) for the current image, (b) the motion compensation signal using the obtained motion parameters, and the reference image of the reconstructed enhancement layer. · If the juxtaposed blocks within the base layer are intra-coded, then the corresponding enhancement layer M×N blocks (or CUs) are also intra-coded. Therein, the intra-prediction signal is derived using the base layer information (see Mode U). For example, ○ The (upsampled / filtered) version of the corresponding base layer reconstruction is used as the intra-prediction signal (see Mode U). ○ The intra-prediction mode is derived based on the intra-prediction mode used within the base layer. And this intra-prediction mode is used for spatial intra-prediction within the enhancement layer. · If the juxtaposed base layer blocks for the M×N enhancement layer blocks (sub-blocks) are merged with previously encoded base layer blocks (or have the same motion parameters), then the M×N enhancement layer (sub) blocks are also merged with the enhancement layer blocks corresponding to the base layer blocks used for merging within the base layer (i.e., the motion parameters are copied from the corresponding enhancement layer blocks) (see Example M).

[0364] [Encoding / Context Modeling of Transform Coefficient Levels] Main Mode: Encode the transform coefficients using different scanning patterns. For the enhancement layer, model the context based on the encoding mode and / or the base layer data, and perform different initializations for the context mode.

[0365] Sub-Example: ·Introduce one or more additional scanning patterns, e.g., horizontal and vertical scanning patterns. Redefine sub-blocks for the additional scanning patterns. Instead of 4×4 sub-blocks, for example, 16×1 or 1×16 sub-blocks may be used. Or 8×2 or 2×8 sub-blocks may be used. The additional scanning pattern may be introduced only for blocks larger than or equal to a specific size, e.g., 8×8 or 16×16 (see Example V). ·(If the coded block flag is equal to 1, hypothetically,) the selected scanning pattern is signaled in the bitstream (see Example N). A fixed context is used to signal the corresponding syntax element. Or the context derivation for the corresponding syntax element can depend on any of the following. ○The gradient of the collocated reconstructed base layer signal or the reconstructed base layer residual. Or an edge detected in the base layer signal. ○The distribution of transform coefficients in the collocated base layer blocks. ·The selected scan can be obtained directly from the base layer signal (without transmitting any additional data) based on the characteristics of the collocated base layer signal (see Example N). ○The gradient of the collocated reconstructed base layer signal or the reconstructed base layer residual. Or an edge detected in the base layer signal. ○The distribution of transform coefficients in the collocated base layer blocks. ·Different scans can be realized in a way that the transform coefficients are reordered after quantization on the encoder side and conventional coding is used. On the decoder side, the transform coefficients are decoded as usual and reordered before scaling and inverse transform (or before scaling and after inverse transform). ·For coding important flags (subgroup flags and / or important flags for a single transform coefficient), the following changes may be used in the enhancement layer. ○ The separation context model is used for all or a subset of the coding modes that use base layer information. It is also possible to use different context models for different modes with base layer information. ○ Context modeling can depend on the data of collocated base layer blocks (e.g., the number of significant transform coefficients for a specific frequency position) (see Example O). ○ A generalized template that evaluates both the number of already coded significant transform coefficient levels within the spatial neighborhood of the coefficient to be coded and the number of significant transform coefficients within the collocated base layer signals at similar frequency positions can be used (see Example O). · For coding the last significant scan position, the following changes can be used within the enhancement layer. ○ The separation context model is used for all or a subset of the coding modes that use base layer information. It is also possible to use different context models for different modes with base layer information (see Example P). ○ Context modeling can depend on the data within the collocated base layer blocks (e.g., the transform coefficient distribution in the base layer, the gradient information of the base layer, the last scan position within the collocated base layer blocks). ○ The last scan position can be coded as the difference relative to the last base layer scan position (see Example S). · Use of different context initialization tables for the base layer and the enhancement layer.

[0366] [Coding of the backward adaptation enhancement layer using base layer data] Main example: Use of base layer data to derive enhancement layer coding parameters.

[0367] Sub - aspect: ·Deriving merge candidates based on a (potentially upsampled) base layer reconstruction. Within the enhancement layer, only the use of merge is signaled. However, in practice, the candidates used to merge the current block are derived based on the reconstructed base layer signal. Thus, for all merge candidates, the magnitude of the error between the (potentially upsampled) base layer signal for the current enhancement layer block and the corresponding prediction signal (derived using motion parameters for the merge candidate) is evaluated for all merge candidates (or a subset thereof). And the merge candidate associated with the smallest error magnitude is selected. Also, the error magnitude can be calculated within the base layer using the reconstructed base layer signal and the reference picture of the base layer (see Example Q). ·Obtaining merge candidates based on a (potentially upsampled) base layer reconstruction. The motion vector difference is not encoded but is inferred based on the reconstructed base layer. For the current block, determine a motion vector predictor and evaluate a defined set of searches located around the motion vector predictor. For each search position, determine the magnitude of the error between the (potentially upsampled) base layer signal for the current enhancement layer block and the replaced reference frame (the replacement is given by the search position). Select the search position / motion vector that results in the smallest error magnitude. The search can be divided into several stages. For example, a full pel search is first performed. Subsequently, a half pel search is performed around the full pel vectors. Subsequently, a quarter pel search is performed around the best full / half pel vectors. Also, the search can be performed within the base layer using the reconstructed base layer signal and the reference picture of the base layer. The discovered motion vector is then scaled according to the resolution change between the base layer and the enhancement layer (see Example Q). ·Deriving an intra prediction mode based on a base layer reconstruction (potentially upsampled). The intra prediction mode is not encoded, but is inferred based on the reconstructed base layer. For each possible intra prediction mode (or a subset thereof), determine the magnitude of the error between the (potentially upsampled) base layer signal for the current enhancement layer block and the intra prediction signal (using the tested prediction nodes). Select the prediction mode that results in the smallest error magnitude. Also, the error magnitude calculation can be done within the base layer using the reconstructed base layer signal and the intra prediction signal within the base layer. Further, an intra block can be implicitly decomposed into 4×4 blocks (or another block size). And for each 4×4 block, a separate intra prediction mode is determined (see Example Q). ·The intra prediction signal can be determined by column alignment or row alignment of boundary samples having the reconstructed base layer signal. To derive the shift between adjacent samples and the current line / column, the error magnitude is calculated between the shifted line / column of adjacent samples and the reconstructed base layer signal. And the shift that results in the smallest error magnitude is selected. As adjacent samples, samples of the (upsampled) base layer or samples of the enhancement layer can be used. Also, the error magnitude can be calculated directly within the base layer (see Example W). ·Using the adaptation method described below for deriving other coding parameters such as block partitioning.

[0368] A more concise overview of the above embodiment is presented below. In particular, the above embodiment is described.

[0369] A1) A scalable video decoder reconstructs (80) base layer signals (200a, 200b, 200c) from an encoded data stream (6), reconstructs (60) an enhancement layer signal (360), Reconfiguration (60) is as follows: The reconfigured base layer signals (200a, 200b, 200c) are subjected to resolution or quality improvement (220) to obtain an inter-layer prediction signal (380). A difference signal between the already reconfigured part (400a or 400b) of the enhancement layer signal and the inter-layer prediction signal (380) is calculated (260). In a first part (440, exemplified in FIG. 46) juxtaposed with the part of the enhancement layer signal (360) to be currently reconfigured, a second part (460) of the difference signal, which is spatially adjacent to the first part and belongs to the already reconfigured part of the enhancement layer signal (360), is spatially predicted (260) to obtain a spatial intra prediction signal. The inter-layer prediction signal (380) and the spatial intra prediction signal are combined (260) to obtain an enhancement layer prediction signal (420). It is configured to include predictively reconstructing (320, 580, 340, 300, 280) the enhancement layer signal (360) using the enhancement layer prediction signal (420). According to Example A1, the base layer signal can be respectively reconstructed from the encoded data stream 6 or sub-stream 6a by the base layer decoding stage 80 in a prediction method based on the aforementioned block having inverse transform decoding, as long as the base layer residual signal 640 / 480 is relevant. However, other alternative reconstructions are also possible. As far as the reconstruction of the enhancement layer signal 360 by the enhancement layer decoding stage 60 is concerned, the resolution or quality improvement that the reconfigured base layer signals 200a, 200b or 200c undergo means, for example, upsampling in the case of resolution improvement, or copying in the case of quality improvement, or tone mapping from n bits to m bits (m > n) in the case of bit depth improvement. The calculation of the difference signal is done on a pixel-by-pixel basis. That is, the pixels with the enhancement layer signal juxtaposed on one side and the prediction signal 380 on the other side are subtracted from each other. And this is done for each pixel position. Spatial prediction of differential signals can be done in some way by transmitting intra prediction parameters such as the intra prediction direction within the encoded data stream 6 or within the sub-stream 6b, and then copying / interpolating the already reconstructed pixels adjacent to the portion of the enhancement layer signal 360 to be currently reconstructed along this intra prediction direction within the current portion of the enhancement layer signal. The combination can mean addition, weighted sum or even more sophisticated combinations, such as combinations that differently weight the contributions in the frequency domain. Predictive reconstruction of the enhancement layer signal 360 using the enhancement layer prediction signal 420 means, as shown in the figure, e...

Claims

Claim 1 Reconstructing (80) a base layer signal from a coded data stream (6) to obtain a reconstructed base layer signal; Predicting (30, 32) a portion of an enhancement layer signal (360) to be currently reconstructed, spatially or temporally, from a previously reconstructed portion of the enhancement layer signal to obtain an enhancement layer intra prediction signal (34); Forming (41) a weighted average of the inter-layer prediction signal obtained from the reconstructed base layer signal (200) and the enhancement layer intra prediction signal such that the weighting between the inter-layer prediction signal and the enhancement layer intra prediction signal varies different spatial frequency components in a portion (28) to be currently reconstructed, to obtain an enhancement layer prediction signal (42); Predictively reconstructing (52) the enhancement layer signal using the enhancement layer prediction signal; being configured to reconstruct (60) the enhancement layer signal (400); A scalable video decoder, characterized by the above. Claim 2 The scalable video decoder according to claim 1, characterized in that, when reconstructing (60) the enhancement layer signal, the reconstructed base layer signal is configured to receive resolution or quality improvement to obtain the inter-layer prediction signal. Claim 3 When forming the weighted average, filtering the inter-layer prediction signal (39) with a first filter (62) and filtering the enhancement layer intra prediction signal (34) with a second filter (64) in the portion to be currently reconstructed to obtain filtered signals, and further configured to add the filtered signals obtained from the first filter and the second filter having different transfer functions. The scalable video decoder according to claim 1 or claim 2, characterized by the above. Claim 4 The scalable video decoder according to claim 3, characterized in that the first filter is a low-pass filter and the second filter is a high-pass filter. Claim 5 The scalable video decoder according to claim 3 or claim 4, characterized in that the first filter and the second filter form a quadrature mirror filter pair.

6. When forming the weighted average, in the part to be currently reconstructed, the inter-layer prediction signal and the enhancement layer internal prediction signal are converted (72, 74) to obtain conversion coefficients (76, 78), and the obtained conversion coefficients are further configured to be superimposed (90) using weighting coefficients (82, 84) with different ratios between different spatial frequency components to obtain the superimposed conversion coefficients, characterized in that it is a scalable video decoder according to claim 1 or claim 2.

7. Further configured to inverse-transform (84) the superimposed conversion coefficients to obtain the enhancement layer prediction signal, characterized in that it is a scalable video decoder according to claim 6.

8. Predictively reconstruct the enhancement layer signal using the enhancement layer prediction signal, extract the conversion coefficient level (59) for the enhancement layer signal from the encoded data stream (6), Perform a sum (52) of the conversion coefficient level and the superimposed conversion coefficients to obtain a converted version of the enhancement layer signal, And further configured such that the converted version of the enhancement layer signal undergoes inverse transformation (84) to obtain the enhancement layer signal, characterized in that it is a scalable video decoder according to claim 6.

9. Characterized in that the weighting coefficients for weighting the inter-layer prediction signal and the enhancement layer internal prediction signal for each spatial frequency component are configured to add up to an equal value (46) for all spatial frequency components, wherein it is a scalable video decoder according to any one of claims 5 to 8.

10. Characterized in that the weighting coefficient for weighting the inter-layer prediction signal corresponds to a low-pass transfer function, and the weighting coefficient for weighting the enhancement layer internal prediction signal corresponds to a high-pass transfer function, wherein it is a scalable video decoder according to any one of claims 5 to 9.

11. When reconstructing (60) the enhancement layer signal, calculate the difference signal (734) between the already reconstructed part of the enhancement layer signal and the inter-layer prediction signal, At a first portion (744) juxtaposed to the portion of the enhancement layer signal to be reconfigured now, spatially adjacent to the first portion, and belonging to the already reconfigured portion of the enhancement layer signal, the different signal is spatially predicted from a second portion (736) of the difference signal to obtain a spatially internal prediction signal (746). The scalable video decoder according to any one of claims 1 to 10, characterized in that it is configured to combine (732) the inter-layer prediction signal and the spatially internal prediction signal to obtain the enhancement layer prediction signal.

12. When reconstructing (60) the enhancement layer signal, a difference signal (734) between the already reconfigured portion of the enhancement layer signal and the inter-layer prediction signal is calculated. At a first portion (744) juxtaposed to the portion of the enhancement layer signal to be reconfigured now, the difference signal is temporally predicted from a second portion (736) of the difference signal belonging to a frame previously reconstructed in the enhancement layer signal to obtain a temporally predicted difference signal (746). The scalable video decoder according to any one of claims 1 to 10, further characterized in that it is configured to combine (732) the inter-layer prediction signal and the temporally predicted difference signal to obtain the enhancement layer prediction signal.

13. Reconstructing (80) a base layer signal from an encoded data stream (6) to obtain a reconstructed base layer signal. Predicting, spatially or temporally, a portion of the enhancement layer signal to be reconfigured now from the already reconfigured portion of the enhancement layer signal to obtain an enhancement layer internal prediction signal. At the portion to be reconfigured now, forming a weighted average of the inter-layer prediction signal obtained from the reconstructed base layer signal and the enhancement layer internal prediction signal (380) such that the weighting between the inter-layer prediction signal and the enhancement layer internal prediction signal varies different spatial frequency components to obtain an enhancement layer prediction signal. Including using the enhancement layer prediction signal to predictively reconstruct the enhancement layer signal. Including reconstructing (60) the enhancement layer signal. A scalable video decoding method, characterized in that it comprises the above steps.

14. Encoding (12) the base layer signal into the encoded data stream so as to enable reconstruction of the base layer signal reconstructed from the encoded data stream (6), Predicting spatially or temporally a portion of the enhancement layer signal to be currently encoded from the already encoded portion of the enhancement layer signal to obtain an enhancement layer intra prediction signal, Forming a weighted average of the inter-layer prediction signal obtained from the reconstructed base layer signal and the enhancement layer intra prediction signal such that the weighting between the inter-layer prediction signal and the enhancement layer intra prediction signal varies different spatial frequency components in the portion to be currently encoded to obtain an enhancement layer prediction signal, Predictively encoding the enhancement layer signal using the enhancement layer prediction signal, being configured to encode (14) the enhancement layer signal, A scalable video encoder, characterized in that.

15. Encoding (12) the base layer signal into the encoded data stream so as to enable reconstruction of the base layer signal reconstructed from the encoded data stream (6), Predicting spatially or temporally a portion of the enhancement layer signal to be currently encoded from the already encoded portion of the enhancement layer signal to obtain an enhancement layer intra prediction signal, Forming a weighted average of the inter-layer prediction signal obtained from the reconstructed base layer signal and the enhancement layer intra prediction signal such that the weighting between the inter-layer prediction signal and the enhancement layer intra prediction signal varies different spatial frequency components in the portion to be currently encoded to obtain an enhancement layer prediction signal, Predictively encoding the enhancement layer signal using the enhancement layer prediction signal, including encoding (14) the enhancement layer signal, A scalable video encoding method, characterized in that.

16. A computer program, characterized by having program code for executing the method according to claim 13 or 15 when operating on a computer.

Citation Information

Patent Citations

  • Scalable encoding method and device, scalable decoding method and device and these program and their recording media

    JP2007028034A

  • Video coding method and apparatus using multi-layer based weighted prediction

    US20060291562A1

  • Image encoding-decoding system and related techniques

    US20070223582A1

  • Video coding and decoding method using weighted prediction and apparatus for the same

    WO2006101354A1

  • Video coding with fine granularity spatial scalability

    WO2007082288A1