Scalable video coding using derivation of subblock subdivision for prediction from base layer
Patent Information
- Application Number
- KR1020247029537
- Authority / Receiving Office
- KR · KR
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2012-10-01
- Filing Date
- 2013-10-01
- Publication Date
- 2026-09-09
- Estimated Expiration
- 2033-10-01
Smart Images

Figure 112024096132247-PAT00019_ABST
Abstract
Description
Technology Field
[0001] The present invention relates to scalable video coding. Background Technology
[0002] In non-scalable coding, intra-coding refers to a coding technique that does not use reference data for already coded pictures, but uses only data from already coded parts of the current picture (e.g., restored samples, coding modes, or symbolic statistics). Intra-coded pictures (or intra-pictures) are used in broadcast bitstreams, for example, to enable decoders to tune to the bitstream at so-called random access points. Intra-pictures are also used to limit error propagation in error-vulnerable environments. Generally, since there are no available pictures that can be used as reference pictures, the first picture of the coded video sequence must be coded as an intra-picture. Often, intra-pictures are also used in scene cuts where temporal prediction generally cannot provide a suitable prediction signal.
[0004] Furthermore, intra-coding modes are also used for specific regions / blocks in so-called inter-pictures, which perform better than inter-coding modes in terms of rate-distortion efficiency. This is the case for flat regions as well as regions where temporal predictions are performed somewhat poorly (occlusions, partially dissolves or fading objects). The problem to be solved
[0005] Therefore, the objective of the present invention is to provide a concept of scalable video coding that achieves higher coding efficiency. means of solving the problem
[0006] This objective is achieved by the subject of the attached independent paragraph. Effects of the invention
[0007] One aspect of the present application is that scalable video coding can be rendered more efficiently by measuring the spatial variation of base layer coding parameters for a base layer signal to derive / select a subblock subdivision from a set of possible subblock subdivisions of an enhancement layer block so that it can be used for enhancement layer prediction. By this method, if any, less signaling overhead should be used to signal the subblock subdivision within the enhancement layer data stream. The subblock subdivision thus selected can be used to predictively code / decode the enhancement layer signal. Brief explanation of the drawing
[0008] Preferred embodiments are described in more detail below in connection with the following drawings. FIG. 1 shows a block diagram of a scalable video encoder in which the embodiments and aspects described herein can be implemented; FIG. 2 shows a block diagram of a scalable video decoder suitable for the scalable video encoder of FIG. 1, in which the embodiments and aspects described herein can be similarly implemented; FIG. 3 shows a block diagram of a more specific embodiment of a scalable video encoder in which the embodiments and aspects described herein can be implemented; FIG. 4 shows a block diagram of a scalable video decoder suitable for the scalable video encoder of FIG. 3, in which the embodiments and aspects described herein can be similarly implemented; Figure 5 shows a schematic diagram of a video and its base layer and enhancement layer versions while additionally indicating the coding / decoding sequence; FIG. 6 shows a schematic diagram of a portion of a layered video signal to illustrate possible prediction modes for an enhancement layer; FIG. 7 shows the formation of an enhancement layer prediction signal using spectrally varying weights between the inter-layer prediction signal and the enhancement layer internal prediction signal according to enhancement; FIG. 8 shows a schematic diagram of syntactic elements that may be included within an enhancement layer substream according to enhancement; FIG. 9 shows a schematic diagram illustrating a possible embodiment of the formation of FIG. 7 according to an embodiment in which the formation / combination is performed in a spatial domain; FIG. 10 shows a schematic diagram illustrating a possible embodiment of the formation of FIG. 7 according to an embodiment in which the formation / combination is performed in the spectral region; FIG. 11 shows a schematic diagram of a layered video signal distribution to illustrate the derivation of spatial intra-prediction parameters from a base layer to an enhancement layer signal according to an embodiment; FIG. 12 shows a schematic diagram illustrated using the derivation of FIG. 11 according to an embodiment; FIG. 13 shows a schematic diagram of a set of spatial intra-prediction parameter candidates into which the one derived from the base layer is inserted according to an embodiment; FIG. 14 shows a schematic diagram of the distribution of layered video signals to illustrate prediction parameter granularity derivation from a base layer according to an embodiment; FIGS. 15a and b schematically illustrate a method for selecting an appropriate subdivision for the current block using spatial variation of base layer motion parameters within the base layer according to two different examples; FIG. 15c schematically illustrates a first possibility (first possibility) of selecting the coarsest among possible sub-block subdivisions for the current enhancement layer block; FIG. 15d schematically illustrates a second possibility (second possibility) of how to select the coarsest among possible sub-block subdivisions for the current enhancement layer block; FIG. 16 schematically shows the distribution of layered video signals to illustrate the use of sub-block subdivision induction for a current enhancement layer block according to an embodiment; FIG. 17 schematically illustrates the distribution of layered video signals to illustrate the use of base layer hints to efficiently code enhancement layer motion parameter data according to the series example; FIG. 18 schematically illustrates a first possibility for improving the efficiency of enhancement layer motion parameter signaling; FIG. 19a schematically illustrates a second possibility of how base layer hints are used to render enhancement layer motion parameter signaling more efficiently; FIG. 19b illustrates the first possibility of sequentially passing the base layer to a list of enhancement layer motion parameter candidates; FIG. 19c illustrates a second possibility of passing the base layer order to a list of enhancement layer motion parameter candidates; FIG. 20 schematically illustrates another possibility of using base layer hints to render enhancement layer motion parameter signaling more efficiently; FIG. 21 schematically illustrates the distribution of layered video signals to illustrate an embodiment in which the sub-block subdivision of the transform factor block is appropriately adjusted to hints derived from the base layer according to the embodiment; FIG. 22 illustrates different possibilities for how to derive appropriate sub-block subdivision of the transformation factor block from the base layer; FIG. 23 shows a block diagram of a much more detailed embodiment of a scalable video decoder in which the aspects and embodiments described herein can be implemented; FIG. 24 shows a block diagram of a scalable video encoder corresponding to the embodiment of FIG. 23, in which the aspects and embodiments described herein can be explained; FIG. 25 shows the generation of an inter-layer intra prediction signal by the sum of a spatial intra prediction using different signals (EH Diff) of adjacent blocks already being coded and a (upsampled / filtered) base layer reconstruction signal (BL Reco); FIG. 26 shows the generation of an inter-layer intra prediction signal by the sum of a spatial intra prediction using restored enhancement layer samples (EH Reco) of adjacent blocks already coded and a (upsampled / filtered) base layer residual signal (BL Resi); FIG. 27 illustrates the generation of an inter-layer intra prediction signal by a frequency-weighted sum of a (upsampled / filtered) base layer reconstruction signal (BL Reco) and a spatial intra prediction using restored enhancement layer samples (EH Reco) of adjacent blocks already coded; FIG. 28 illustrates the enhancement layer signals and base used in the explanation; FIG. 29 illustrates the motion compensation prediction of the enhancement layer; FIG. 30 illustrates a prediction using enhancement layer restoration and base layer residue; Figure 31 illustrates the prediction using the EL difference signal and BL restoration; Figure 32 illustrates the prediction using the 2-hypothesis of the EL difference signal and BL restoration; Figure 33 illustrates predictions using BL restoration and EL restoration; FIG. 34 shows an example of the decomposition of a picture into square blocks and a corresponding quad tree structure; FIG. 35 illustrates the decomposition of an acceptable square block into sub-blocks in a preferred embodiment; FIG. 36 illustrates the positions of motion vector predictors, where (a) represents the positions of spatial candidates and (b) represents the positions of temporal candidates; FIG. 37 shows a block merging algorithm (a) and a redundancy check (b) performed on spatial candidates; FIG. 38 shows a block merging algorithm (a) and a redundancy check (b) performed on spatial candidates; FIG. 39 illustrates the scan directions (diagonal, vertical, horizontal) for 4x4 transformation blocks; FIG. 40 illustrates the scan directions (diagonal, vertical, horizontal) for 8x8 transformation blocks. Shaded areas define effective sub-groups; Fig. 41 is a city of 16x16 transformations, where only diagonal scans are defined; Fig. 42 illustrates a vertical scan for a 16x16 transformation as proposed in JCTVC-G703; FIG. 43 illustrates the realization of vertical and horizontal scans for 16x16 transformation blocks. The coefficient subgroups are defined as a single column or a single column, respectively; Figure 45 illustrates backward-adaptive enhancement layer intra-prediction using restored base layer samples and adjacent restored enhancement layer samples. FIG. 46 schematically shows an enhancement layer picture / frame for illustrating different signal spatial interpolation according to an embodiment. Specific details for implementing the invention
[0009] In scalable coding, the concept of intra-coding (coding of intra-pictures and coding of intra-blocks in inter-pictures) can be extended to all pictures belonging to the same access unit or time instant. Thus, intra-coding modes for spatial or quality enhancement layers may utilize inter-layer prediction from lower-layer pictures in the same time instant to increase coding efficiency. This means that not only parts already coded within the current enhancement layer picture, but also lower-layer pictures already coded in the same time instant can be utilized. The latter concept is also referred to as inter-layer intra prediction.
[0011] In modern hybrid video coding standards (such as H.264 / AVC or HEVC), pictures in a video sequence are divided into blocks of samples. The block size may be fixed, or the coding approach may provide a hierarchical structure that allows blocks to be further subdivided into blocks with smaller block sizes. Reconstruction of a block is generally achieved by generating a prediction signal for the block and adding the transmitted residual signal. The residual signal is typically transmitted using transform coding, which means that the quantization indices of the transform coefficients (also referred to as transform coefficient levels) are transmitted using entropy coding techniques; from the decoder's perspective, these transmitted transform coefficient levels are scaled and inverse transformed to obtain the residual signal added to the prediction signal. The residual signal is generated by intra prediction (using only data already transmitted for the current time instance) or by inter prediction (using data already transmitted for different time instances).
[0013] When inter-prediction is used, the prediction block is derived by motion-compensated prediction using samples of already restored frames. This can be performed by unidirectional prediction (using a single set of motion parameters and a single reference picture), or the prediction signal can be generated by multi-hypothesis prediction. In the latter case, two or more prediction signals are superimposed, that is, for example, configured so that a weighted average forms the final prediction signal. The multiple prediction signals (superimposed) can be generated using different motion parameters (e.g., different reference pictures or motion vectors) for different estimates. For unidirectional prediction, it is also possible to add a constant offset (compensation value) and multiply samples of motion-compensated prediction signals with a constant factor to form the final prediction signal. Such scaling and offset modification may also be used for all or selected estimates in multi-hypothesis prediction.
[0015] In current state-of-the-art video coding technologies, intra-predicted signals for a block are obtained from prediction samples derived from the spatial neighborhood of the current block (recovered prior to the current block based on the block processing order). Most recently, various prediction methods are utilized to perform predictions in spatial domains. There is micro-particle directional prediction, where filtered or single-filtered samples of adjacent blocks are extended at specific angles to generate a prediction signal. Furthermore, there are plane-based and DC-based prediction modes that utilize adjacent block samples to generate flat prediction planes or DC prediction blocks.
[0017] In older video coding standards (e.g., H.263, MPEG-4), intra-prediction was performed in the transform domain. In these cases, the transmitted coefficients were inversely quantized. For the transform coefficients, the values were predicted using the corresponding reconstructed transform coefficients of adjacent blocks. The inversely quantized transform coefficients were added to the predicted transform coefficient values, and the reconstructed transform coefficients were used as input for the inverse transform. The output of the inverse transform formed the final reconstructed signal for the block.
[0019] In scalable video coding, base layer information can be utilized to support the prediction process for the enhancement layer. In the latest video coding standards for scalable coding, such as the SVC extension of H.264 / AVC, there is an additional mode to improve the coding efficiency of intra prediction in the enhancement layer. This mode is supported only when coexisting samples from lower layers are coded using the intra prediction mode. When this mode is selected for a macroblock in the quality enhancement layer, the prediction signal is composed of coexisting samples of the lower layer signal restored before the block separation filter operation (deblocking filter operation). If the inter-layer intra prediction mode is selected in the spatial enhancement layer, the prediction signal is generated by upsampling the coexisting restored base layer signal (after the block separation filter operation). FIR filters are used for upsampling. Generally, for the inter-layer intra prediction mode, additional residual signals are transmitted by transform coding. The transmission of the residual signal may be omitted if it is signaled correspondingly in the bitstream (it is assumed to be equal to 0). The final restored signal is obtained by adding the restored residual signal (obtained by scaling the transmitted transformation coefficient levels and applying an inverse spatial transformation) to the predicted signal.
[0021] However, it would be desirable to be able to achieve higher coding efficiency in scalable video coding.
[0023] Therefore, the objective of the present invention is to provide a concept of scalable video coding that achieves higher coding efficiency.
[0025] This objective is achieved by the subject of the attached independent paragraph.
[0027] One aspect of the present application is that a better predictor for predictively coding an enhancement layer signal in scalable video coding can be achieved by forming an enhancement layer prediction signal from an inter-layer prediction signal and an enhancement layer internal prediction signal weighted in different ways for different spatial frequency components, that is, by forming a weighted average of the enhancement layer internal prediction signal and the inter-layer prediction signal in the portion currently to be restored to obtain an enhancement layer prediction signal such that the weights contributing to the enhancement layer prediction signal of the enhancement layer internal prediction signal and the inter-layer prediction signal vary across different spatial frequency components. By this manner, it is feasible to construct an enhancement layer prediction signal from the enhancement layer internal prediction signal and the inter-layer prediction signal in a manner optimized for the spectral features of the individual contributing components, namely, on the one hand, the inter-layer prediction signal and on the other, the enhancement layer internal prediction signal. For example, due to the resolution or quality improvement obtained from the reconstructed base layer signal, the inter-layer prediction signal may be more accurate at lower frequencies (low frequencies) compared to high frequencies. As far as the enhancement layer internal prediction signal is concerned, its characteristics may be the opposite; that is, its accuracy may be increased for high frequencies compared to low frequencies. In this example, the contribution of the inter-layer prediction signal to the enhancement layer prediction signal must, by respective weighting, exceed the contribution of the enhancement layer internal prediction signal to the enhancement layer prediction signal at low frequencies, and, as far as high frequencies are concerned, keep the contribution of the enhancement layer internal prediction signal to the enhancement layer prediction signal below a certain level.Through this method, a more accurate enhancement layer prediction signal is achieved, thereby deriving a higher compression ratio and increasing coding efficiency.
[0029] By various embodiments, different possibilities are described to construct the described concept into any scalable video coding-based concept. For example, the formation of a weighted average may be formed in a spatial domain or a transformation domain, respectively. Performing a spectrally weighted average requires transformations performed on individual contributions, namely the inter-layer prediction signal and the enhancement layer internal prediction signal, but avoids spectrally filtering either the enhancement layer internal prediction signal or the inter-layer prediction signal, such as FIR or IIR filtering, in which the spatial domain is involved. However, performing the formation of a spectrally weighted average in the spatial domain avoids bypassing individual contributions to the weighted average through the transformation domain. The decision regarding which region is actually selected to perform the formation of the spectrumally weighted average may depend on whether the scalable video data stream contains residual signals in the form of transform factors for the portion currently to be constructed in the enhancement layer signal: if not, bypassing through the transform region may be discontinued, whereas in the case of present residual signals, bypassing through the transform region is much more advantageous because it directly enables the residual signals transmitted from the transform region to be added to the spectrumally weighted average in the transform region.
[0031] One aspect of the present application is that information available from coding / decoding the base layer, namely base-layer hints, can be used to render the motion-compensation prediction of the enhancement layer more efficiently by coding the enhancement layer motion parameters more efficiently. In particular, a set of motion parameter candidates collected from already restored adjacent blocks of the enhancement layer signal frame can be expanded by a set of one or more base layer motion parameters of a block of the base layer signal coexisting in the block of the enhancement layer signal frame, thereby improving the available quality of the set of motion parameter candidates based on which the motion compensation prediction of the block of the enhancement layer signal can be performed by using a motion parameter candidate selected for prediction and selecting one of the motion parameter candidates of the expanded set of motion parameter candidates. Additionally or alternatively, the list of motion parameter candidates of the enhancement layer signal can be sorted based on the base layer motion parameters involved in coding / decoding the base layer. By this method, the probability distribution for selecting enhancement layer motion parameters from an ordered motion parameter candidate list is densed, for example, so that clearly signaling index syntax elements can be coded using fewer bits, such as by using entropy coding. Furthermore, additionally or alternatively, the index used to code / decode the base layer can function as a basis for determining the index as a motion parameter candidate list for the enhancement layer.By this method, any signaling of the index to the enhancement layer can be completely avoided, or only the deviation of the prediction determined for the index can be transmitted within the enhancement layer substream, thus improving coding efficiency.
[0033] One aspect of the present application is that scalable video coding can be rendered more efficiently by deriving / selecting a subblock subdivision to be used for enhancement layer prediction from a set of possible subblock subdivisions of an enhancement layer block by measuring spatial variation of base layer coding parameters over the base layer signal. In this manner, less signalization overhead must be consumed in signaling this subblock subdivision within the enhancement layer data stream. Thus, the selected subblock subdivision can be used to predictively code / decode the enhancement layer signal.
[0035] One aspect of the present application is that the subblock-based coding of the transform coefficient blocks of the enhancement layer can be rendered more efficiently when the subblock subdivision of individual transform coefficient blocks is controlled based on the base layer residual signal or the base layer signal. In particular, by utilizing a particular base layer hint, the subblocks can be formed longer along a spatial frequency axis that crosses an observable edge extension from the base layer signal or the base layer residual signal. By this manner, it is feasible to adapt the shape of the subblocks to the estimated distribution of the energy of the transform coefficients of the enhancement layer transform coefficient blocks in such a way that, at reduced probability, some subblocks have a similar number of significant transform coefficients on one hand and insignificant transform coefficients on the other, whereas at increased probability, each subblock can be filled almost entirely with significant, i.e., transform coefficients that are not quantized to zero, or less significant transform coefficients, i.e., transform coefficients that are only quantized to zero. However, because subblocks without valid transformation coefficients can be efficiently signaled within the data stream as if by the use of only a single flag, and subblocks almost completely filled with valid transformation coefficients can be placed there, there is no need to waste signaling quantity to code less valid transformation coefficients, and coding efficiency for coding the transformation coefficient blocks of the enhancement layer is increased.
[0037] One aspect of the present application is that the coding efficiency of scalable video coding can be increased by using intra-predictors of coexisting blocks of base layer signals to replace spatially lost intra-predictors in the spatial vicinity of the current block of the enhancement layer. By this method, the coding efficiency for coding spatial intra-predictors is increased due to the improved prediction quality of the set of intra-predictors of the enhancement layer, or, more specifically, the increased possibility that appropriate predictors for intra-predictors for the intra-predictors of the enhancement layer blocks are available, thereby increasing the possibility that, on average, the signaling of the intra-predictors of individual enhancement layer blocks can be performed with fewer bits.
[0039] More advantageous practices are explained in the dependent terms.
[0041] FIG. 1 illustrates, in a general manner, an embodiment of a scalable video encoder that may be equipped in the embodiments described below. The scalable video encoder of FIG. 1 receives a video (4) to be encoded, generally indicated by the reference numeral (2). The scalable video encoder (2) is configured to encode the video (4) into a data stream (6) in a scalable manner. The data stream (6) comprises a first portion (6a) having a video (4) encoded in a first amount of information content, and an additional portion (6b) having a video (4) encoded in a larger amount of information content than one of the portions (6a). The amounts of information content of the portions (6a and 6b) may differ, for example, in quality and fidelity, i.e., in pixel-unit deviation from the original video (4), and / or in spatial resolution. However, other forms of differences in the amounts of information content may apply, for example, such as color fidelity or similar. While part (6b) may be a canceled enhancement layer data stream or enhancement layer substream, part (6a) may be a canceled base layer data stream or base layer substream.
[0043] A scalable video encoder (2) is configured to use redundancies between versions (8a and 8b) of video (4) recoverable from a base layer substream (6a), each without an enhancement layer substream (6b) on one hand and without both substreams (6a and 6b) on the other. To do so, the scalable video encoder (2) may use inter-layer prediction.
[0045] As shown in FIG. 1, the scalable video encoder (2) can alternately receive two versions (4a and 4b) of the video (4), and the two versions differ in the amount of information content, as do the base layer and enhancement layer substreams (6a and 6b). Thus, for example, the scalable video encoder (2) will be configured to generate substreams (6a and 6b) such that the base layer substream (6a) has versions (4a) encoded therein, while the enhancement layer data stream (6b) is encoded into version (4b) using inter-layer prediction based on the base layer substream (6b). Both encodings of the substreams (6a and 6b) may be lost.
[0047] Even if the scalable video encoder (2) receives only the original version of the video (4), it is configured to internally derive the same from two versions (4a and 4b), for example, by obtaining a base layer version (4a) by spatial down-scaling from a high bit depth to a low bit depth and / or tone mapping.
[0049] FIG. 2 shows a scalable video decoder suitable for the scalable video encoder of FIG. 1 in the same manner suitable for working with any of the embodiments described below. The scalable video decoder of FIG. 2 is generally denoted by the reference numeral (10) and is configured to decode a coded data stream (6) to restore from an enhancement layer version (8b) of the video when both parts (6a and 6b) of the data stream (6) arrive at the scalable video decoder (10) in an intact manner, or, for example, from a base layer version (8a) when part (6b) is unavailable due to transmission loss or similar things. It is configured such that the scalable video decoder (10) can restore version (8a) from only the base layer substream (6a), and can restore version (8b) from both parts (6a and 6b) using inter-layer prediction.
[0051] Before describing the details of the embodiments of the present application in the details below, that is, before showing how the embodiments of FIGS. 1 and 2 are more specifically realized, more detailed implementations of the scalable video encoder and decoder of FIGS. 1 and 2 are described with respect to FIGS. 3 and 4. FIG. 3 shows a scalable video (2) comprising a multiplexer (16), an enhancement layer coder (14), and a base layer coder (12). The base layer coder (12) is configured to encode a base layer version (4a) of the inbound video, whereas the enhancement layer coder (14) is configured to encode an enhancement layer version (4b) of the video. Thus, the multiplexer (16) receives an enhancement layer substream (6b) from the enhancement layer coder (14) and a base layer substream (6a) from the base layer coder (12) and multiplexes both into a coded data stream (6) at its output.
[0053] As shown in FIG. 3, both coders (12 and 14) may be prediction coders that utilize spatial and / or temporal prediction to encode individual inbound versions (4a and 4b) into individual substreams (6a and 6b), for example. In particular, the coders (12 and 14) may each be hybrid video block coders. This means that one of the coders (12 and 14) may be configured to encode individual inbound versions of video on a block-by-block basis, for example, frames or pictures of individual video versions (4a and 4b), during a selection between different prediction modes for each block of blocks, each of which is subdivided. The different prediction modes of the base layer coder (12) may include spatial and / or temporal prediction modes, while the enhancement layer coder (14) may additionally support inter-layer prediction modes. Subdivision into blocks may differ between the base layer and the enhancement layer. Prediction modes, prediction parameters for the prediction modes selected for various blocks, prediction residues, and, optionally, block subdivisions of individual video versions may be described by individual coders (12, 14) using individual syntax, including syntax elements that can be coded into individual substreams (6a, 6b) using entropy coding. Inter-layer prediction may be used in one or more occasions to predict samples of the enhancement layer video, prediction modes, prediction parameters, and / or block subdivisions, for example, for some of the examples just mentioned. Thus, both the base layer coder (12) and the enhancement layer coder (14) may each include a prediction coder (18a, 18b) followed by an entropy coder (19a, 19b).On the other hand, the prediction coders (18a, b) each form a stream of syntactic elements using prediction coding from the inbound versions (4a and 4b), and the entropy coder entropy encodes the syntactic elements output by the individual prediction coders. As just mentioned, the inter-layer prediction of the encoder (2) may be related to different occurrences in the encoding procedure of the enhancement layer, and thus the prediction coder (18b) is shown to be connected to one or more of the prediction coders (18a), their outputs, and the entropy coder (19a). Similarly, the entropy coder (19b) may optionally take advantage of the inter-layer prediction by predicting the contexts used for entropy coding from the base layer, for example, and thus the entropy coder (19b) is shown to be optionally connected to any of the elements of the base layer coder (12).
[0055] In the same manner as Fig. 2 relating to Fig. 1, Fig. 4 shows a possible embodiment of a scalable video decoder (10) corresponding to the scalable video encoder of Fig. 3. Accordingly, the scalable video decoder (10) of Fig. 4 includes a demultiplexer (40) that receives a data stream (6) to obtain substreams (6a and 6b), a base layer decoder (80) configured to decode the base layer substream (6a), and an enhancement layer decoder (60) configured to decode the enhancement layer substream (6b). As shown, the decoder (60) is connected to the base layer decoder (80) to receive information from it in order to take advantage of inter-layer prediction. In this manner, the base layer decoder (80) can recover a base layer version (8a) from a base layer substream (6a), and the enhancement layer decoder (60) is configured to recover an enhancement layer version (8b) of the video using an enhancement layer substream (6b). Similar to the scalable video encoder of FIG. 3, each of the base layer and enhancement layer decoders (60 and 80) may internally include an entropy decoder (100, 320) followed by a prediction decoder (102, 322), respectively.
[0057] To simplify understanding of the following embodiments, FIG. 5 illustrates different versions of video (4), namely, base layer versions (4a and 8a) that differ from each other only by coding loss, and enhancement layer versions (4b and 8b) that differ from each other only by coding loss. As shown, the base layer and enhancement layer signals can each be composed of a sequence of pictures (22a and 22b). They are illustrated in FIG. 5 as being registered to each other along the temporal axis (24), namely, the picture (22a) of the base layer version next to the picture (22b) of the enhancement layer signal that corresponds temporally. As described above, the picture (22b) may have a higher spatial resolution and / or represent video (4) with high fidelity, such as in the high bit depth of the sample values of the pictures. By means of continuous and dotted lines, the coding / decoding order is shown to be defined among the pictures (22a, 22b). According to the example illustrated in FIG. 5, the coding / decoding order traverses pictures (22a) and (22b) in such a way that the base layer picture (22a) of a specific time stamp / instance is traversed ahead of the enhancement layer picture (22b) of the same time stamp of the enhancement layer signal. With respect to the temporal axis (24), the pictures (22a, 2b) may be traversed by the coding / decoding order (26) in the representation time order, but an order deviating from the representation time order of the pictures (22a, 22b) may also be implemented. Neither the encoder nor the decoder (10, 2) needs to encode / decode sequentially along the coding / decoding order (26). Rather, parallel coding / decoding can be used.The coding / decoding order (26) can define availability between adjacent enhancement layer signals and base parts spatially, temporally, and / or in an inter-layer sense, such that the parts available for the current enhancement layer part are defined through the coding / decoding order at the time of coding / decoding the current part of the enhancement layer. Thus, only adjacent parts available according to this coding / decoding order (26) are used for prediction by the encoder so that the decoder has access to the same source of information to re-perform the prediction.
[0059] In connection with the following drawings, as described above with respect to FIGS. 1 through 4, how a scalable video encoder or decoder can be implemented to form an embodiment of the present application according to one aspect of the application is described. Possible embodiments of the aspect now described are discussed below using the designation "Aspect C".
[0061] In particular, FIG. 6 illustrates a picture (22b) of an enhancement layer signal, indicated here using the reference numeral (360), and a picture (22a) of a base layer signal, indicated here using the reference numeral (200). Temporarily corresponding pictures of different layers are shown in a manner registered with respect to the temporal axis (24). Using hatching, parts within the base and enhancement layer signals (200 and 36) are distinguished from parts that have already been coded / decoded according to the coding / decoding order and parts that have not yet been coded or decoded according to the coding / decoding order shown in FIG. 5. FIG. 6 also shows a part (28) of the enhancement layer signal (360) that is currently being coded / decoded.
[0063] According to the embodiments described herein, the prediction of part (28) utilizes both inter-layer prediction from the base layer and intra-layer prediction within the enhancement layer itself to predict part (28). However, the predictions are combined in such a spectral change manner that these predictions contribute to the final predictor of part (28), particularly in a spectral change manner where both contributions vary spectrally.
[0065] In particular, part (28) is spatially or temporally predicted from an already restored part of the enhancement layer signal (400), i.e., any part illustrated as a hatching within the enhancement layer signal (400) in FIG. 6. While temporal prediction is illustrated using arrow (32), spatial prediction is illustrated using arrow (30). Temporal prediction may include, for example, motion compensation prediction as information on the motion vector is transmitted within the enhancement layer substream for the current part (28), and the motion vector represents the displacement of a part of the reference picture of the enhancement layer signal (400) to be replicated to obtain the temporal prediction of the current part (28). Spatial prediction (30) may include extrapolating already coded / decoded parts of the spatially adjacent current part (28) and the spatially adjacent picture (22b) to the current part (28). For this reason, intra-prediction information, such as extrapolation (or angular) directions, can be signaled within the enhancement layer substream for the current part (28). A combination of spatial and temporal predictions (30 and 32) may also be used. In any case, the enhancement layer internal prediction signal (34) is obtained by doing so as shown in FIG. 7.
[0067] To obtain another prediction of the current portion (28), an inter-layer prediction is used. For this reason, the base layer signal (200) is subject to resolution or quality improvement of the portion (36) that spatially and temporally corresponds to the current portion of the enhancement layer signal (400), and this is to obtain a potentially resolution-increased inter-layer prediction signal for the current portion (28) with an improvement procedure illustrated by the arrow (38) in FIG. 6, which derives the inter-layer prediction signal (39) as shown in FIG. 7.
[0069] Accordingly, two prediction contributions (34 and 39) exist for the current part (28), and the weighted average of both contributions is formed to obtain the enhancement layer prediction signal (42) for the current part (28) in such a way that the weights contributing to the enhancement layer prediction signal (42) by the inter-layer signal and the enhancement layer internal prediction signal vary differently for the spatial frequency components as schematically shown in FIG. 7, where, for example, the graph shows the case where the weights contributing to the final prediction signal by the prediction signals (34 and 38) for all spatial frequency components are added up to the same value (46) for all spectral components, but with a spectrum that changes the ratio between the weights applied to the prediction signal (39) and the weights applied to the prediction signal (34).
[0071] On the other hand, the prediction signal (42) can be directly utilized by the enhancement layer signal (400) in the current part (28), and alternatively, the residual signal may exist within the enhancement layer substream (6b) for the current part (28) derived by combining (50) with the prediction signal (42), as shown in FIG. 7, for example, as a restored version (54) of the current part (28). As an intermediate note, both scalable video encoders and decoders may be hybrid video decoders / encoders that use transform coding and prediction coding to encode / decode the prediction residual.
[0073] To summarize the description of FIGS. 6 and 7, the enhancement layer substream (6b) may include intra-prediction parameters (56) for controlling spatial and / or temporal predictions (30, 32) for the current portion (28), and optionally, residual information (59) for signaling residual signals (48) and weighting parameters (58) for controlling the formation of a spectral weighted average (41). While a scalable video encoder determines all of these parameters (56, 58, 59) and thus inserts the same into the enhancement layer substream (6b), a scalable video decoder uses the same to restore the current portion (28) as described above. All of these elements (56, 58, and 59) can be subject to some quantization, and thus a scalable video encoder can determine these quantized parameters / elements, i.e., using a rate / distortion cost function. Interestingly, the encoder (2) uses the determined parameters / elements (56, 58, and 59) to function as a basis for some prediction, for example, for a portion of the enhancement layer signal (400) that follows in coding / decoding order, and to obtain a restored version (54) for the current portion.
[0075] Different possibilities exist for the weighting parameters (58) and whether they control the formation of the spectrally weighted average in (41). For example, the weighting parameters (58) may signal only one of two states for the current part (28), namely, one state activating the formation of the spectrally weighted average described so far, and the other state deactivating the contribution of the inter-layer prediction signal (38), so that the final enhancement layer prediction signal (42) is, in such a case, made only by the enhancement layer internal prediction signal (34). Alternatively, the weighting parameter (58) for the current part (28) may switch between the inter-layer prediction signal (39) that activates the formation of the spectrally weighted average on one hand and forms the enhancement layer prediction signal (42) on the other. The weighting parameter (58) may also be designed to signal one of the three states / alternatives just mentioned. Alternatively, or additionally, weighting parameters (58) may control the spectrally weighted average formation (41) for the current part (28) with respect to the spectral variation of the ratio between the weights of the prediction signals (34 and 39) contributing to the final prediction signal (42). Later, it will be explained that, before adding the same, such as using a high pass and / or low pass filter, the spectrally weighted average formation (41) may include filtering both or one of the prediction signals (34 and 39), and in such case, the weighting parameters (58) may signal the filter characteristics for the filter or filters to be used for the prediction of the current part (28).Alternatively, it is explained that in step (41), spectral weighting can be achieved by individual spectral component weighting in the transformation area, and thus in this case, weighting parameters (58) can signal / set these individual spectral component weighting values.
[0077] Additionally or alternatively, the weighting parameter for the current part (28) can signal whether spectral weighting is performed in the transformation region or the spatial region in step (step) 41.
[0079] FIG. 9 illustrates an embodiment for performing a spectrally weighted average formation in a spatial domain. It is illustrated that the predicted signals (39 and 34) are obtained in the form of individual pixel batches that correspond to the pixel raster of the current portion (28). To perform the spectrally weighted average formation, it is shown that both pixel batches of both predicted signals (34 and 39) are subject to filtering. FIG. 9 illustrates filtering, for example, by showing filter kernels (62 and 64) crossing the pixel batches of the predicted signals (34 and 39) to perform FIR filtering. However, IIR filtering is also feasible. Furthermore, only one of the predicted signals (34 and 39) may be subject to filtering. The transfer functions of both filters (62 and 64) are different from each other so that adding up (66) the filtering results of pixel batches of the prediction signals (39 and 34) yields the result of forming a spectrally weighted average, i.e., the enhancement layer prediction signal (42). In other words, the addition (66) will simply add the samples coexisting within the prediction signals (39 and 34) as if they were filtered using the filters (62 and 64), respectively. (62) through (66) will yield the formation of a spectrally weighted average (41). FIG. 9 illustrates a case where residual information (59) exists in the form of transformation coefficients, so that the residual signal (48) is signaled in the transform domain, and the inverse transform (68) can be used to derive in the spatial domain in the form of a pixel arrangement (70), so that the combination (52) deriving the restored version (55) can be realized by simple pixel-wise addition of the enhancement layer prediction signal (42) and the residual signal arrangement (70).Again, it is recalled that the above prediction is performed by scalable video encoders and decoders using predictions for restoration in decoders and encoders.
[0081] FIG. 10 illustrates, by way of example, how to perform spectral weighted average formation in the transformation region. Here, pixel arrangements of the prediction signals (39 and 34) are each subject to transformation (72 and 74), thereby deriving spectral decompositions (76 and 78), respectively. Each spectral decomposition (76 and 78) consists of a arrangement of transformation coefficients having one transformation coefficient per spectral component. Each transformation coefficient block (76 and 78) is multiplied by a corresponding block of weights, namely blocks (82 and 84). Thus, in each spectral component, the transformation coefficients of blocks (76 and 78) are individually weighted. In each spectral component, the weight values of blocks (82 and 84) may be added up to a value common to all spectral components, but this is not mandatory. In fact, the multiplication (86) between blocks (76 and 82) and the multiplication (88) between block (78) and block (84), each represent spectral filtering in the transform region, and the transform coefficient / spectral component-wise adding (90) completes the formation of a spectrally weighted average (41), which is to derive a transform region version of the enhancement layer prediction signal (42) in the form of a transform coefficient block. As illustrated in FIG. 10, where the residual signal (59) signals the residual signal in the form of a transform coefficient block, the same can be combined with the transform coefficient block representing the enhancement layer prediction signal (42) to derive a restored version of the current part (28) in the transform region, whether the same is simply added by transform coefficients or otherwise (52). Accordingly, the inverse transformation (84) applied to the additional result of the combination (52) yields a pixel arrangement that restores the current part (28), i.e., the restored version (54).
[0083] As described above, parameters existing within the enhancement layer substream (6b) for the current part (28), such as residual information (59), or weighting parameters (58), can signal whether average formation (41) is performed within the transformation region shown in FIG. 10 or the spatial region according to FIG. 9. For example, if residual information (59) indicates the absence of any transformation factor block for the current part (28), the spatial region may be used, or the weighting parameters (58) may switch between both regions regardless of whether residual information (59) includes transformation factors.
[0085] Later, to obtain a layer-internal enhancement layer prediction signal, a difference signal can be calculated and managed between the inter-layer prediction signal and the already restored portion of the enhancement layer signal. Spatially predicting the difference signal from the first portion (first portion) associated with the portion of the enhancement layer signal currently to be restored, from the second portion (second portion) of the difference signal belonging to the already restored portion of the enhancement layer signal and spatially adjacent to the first portion, can be used to spatially predict the difference signal. Alternatively, temporally predicting the difference signal from the first portion associated with the portion of the enhancement layer signal currently to be restored, from the second portion of the difference signal belonging to previously restored frames of the enhancement layer signal, can be used to obtain a temporally predicted difference signal. The combination of the inter-layer prediction signal and the predicted difference signal can be used to obtain a layer-internal enhancement layer prediction signal combined with the inter-layer prediction signal.
[0087] In connection with the following drawings, it is explained how a scalable video encoder or decoder, as described above in connection with FIGS. 1 to 4, can be implemented to form an embodiment of the present application according to another aspect of the application.
[0089] To illustrate this perspective, FIG. 11 is provided by reference. FIG. 11 illustrates the possibility of performing a spatial prediction (30) of the current part (28). The following description of FIG. 11 may be combined with the description of FIG. 6 through 10. In particular, the perspective described herein will be further described later with respect to illustrated embodiments referred to as aspects X and Y.
[0091] The situation shown in FIG. 11 corresponds to one shown in FIG. 6. This is that base layer and enhancement layer signals (200 and 400) are shown together with already coded / decoded portions illustrated using hatching. The portion currently to be coded / decoded within the enhancement layer signal (400) has adjacent blocks (92 and 94), which are exemplarily depicted as block (94) to the left of and block (92) above the current portion (28), with both blocks (92 and 94) having the same size as the current block (28). However, matching sizes are not mandatory. Rather, the portions of the blocks into which the picture (22b) of the enhancement layer signal (400) is subdivided may have different sizes. They are not even limited to quadratic shapes. They may be square or other shapes. The current block (28) has additional adjacent blocks that are not specifically depicted in FIG. 11, but are not yet decoded / coded, that is, they follow the coding / decoding order and are not available for prediction. Beyond this—exemplarily diagonally across the top-left corner of the current block (28)—thereby may be blocks other than the blocks (92 and 94) that are already coded / decoded in the coding / decoding order, such as block (96), which is adjacent to the current block (28), but blocks (92 and 94) are predetermined adjacent blocks that serve to predict the intra prediction parameters for the current block (28) that are the subject of the intra prediction (30) in the example considered here. The number of such predetermined adjacent blocks is not limited to two. It may be just one or more.
[0093] Scalable video encoders and scalable video decoders can determine a set of predetermined adjacent blocks, where 92 and 94, from a set of already coded adjacent blocks, where 92 to 96, depending on a predetermined sample location (98) within the current part (28), for example, its top-left sample. For example, such already coded adjacent blocks of only the current part (28) can form a set of "predetermined adjacent blocks" that includes sample locations immediately adjacent to the predetermined sample location (98). In any case, the adjacent already coded / decoded blocks include samples (102) adjacent to the current block (28) based on sample values in which regions of the current block (28) are spatially predicted. For this reason, spatial prediction parameters such as (56) are signaled in the enhancement layer substream (6b). For example, the spatial prediction parameter for the current block (28) indicates the spatial direction according to the replication of the sample value of the sample (102) into the area of the current block (28).
[0095] In any case, when spatially predicting the current block (28), the scalable video decoder / encoder has a base layer (200) that has already been restored (and encoded in the case of an encoder) using a base layer substream (6a), insofar as the relevant spatial corresponding region of the temporally corresponding picture (22a) is related, for example, by using a block-by-block selection between spatial and temporal prediction modes and by using block-by-block prediction, as described above.
[0097] In FIG. 11, a current block (28) is exemplarily described as being located around and in a locally corresponding area to some blocks (104) into which a time-aligned picture (22a) of a base layer signal (200) is subdivided. As with spatially predicted blocks within an enhancement-layer signal (400), spatial prediction parameters are signaled or included within a base layer substream for such blocks (104) within the base layer signal (200), where the selection of a spatial prediction mode is signaled.
[0099] In order to allow the recovery of an enhancement layer signal from a data stream coded for a block (28), where, for example, a spatial intra-layer prediction (30) is selected, the intra-prediction parameters are used and coded within the bitstream according to the following:
[0101] Intra-predictor parameters are often coded using the concept of most probable intra-predictor parameters, which are somewhat small subsets of all possible intra-predictor parameters. The set of most probable intra-predictor parameters may contain one, two, or three intra-predictor parameters, for example, while the set of all possible intra-predictor parameters may contain, for example, 35 intra-predictor parameters. If an intra-predictor parameter is included in the set of most probable intra-predictor parameters, it can be signaled within a bitstream with a small number of bits. If an intra-predictor parameter is not included in the set of most probable intra-predictor parameters, its signaling within the bitstream requires more bits. Thus, the amount of bits consumed for a syntax element to signal an intra-predictor parameter for the current intra-predictor block depends on the quality of the set of most probable, or perhaps favorable, intra-predictor parameters. Based on this concept, it is assumed that, on average, a lower number of bits is required to code the intra-predictor parameters, and that the most probable set of intra-predictor parameters can be appropriately derived.
[0103] Generally, the most probable set of intra-predictors is selected by including, for example, intra-predictors often additionally used in the form of default parameters and / or intra-predictors of directly adjacent blocks. For example, because the main gradient directions of adjacent blocks are similar, it is generally advantageous to include the intra-predictors of adjacent blocks in the set of most probable intra-predictors.
[0105] However, adjacent blocks are not coded in spatial intra-prediction mode, and such parameters are not available in terms of the decoder.
[0107] However, in scalable coding, it is possible to use the intra-predicted parameters of co-located base layer blocks, and thus, according to the perspective described below, this environment is utilized by using the intra-predicted parameters of co-located base layer blocks when adjacent blocks are not coded in spatial intra-predicted mode.
[0109] Thus, according to Fig. 11, a set of possible favorable intra-prediction parameters for the current enhancement layer block is created by examining the intra-prediction parameters of predetermined adjacent blocks, and, for example, by exceptionally reclassifying to blocks coexisting in the base layer in cases where none of the predetermined adjacent blocks have suitable intra-prediction parameters associated with them because, since individual predetermined adjacent blocks are not coded in the intra-prediction mode.
[0111] Above all, a predetermined block, such as a block (92 or 94) of the current block (28), is checked to see if the same is predicted using a spatial intra-prediction mode, that is, whether a spatial intra-prediction mode has been selected for an adjacent block. Depending on that, the intra-prediction parameters of the adjacent block are included in the set of probably advantageous intra-prediction parameters for the current block (28), or as a substitute, if any, in the intra-prediction parameters of the coexisting block (108) of the base layer.
[0113] For example, if individual predetermined adjacent blocks are not spatial intra-prediction blocks, instead of using default predictors or similar ones, the intra-prediction parameters of the block (108) of the base layer signal (200) are included in a set of probably advantageous inter-prediction parameters for the current block (28) that coexist with the current block (28).
[0115] For example, the coexisting block (108) is determined using the predetermined sample location (98) of the current block (28), that is, the block (108) covers a location (106) that locally corresponds to the predetermined sample location (98) within the temporally aligned picture (22a) of the base layer signal (200). Naturally, additional verification may be performed to determine whether the coexisting block (108) within the base layer signal (200) is actually a spatially intra-predicted block. In the case of FIG. 11, such a case is illustrated exemplarily. However, if the coexisting block is not coded in the intra-predicted mode, the set of probably advantageous intra-predicted parameters may be left without any contribution to the predetermined adjacent block, or default intra-predicted parameters may be used as a substitute, that is, the default intra-predicted parameters are inserted as probably advantageous intra-predicted parameters.
[0117] So, if a block (108) coexisting with the current block (28) is spatially predicted, its intra-prediction parameter signaled within the base layer substream (6a) can be used as a substitute for any of the predetermined adjacent blocks (92 or 94) of the current block (28), which has no intra-prediction parameter because it is the same as coded using another prediction mode, like the temporal prediction mode.
[0119] According to another embodiment, in certain cases, even if each predetermined adjacent block is in an intra-prediction mode, the intra-prediction parameter of the predetermined adjacent block is substituted by the intra-prediction parameter of the coexisting base layer block. For example, additional verification may be performed on any of the predetermined adjacent blocks in the intra-prediction mode regarding whether the intra-prediction palmiers satisfy a specific criterion. If the specific criterion is not satisfied by the intra-prediction parameter of the adjacent block, but the same criterion is satisfied by the intra-prediction parameter of the coexisting base layer block, substitution is performed despite the very adjacent block being intra-coded. For example, if the intra-prediction parameter of the adjacent block does not represent an angular intra-prediction mode (e.g., DC or planar intra-prediction mode), and the intra-prediction parameter of the coexisting base layer block represents an angular intra-prediction mode, the intra-prediction parameter of the adjacent block may be substituted by the intra-prediction parameter of the base layer block.
[0121] The inter-predictor parameter for the current block (28) is determined based on the probabilistically advantageous intra-predictor parameters and the syntactic element present in the coded data stream, such as the enhancement layer substream (6b) for the current block (28). This means that in the case of the inter-predictor parameter for the current block (28) being a member of the set of probabilistically advantageous intra-predictor parameters, the syntactic element can be coded using fewer bits than in the case of the inter-predictor parameter being a member of the set of probabilistically advantageous intra-predictor parameters being a member of the remainder of the set of possible intra-predictor parameters that is disjoint to the set of probabilistically advantageous intra-predictor parameters.
[0123] The set of possible intra-predicted parameters may encompass several angular direction modes in which the current block is filled by copying from already coded / decoded adjacent samples by replicating along the angular direction of each mode / parameter, one DC mode in which the samples of the current block are set to a constant value determined based on already coded / decoded adjacent samples, for example by some averaging, and a plane mode in which the samples of the current block are set to a value distribution following the slopes of the linear functions x and y, for example, intercept determined based on already coded / decoded adjacent samples.
[0125] FIG. 12 illustrates the possibility of how a spatial prediction parameter substitute obtained from coexisting blocks (108) of the base layer can be utilized along syntactic elements signaled in the enhancement layer substream. FIG. 12, in an extended manner, shows the current block (28) along predetermined adjacent blocks (92 and 94) and adjacent already coded / decoded samples (102). FIG. 12 exemplarily illustrates the angular direction (112) as indicated by the spatial prediction parameter of the coexisting blocks (108).
[0127] The syntax element (114) signaled within the enhancement layer substream (6b) for the current block (28) signals an index (118) that is conditionally coded into a resulting list (122) of possible favorable intra-prediction parameters, which is exemplarily illustrated in the angle direction (124) as shown in FIG. 13, for example, or, if the actual intra-prediction parameter (116) is not in the most probable set (122), it identifies the actual intra-prediction parameter (116) in an index (123) into a list of possible intra-prediction modes (125) that excludes candidates of the possible list (122)—as shown in 127. If the actual intra-prediction parameter is located in the list (122), the coding of the syntax element may consume fewer bits. A syntax element may include, for example, a flag and an index field, the flag indicating whether the index points to a list (122) or a list (125) that includes or excludes members of the list (122), or the syntax element may include a field identifying one of an escape code or a member (124) of the list (122), in the case of an escape code, the second field identifies a member from the list (125) that includes or excludes the list (122). The order of the members (124) within the list (122) may be predetermined, for example, by default rules.
[0129] Thus, the scalable video decoder can obtain or retrieve a syntax element (114) from the enhancement layer substream (6b), and the scalable video encoder can input the syntax element (114) into the same, and the syntax element (114) can subsequently be used, for example, to index a spatial prediction parameter from the list (122). In forming the list (122), the substitution described above can be performed to determine whether the predetermined adjacent blocks (92 and 94) are of the same spatial prediction coding mode type. As described, if not, the coexisting block (108) is, for example, checked in turn to determine whether the same is a spatially predicted block, and if so, the spatial prediction parameter of the same, used to spatially predict this coexisting block (108), such as the angle direction (112), is included in the list (122). If the base layer block (108) also does not contain a suitable intra-predicted parameter, the list (122) may be left without any contribution from each predetermined adjacent block (92 or 94). For example, due to inter-predicted, both the coexisting block (108) and the predetermined adjacent blocks (92, 98) may have at least one of the members (124) determined unconditionally using a default intra-predicted parameter to avoid the list (122) being left empty due to a lack of suitable intra-predicted parameters. Alternatively, the list (122) may be allowed to be left empty.
[0131] Naturally, the perspective described in relation to FIGS. 11 to 13 can be combined with the perspective described above in relation to FIGS. 6 to 10. The intra prediction obtained using spatial intra prediction parameters derived through detour of the base layer according to FIGS. 11 to 13 can specifically represent the enhancement layer internal prediction signal (34) of the perspective of FIGS. 6 to 10 in order to be combined with the inter-layer prediction signal (38) described above in a spatially weighted manner.
[0133] With respect to the following drawings, how a scalable video encoder or decoder is implemented to form an embodiment of the present application according to a further aspect of the application as described above with respect to FIGS. 1 through 4 is described. Later, some additional embodiments for aspects described subsequently hereafter are presented by reference to aspects T and U.
[0135] Reference is made to FIG. 14, which shows pictures (22b and 22a) of the base layer signal (200) and the enhancement layer signal (400), respectively, in a temporarily registered manner. The part currently to be coded / decoded is shown in 28. According to this aspect, the base layer signal (200) is predictively restored by a scalable video decoder and predictively encoded by a scalable video encoder using base layer coding parameters that vary spatially with respect to the base layer signal. Spatial variation using the hatched part (132) within the base layer coding parameters that predictively code / restore a constant base layer signal (200), which is surrounded by non-hatched regions where the base layer coding parameters change during the transition from the hatched part (132) to the non-hatched region, is illustrated in FIG. 14. According to the aspect described above, the enhancement layer signal (400) is encoded / restored in blocks. The current part (28) is such a block. According to the view described above, the subblock subdivision for the current part (28) is selected from a set of possible subblock subdivisions based on spatial variation of base layer coding parameters within the coexisting part (134) of the base layer signal (200), that is, within the temporal coexisting part of the temporal corresponding picture (22a) of the base layer signal (200).In particular, instead of signaling within the enhancement layer substream (6b) subdivision information for the current part (28), the description suggests selecting a subblock subdivision from the set of possible subblock subdivisions of the current part (28) such that the selected subblock subdivision is the coarsest among the set of possible subblock subdivisions for subdividing the base layer signal (200), so that when the base layer coding parameters within each subblock subdivision are transmitted to the coexisting part (134) of the base layer signal, the base layer coding parameters within each subblock subdivision are sufficiently similar to one another. For the sake of understanding, reference is made to FIG. 15a. FIG. 15a shows a part (28) that records, using hatching, the spatial variation of base layer coding parameters within the coexisting part (134). In particular, part (28) shows three different subblock subdivisions applied to the block (28). In particular, a quad-tree subdivision is used exemplarily in the case of FIG. 15a. It is that the set of possible sub-block subdivisions is a quad-tree subdivision, or is defined by it, and that three instantiations of the sub-block subdivisions of the part (28) described in FIG. 15a belong to different hierarchical levels of the quad-tree subdivision of the block (28). From bottom to top, the level or coarseness of the subdivision of the block (28) into sub-blocks increases. At the highest level, the part (28) is left as is. At the next lower level, the block (28) is subdivided into four sub-blocks, and at least one of the next is further subdivided into four sub-blocks at the next lower level, and so on. In FIG. 15a, at each level, the quad-tree subdivision is selected where the number of sub-blocks is smallest, even though there are no sub-blocks that overlap with the base layer coding parameter change boundary.It can be seen in the case of FIG. 15a, and the quad-tree subdivision of block (28) that must be selected to subdivide block (28) is the lowest one shown in FIG. 15a. Here, the base layer coding parameters of the base layer are constant within each part coexisting in each sub-block of the sub-block subdivision.
[0137] Therefore, subdivision information for block (28) does not need to be signaled within the enhancement layer substream (6b), thus increasing coding efficiency. Furthermore, as described, the method for obtaining subdivision is applicable regardless of any registration of the current part (28) position with respect to any grid or sample batch of the base layer signal (200). In particular, subdivision derivation acts in the case of fractional spatial resolution ratios between the base layer and the enhancement layer.
[0139] Based on the sub-block subdivision of the part (28) determined in this way, the part (28) can be predictively restored / coded. Regarding the above description, it should be noted that there are different possibilities for "measure" the coarseness of the different available sub-block subdivisions of the current block (28). For example, the measure of coarseness can be determined based on the number of sub-blocks: the more sub-blocks each sub-block subdivision has, the lower the level. This definition was not clearly applied in the case of FIG. 15a, where the "measure of coarseness" is determined by the combination of the smallest size among all sub-blocks of each sub-block subdivision and the number of sub-blocks of each sub-block subdivision.
[0141] For completeness, FIG. 15b exemplarily illustrates a case where a possible sub-block subdivision is selected from the available subdivisions for the current block (28) by exemplarily using the subdivision of FIG. 35 as an available set. Different hatchings (and non-hatchings) show regions within each coexisting region in a base layer signal having the same base layer coding parameters associated with them.
[0143] As described above, the selection just described can be executed by traversing possible subblock subdivisions in some sequential order, such as the order of increasing or decreasing roughness levels, and selecting a possible subblock subdivision from which the possible subblock subdivisions within each subblock of that individual subblock are sufficiently similar to one another is applied no further (in the case of using traversal based on increasing roughness) or first (in the case of using traversal based on decreasing roughness levels). Alternatively, all possible subdivisions can be tested.
[0145] Although the broad term "base layer coding parameters" was used in the description of FIG. 14 and FIG. 15a, b, in a preferred embodiment, these base layer coding parameters represent base layer prediction parameters, that is, parameters related to the prediction formation of the base layer signal, but not related to the formation of the prediction residue. Accordingly, base layer coding parameters include, for example, prediction modes distinguishing between spatial prediction and temporal prediction, prediction parameters for blocks / parts of the base layer signal assigned to spatial prediction, such as each direction, and prediction parameters for blocks / parts of the base layer signal assigned to temporal prediction, such as motion parameters or similar ones.
[0147] Interestingly, however, the "sufficiency" of base-layer coding parameter similarity within a specific subblock may be determined or defined beyond merely a subset of base-layer coding parameters. For example, similarity may be determined based solely on prediction modes. Alternatively, prediction parameters that further adjust spatial and / or temporal predictions may form the parameters upon which the similarity of base-layer coding parameters within a specific subblock depends.
[0149] Furthermore, as already mentioned above, in order to be sufficiently similar to one another, the base layer coding parameters within a specific subblock may need to be completely identical to each other within that subblock. Alternatively, the similarity measurement used may need to be within a certain interval to satisfy the criteria for "similarity."
[0151] As described above, the selected subblock subdivision is not merely an amount that can be predicted or transmitted from the base layer signal. Rather, the base layer coding parameters themselves can be transmitted onto the enhancement layer signal to derive, based thereon, enhancement layer coding parameters for the subblocks of the subblock subdivision obtained by transmitting the selected subblock subdivision from the base layer signal to the enhancement layer signal. Insofar as motion parameters are involved, for example, scaling can be used to account for the transition from the base layer to the enhancement layer. Preferably, only such parts or syntactic elements of the base layer's predicted parameters are used to set the subblocks of the current partial subblock subdivision obtained from the base layer, which affect the similarity measurement. The fact that these syntactic elements of the prediction parameters within each subblock of the subblock subdivision selected by such measurement are somehow similar to one another guarantees that the syntactic elements of the base layer prediction parameters used to predict the corresponding prediction parameters of the subblocks of the current part (308) are similar or more identical to one another, and thus, in the first case allowing some changes, some meaningful “mean” of the syntactic elements of the base layer prediction parameters corresponding to the part of the base layer signal covered by each subblock may be used as a predictor for the corresponding subblock. However, even though mode-specific base layer prediction parameters participate in the similarity measurement decision, the part of the syntactic elements contributing to the similarity measurement may be used to predict the prediction parameters of the subblocks of the enhancement layer subdivision in addition to the subdivision movement itself, as if merely predicting or pre-setting the modes of the subblocks of the current part (28).
[0153] One such possibility that does not use only the subdivided inter-layer prediction from the base layer to the enhancement layer will be described in relation to the following figures, FIG. 16. FIG. 16 shows pictures (22b) of the enhancement layer signal (400) and pictures (22a) of the base layer signal (200) in a manner registered along the representation time axis (24).
[0155] According to the embodiment of FIG. 16, the base layer signal (200) is predictively encoded by the use of a scalable video encoder and predictively restored by a scalable video decoder by subdividing frames (22a) of the base layer signal (200) into inter-blocks and intra-blocks. According to the example of FIG. 16, the subsequent subdivision is performed in a two-stage manner: among other things, the frames (22a) are regularly subdivided into maximum code units or maximum blocks, indicated by the reference numeral (302) in FIG. 16, using a double line along its circumference. Then, each such maximum block (302) becomes subject to hierarchical quad-tree subdivision into coding units forming the aforementioned intra-blocks and inter-blocks. Thus, they are the leaves of the quad-tree subdivisions of the maximum blocks (302). In FIG. 16, the reference numeral (304) is used to denote these leaf blocks or coding units. General, continuous lines are used to denote the perimeter of these coding units. While spatial intra-prediction is used for intra-blocks, temporal inter-prediction is used for inter-blocks. The prediction parameters associated with spatial intra and temporal inter-prediction, however, are set to units of smaller blocks into which the intra- and inter-blocks or coding units are subdivided. Such subdivision is exemplarily illustrated for one of the coding units (304) in FIG. 16 using the reference numeral (306) representing the smaller blocks. The smaller blocks (304) are described using dashed lines. In the case of the embodiment of FIG. 16, the spatial video encoder has the opportunity to select between spatial prediction on one side and temporal prediction on the other side for each coding unit (304) of the base layer.However, as far as the enhancement layer signal is concerned, the degrees of freedom are increased. In particular, here, the frames (22b) of the enhancement layer signal (400) are assigned to each of the sets of prediction modes, including spatial intra-prediction and temporal inter-prediction, as well as inter-layer prediction as described in more detail below, in the coding units into which the frames (22b) of the enhancement layer signal (400) are subdivided. This subdivision into these coding units can be performed in a similar manner as described for the base layer signal: above all, the frames (22b) can be regularly subdivided into columns and rows of maximum blocks described using double-lines (double lines) that are subdivided into coding units described using continuous lines, which is common in the hierarchical quad-tree subdivision process.
[0157] One such coding unit (308) of the current picture (22b) of the enhancement layer signal (400) is exemplarily assumed to be assigned to an inter-layer prediction mode and is illustrated using hatching. In a manner similar to FIGS. 14, 15a and 15b, FIG. 16 shows in (312) how the subdivision of the coding unit (308) is predictively derived by a local shift from the base layer signal. In particular, the local region superimposed by the coding unit (308) in (312) is shown. Within this region, dotted lines indicate boundaries between adjacent blocks of the base layer signal, or, more generally, boundaries where the coding parameters of the base layer may change. These boundaries, as such, may be the boundaries of the prediction blocks (306) of the base layer signal (200), and may partially coincide with the boundaries of the more adjacent maximum coding units (302) or adjacent coding units (304) of the base layer signal (200), respectively. The dashed lines of (312) represent the subdivision of the current coding units (308) into prediction blocks as induced / selected by local movement from the base layer signal (200). Details regarding local movement have been described above.
[0159] As already mentioned above, according to the embodiment of FIG. 16, only subdivision into prediction blocks is adopted from the base layer. Rather, as used within the region (312), the prediction parameters of the base layer signal can also be used to derive prediction parameters to be used to perform predictions on the prediction blocks of the coding unit (308) of the enhancement layer signal (400).
[0161] In particular, according to the embodiment of FIG. 16, not only are subdivisions into prediction blocks derived from the base layer signal, but prediction modes used in the base layer signal (200) are also derived, for the purpose of coding / restoring each region locally covered by each sub-block of the derived subdivision. One example is as follows: In order to derive the subdivision of the coding unit (308) according to the above description, prediction modes and mode-specific prediction parameters used in connection with the base layer signal (200) along with the relevant ones can be used to determine the "similarity" discussed above. Thus, different hatchings shown in FIG. 16 can correspond to different prediction blocks (306) of the base layer, each of which can have an intra or inter prediction mode, that is, a spatial or temporal prediction mode associated with it. As described above, in order to be "sufficiently similar," the prediction modes used within the regions coexisting in each sub-block of the coding unit (308), and the prediction parameters specific to each prediction mode within the sub-region, may be completely identical to each other. Alternatively, some modifications may be allowed.
[0163] In particular, according to the embodiment of FIG. 16, all blocks shown by the hatching extending from the top left to the bottom right can be set as intra-prediction blocks of the coding unit (308) because the locally corresponding part of the base layer signal is covered by prediction blocks (306) having a spatial intra-prediction mode associated with it, whereas other parts, i.e., the hatched part extending from the bottom left to the top right, can be set as inter-prediction blocks because the locally corresponding part of the base layer signal is covered by prediction blocks (306) having a temporal inter-prediction mode associated with it.
[0165] According to an alternative embodiment, the deviation of prediction details for performing prediction within the coding unit (308) may be stopped here, that is, limited to the allocation of these prediction blocks to being coded using temporal prediction and to being coded using non-temporal or spatial prediction, and the induction of subdivision into the prediction blocks of the coding unit (308), which does not follow the embodiment of FIG. 16.
[0167] According to the latter embodiment, all prediction blocks of the coding unit (308) having the non-temporal prediction assigned thereto are subject to non-temporal, spatial intra-prediction while using prediction parameters derived from prediction parameters of locally matching intra-blocks of the base layer signal (200), such as enhancement layer prediction parameters of these non-temporal mode blocks. Such derivation may include spatial prediction parameters of locally coexisting intra-blocks of the base layer signal (200). Such spatial prediction parameters may be, for example, an indication of the angular direction in which the spatial prediction is performed. As described above, the definition of similarity by itself requires that the spatial base layer prediction parameters overlapped by each non-temporal prediction block of the coding unit (308) are identical to each other, or that for each non-temporal prediction block of the coding unit (308), some averaging of the spatial base layer prediction parameters overlapped by each non-temporal prediction block is used to derive the prediction parameters of each non-temporal prediction block.
[0169] Alternatively, all prediction blocks of a coding unit (308) having a non-temporal prediction mode assigned thereto may be subject to inter-layer prediction in the following manner: above all, within regions spatially coexisting with the non-temporal prediction mode prediction blocks of the coding unit (308), the base layer signal is subject to resolution or quality improvement to obtain an inter-layer prediction signal, and then these prediction blocks of the coding unit (308) are predicted using the inter-layer prediction signal.
[0171] The scalable video decoder and encoder may, by default, subject the entire coding unit (308) to inter-layer prediction or spatial prediction, respectively. Alternatively, the scalable video encoder / decoder may support signaling within coded video data stream signals in which a version involving non-temporal prediction mode prediction blocks of the coding unit (308) is utilized, and both alternatives. In particular, the decision between the two alternatives may be signaled within a data stream of any granularity, for example, individually for the coding unit (308).
[0173] Insofar as other prediction blocks of the coding unit (308) are involved, the same may be subject to temporal inter-prediction using prediction parameters that can be derived from the prediction parameters of locally matching inter-blocks, just as in the case of non-temporal prediction mode prediction blocks. The derivation may, in turn, be related to motion vectors assigned to corresponding parts of the base layer signal.
[0175] For all other coding units utilizing either the temporal inter-prediction mode or the spatial intra-prediction mode assigned thereto, the same is subject to spatial prediction or temporal prediction in the following manner: in particular, when the same is assigned to each coding unit, it is further subdivided into prediction blocks having a prediction mode assigned to the same prediction mode and common to all prediction blocks within the coding unit. It is such that coding units having a spatial intra-prediction mode or a temporal inter-prediction mode associated with it, which are different from coding units such as the coding unit (308) to which the inter-layer prediction mode is associated, are subdivided into prediction blocks of the same prediction mode, that is, prediction modes inherited from each coding unit derived by the subdivision of each coding unit.
[0177] The subdivision of all coding units including (308) may be a quad-tree subdivision into prediction blocks.
[0179] An additional difference between coding units of inter-layer prediction mode and coding units of spatial intra-prediction mode or temporal inter-prediction mode, such as coding unit (308), is that when prediction blocks of temporal inter-prediction mode coding units or spatial intra-prediction mode coding units are targeted for temporal prediction and spatial prediction, respectively, prediction parameters are set without any dependence on the base layer signal (200), for example, in a manner signaled within the enhancement layer substream (6b). Subdivisions of coding units other than those having the inter-layer prediction mode associated with it, such as coding unit (308), may be signaled within the enhancement layer signal (6b). This is because inter-layer prediction mode coding units, such as 308, which have the advantage of low bit rate signaling, are required: according to the embodiment, the mode indicator itself for coding unit (308) does not need to be signaled within the enhancement layer substream. Optionally, additional parameters may be transmitted to coding unit (308) as prediction parameter residues for individual prediction blocks. Additionally or alternatively, the predicted residue for the coding unit (308) may be transmitted / signaled within the enhancement layer substream (6b). While the scalable video decoder retrieves information from the enhancement layer substream, the scalable video encoder according to the present embodiment determines these parameters and inputs the same into the enhancement layer substream (6b).
[0181] In other words, the prediction of the base layer signal (200) can be performed using base layer coding parameters in a manner similar to how the same thing changes spatially across the base layer signal (200) in units of base layer blocks (304). The prediction modes available for the base layer may include, for example, spatial and temporal predictions. The base layer coding parameters may further include individual prediction mode parameters, such as motion vectors with respect to the temporal prediction blocks (304) and angular directions with respect to the spatial prediction blocks (304). The later individual prediction mode parameters may change according to the base layer signal in units smaller than the base layer blocks (34), that is, in the previously mentioned prediction blocks (306). To satisfy the sufficient similarity requirement mentioned above, it may be a requirement that the prediction modes of all base layer blocks (304) overlapping the area of each possible sub-block subdivision are identical to one another. Only then can each sub-block subdivision be finally selected to obtain the selected sub-block subdivision. The above requirement, however, may be stricter: the individual prediction parameters of the prediction modes of the prediction blocks overlapping the common regions of each subblock subdivision may be identical to one another. Subblock subdivisions satisfying these requirements regarding corresponding regions within the base layer signal and each subblock of each subblock subdivision may be selected only to obtain the finally selected subblock subdivision.
[0183] In particular, as briefly summarized above, there are different possibilities regarding how to make a selection from a set of possible subblock divisions. To explain this in more detail, FIGS. 15c and FIGS. 15d are referenced. Let us imagine that the set (352) passes through all possible subblock divisions (354) of the current block (28). Naturally, FIGS. 15c is merely an illustrative example. The set (352) of possible or available subblock divisions of the current block (28) can be signaled within a coded data stream, for example, for a sequence of pictures or similar, or can be known to a scalable video encoder and a scalable video decoder by default. According to the example of FIG. 15c, each member of the set (352), i.e., each available subblock subdivision (354), is subject to check (356), which checks whether the regions subdivided by transmitting each subblock subdivision (354) from the enhancement layer to the base layer in a coexisting portion (108) of the base layer signal are overlapped by base layer coding parameters, prediction blocks (306), and coding units (304) that satisfy sufficient similarity requirements. For example, refer to the exemplary subdivision labeled (354). According to this exemplary available subblock subdivision, the current block (28) is subdivided into four quadrants / subblocks (358), and the upper left subblock corresponds to the region (362) in the base layer.
[0185] Clearly, this region (362) overlaps with four blocks of the base layer, namely two coding units (304) and two prediction blocks (306) that are not further subdivided into prediction blocks, and thus represent the prediction blocks themselves. Thus, if the base layer coding parameters of all these prediction blocks overlapping the region (362) satisfy the similarity criteria, this is an additional case for all subblocks / quadrants of the base layer coding parameters and possible subblock subdivisions overlapping their corresponding regions, and then these possible subblock subdivisions (354) belong to the set (364) of subblock subdivisions and satisfy the sufficiency requirement for all regions covered by the subblocks of each subblock subdivision. Among this set (364), the coarsest subdivision is selected as shown by the arrow (366), thereby obtaining the subblock subdivision (368) selected from the set (352). Clearly, it is desirable to attempt to avoid performing verification (356) for all members of set (352), and thus, as shown in FIG. 15d and described above, possible subdivisions (354) can be traversed in order of increasing or decreasing coarseness. Such traversal is illustrated using a double-headed arrow (372). FIG. 15d illustrates that the level or measurement of coarseness may be the same for at least some of the available subblock subdivisions. In other words, the order according to the increasing or decreasing level of coarseness may be ambiguous. However, this hinders the search for the "coarsest subblock subdivision" belonging to set (364), because such equally coarse possible subblock subdivisions may belong to set (364).Accordingly, as soon as the result of the criterion check (356) changes from satisfied to unsatisfied, it is discovered, and when the roughest possible sub-block subdivision (368) traverses in the direction of the level of increasing roughness, it is together with the next one of the final traversed possible sub-block subdivisions, which is the sub-block subdivision (354) of the selected sub-block, or when switching from unsatisfied to satisfied while traversing along the direction of decreasing roughness level, it is together with the most recently traversed sub-block subdivision, which is the sub-block subdivision (368).
[0187] With respect to the following drawings, it is explained how a scalable video encoder or decoder, as described above with respect to FIGS. 1 through 4, can be implemented to form an embodiment of the present application according to a further aspect of the application. Possible embodiments of the aspects described below are referred to as aspects K, A, and M and are presented below.
[0189] To explain the above perspective, reference is made to FIG. 17. FIG. 17 illustrates the possibility of a temporal prediction (32) of the current part (28). The following description of FIG. 17 may be combined with the description related to FIG. 6 through 10 insofar as the combination of inter-layer prediction signals is involved, or may be combined with the description related to FIG. 11 through 13 as a temporal inter-layer prediction mode.
[0191] The situation shown in FIG. 17 corresponds to that shown in FIG. 6. It shows base layer and enhancement layer signals (200 and 400), in which already coded / decoded portions are depicted using hatching. Within the enhancement layer signal (400), the currently coded / decoded portion has adjacent blocks (92 and 94), where, exemplarily, block (94) is depicted above block (92) and to the left of the current portion (28), and both blocks (92 and 94) have the same size as the current block (28), exemplarily. However, matching sizes are not mandatory. Rather, the portions of the blocks into which the picture (22b) of the enhancement layer signal (400) is subdivided are subdivided into different sizes. They are not limited to a rectangular shape. They may be rectangular or other shapes. The current block (28) has additional adjacent blocks not specifically depicted in FIG. 17, which, however, have not yet been decoded / coded, that is, they follow the coding / decoding order and are not available for prediction. Beyond this, there may be other blocks in addition to the blocks (92 and 94) that have already been coded / decoded according to the coding / decoding order, such as block (96), which are adjacent to the current block (28)—exemplarily diagonally to the top-left corner of the current block (28), and blocks (92 and 94) are predetermined adjacent blocks that serve to predict the inter-prediction parameters for the current block (28), which is the subject of the inter-prediction (30) in the example considered herein. The number of such predetermined adjacent blocks is not limited to two. It may be higher or just one. A discussion of possible embodiments is presented in relation to FIGs. 36 through 38.
[0193] Scalable video encoders and scalable video decoders can determine a set of predetermined adjacent blocks, for example, blocks (92 to 96) that depend on a predetermined sample location (98) within the current part (28), such as the top-left sample, wherein the blocks (92, 94) are blocks from a set of already coded adjacent blocks. For example, only the already coded adjacent blocks of the current part (28) can form a set of "predetermined adjacent blocks" that includes sample locations immediately adjacent to the predetermined sample location (98). Additional possibilities are described in relation to FIGS. 36 to 38.
[0195] In any case, a portion (502) of a previously coded / decoded picture (22b) of an enhancement layer signal (400), which is placed from a coexisting location of the current block (28) by a motion vector (504) according to the decoding / coding order, includes restored sample values based on which sample values of the portion (28) can be predicted by interpolation or replication. For this reason, the motion vector (504) is signaled in the enhancement layer substream (6b). For example, a temporal prediction parameter for the current block (28) represents a displacement vector (506) indicating the placement of the portion (502) from a coexisting portion of the portion (28) in the reference picture (22b) to be optionally replicated by interpolation on the samples of the portion (28).
[0197] In any case, when predicting the current block (28) in time, the scalable video decoder / encoder has already restored (and encoded in the case of the encoder) the base layer (200) using the base layer substream (6a), using block-by-block prediction as described above, for example, by using a block-by-block selection between spatial and temporal prediction modes, at least with respect to the relevant spatial correspondence region of the temporal correspondence picture (22a).
[0199] In FIG. 17, several blocks (104) into which a time-aligned picture (22a) of a base layer signal (200) is subdivided are illustrated as an example, which corresponds locally to the current part (28) and is located in the surrounding area. In the base layer substream (6a) for the blocks (104) in the base layer signal (200), spatial prediction parameters are included or signaled, such as in the very case having spatial prediction blocks in the enhancement-layer signal (400), where the selection of the spatial prediction mode is signaled.
[0201] For example, to allow the restoration of an enhancement layer signal from a data stream coded for a selected block (28) where temporal intra-layer prediction (32) is selected, inter-prediction parameters such as motion parameters are determined and used according to any of the following methods:
[0203] The first possibility is described with respect to FIG. 18. In particular, first, a set (512) of motion parameter candidates (514) is gathered or generated from adjacent already restored blocks of the frame, such as predetermined blocks (92 and 94). The motion parameters may be motion vectors. The motion vectors of blocks (92 and 94) are each symbolized using arrows (516 and 518) having 1 and 2 written thereon. As shown, these motion parameters (516 and 518) can directly form a candidate (514). Some candidates may be formed by combining motion vectors such as 518 and 516, as shown in FIG. 18.
[0205] Furthermore, a set (522) of one or more base layer motion parameters (524) of a block (108) of a base layer signal (200) coexisting in part (28) is gathered or generated from the base layer motion parameters. In other words, motion parameters associated with a block (108) coexisting in the base layer are used to derive one or more base layer motion parameters (524).
[0207] One or more base layer motion parameters (524), or a scaled version thereof, are added (526) to a set (512) of motion parameter candidates (514) to obtain an expanded set (528) of motion parameter candidates. This can be done in any manifold way, such as simply attaching the base layer motion parameters (524) to the end of the list of candidates (514), or in a different way, as illustrated in FIG. 19a.
[0209] At least one of the motion parameter candidates (532) of the extended motion parameter candidate set (528) is selected, and the temporal prediction (32) of the part (28) by motion compensation prediction is performed using the selected one of the motion parameter candidates of the extended motion parameter candidate set. The selection (534) may be signaled within a data stream such as a substream (6b) of the part (28) in the manner of an index (536) to the list / set (528), or it may be performed differently as described with respect to FIG. 19a.
[0211] As described above, it can be determined whether the base layer motion parameter (523) is coded in a coded data stream such as the base layer substream (6a) using merging, and if the base layer motion parameter (523) is coded in a coded data stream using merging, the addition (526) can be suppressed.
[0213] The motion parameters mentioned according to FIG. 18 may relate only to motion vectors (motion vector prediction), or to a complete set of motion parameters including the number of motion hypotheses per block, reference indices, and splitting information (merging). Thus, the "scaled version" may be derived from the scaling of motion parameters used in the base layer signal according to the ratio of spatial resolution between the base and enhancement layer signals in the case of spatial scalability. The coding / decoding of base layer motion parameters of the base layer signal by the coded data stream method may include motion vector prediction such as spatial, temporal, or merging.
[0215] The incorporation (526) of motion parameters (523) used in the coexisting portion (108) of the base layer signal into a set (528) of merged / motion vector candidates (532) enables very effective indexing among one or more inter-layer candidates (524) and intra-layer candidates (514). The selection (534) may include explicit signaling to an index into an extended set / list of motion parameter candidates in an enhancement layer signal (6b) per prediction block, per coding unit, or similar. Alternatively, the selection index (536) may be inferred from inter-layer information or other information of the enhancement layer signal (6b).
[0217] According to the possibility of FIG. 19a, the formation (542) of the final motion parameter candidate list for the enhancement layer signal for part (28) is performed only optionally, as described in relation to FIG. 18. It is possible that the same thing (528) or (512) is the same thing. However, the list (528 / 512) is sorted based on base layer motion parameters, such as the example of motion parameters represented by the motion vector (523) of the coexisting base layer block (108) (544). For example, the rank of the members, i.e., motion parameter candidates, or (532) or (514) of the list (528 / 512) is determined based on the deviation of each of the same thing from the potentially scaled version of the motion parameter (523). The greater the deviation, the lower the rank (523 / 512) of each member in the sorted list (528 / 512'). Alignment (544) may thus include determining the deviation measurement per member (532 / 514) of the list (528 / 512). The selection (534) of a candidate (532 / 512) within the aligned list (528 / 512') is performed by being controlled via an explicit signaled index syntax element (536) in the coded data stream, which is to obtain an enhancement layer motion parameter from the aligned motion parameter candidate list (528 / 512') for a portion (28) of the enhancement layer signal, and subsequently, the temporal prediction (32) by motion compensation prediction of the portion (28) of the enhancement layer signal is performed using the selected motion parameter pointed to (534) by the index (536).
[0219] Regarding the motion parameters mentioned in FIG. 19a, the same as mentioned above applies in relation to FIG. 18. Decoding of base layer motion parameters (520) from a coded data stream may (optionally) include spatial or temporal motion vector prediction or merging. Ordering is performed according to a measure that measures the difference between the base layer motion parameters of the base layer signal and each enhancement layer motion parameter candidate, in relation to the block of base layer signals coexisting in the current block of the enhancement layer signal, as just mentioned. This means that, for the current block of the enhancement layer signal, a list of enhancement layer motion parameter candidates can be determined first. Then, ordering is performed as just mentioned. Subsequently, the selection is performed by explicit signaling.
[0221] Alignment (544) is, alternatively, performed according to a measurement that measures the difference between base layer motion parameters (546) of spatially and / or temporally adjacent blocks (548) in the base layer and base layer motion parameters (523) of the base layer signal that are associated with the block (108) of the base layer signal that coexists in the current block of the enhancement layer signal. The alignment determined in the base layer is transmitted to the enhancement layer, and the enhancement layer motion parameter candidates are aligned in the same manner as the alignment is predetermined for the corresponding base layer candidates. In this regard, it may be said that the base layer motion parameter (546) corresponds to the enhancement layer motion parameter of the adjacent enhancement layer block (92, 94) when the relevant base layer block (548) is spatially and temporally associated with the adjacent enhancement layer block (92 and 94) that is associated with the enhancement layer motion parameters considered. Alternatively, when the adjacency relationship between the block (108) coexisting in the current enhancement layer block (28) and the associated base layer block (548) (left adjacency, top adjacency, A1, A2, B1, B2, B0, or FIG. 36 through 38 for further embodiments) is the same as the adjacency relationship between the current enhancement layer block (28) and each enhancement layer adjacency block (92, 94), it can be said that the base layer motion parameter (546) corresponds to the enhancement layer motion parameter of the adjacency layer block (92, 94).
[0223] To explain this in more detail, reference is made to FIG. 19b. FIG. 19b illustrates the first of the just-mentioned alternatives for inducing enhancement layer alignment for a list of motion parameter candidates by using base layer hints. FIG. 19b shows the current block (28) and three different predetermined sample locations of the same: the top-left sample (581), the bottom-left sample (583), and the top-right sample (585). The above example is interpreted merely for illustration purposes. Let us exemplarily imagine a set of predetermined adjacent blocks passing through four types of adjacents: an adjacent block (94a) located above sample location (581) and covering the immediately adjacent sample location (587), and an adjacent block (94b) located immediately above sample location (585) and covering or including the adjacent sample location (589). Similarly, adjacent blocks (92a and 92b) are blocks located to the left of sample locations (581 and 583), immediately adjacent to and containing sample locations (591 and 593).
[0225] According to the alternative of FIG. 19b, for each predetermined adjacent block (92a, b), a coexisting block of the base layer is determined. For example, for this reason, the top-left sample (595) of each adjacent block is used, and the current block (28) regarding the top-left sample (581) is formally referred to in FIG. 19a. This is illustrated in FIG. 19b using dashed arrows. By this measurement, for each of the predetermined adjacent blocks, a corresponding block (597) can be additionally found in the coexisting block (108) that coexists with the current block (28). Using their respective differences for the motion parameters (m1, m2, m3, and m4) of coexisting base layer blocks (597) and the base layer motion parameter m of coexisting base layer block (108), the enhancement layer motion parameters of predetermined adjacent blocks (92a, b and 94a, b) are sorted within a list (528 or 512). For example, the greater the distance of either m1 - m4, the greater the corresponding enhancement layer motion parameters M1 - M4 may be, i.e., higher indices may be required at the same index from the list (528 / 512'). For distance measurement, an absolute difference may be used. In a similar manner, motion parameter candidates (532 or 514) may be rearranged within a list regarding their ranks, which are combinations of enhancement layer motion parameters M1 - M4.
[0227] FIG. 19c illustrates an alternative in which corresponding blocks in the base layer are determined in a different way. In particular, FIG. 19c shows the coexisting blocks (108) of the current block (28) and the predetermined adjacent blocks (92a,b and 94a,b) of the current block (28). According to the embodiment of FIG. 19c, the base layer blocks corresponding to those of the current block (28), namely (92a,b and 94a,b), are determined in such a way that these base layer blocks are related to the enhancement layer adjacent blocks (92a,b and 94a,b) using the same adjacent determination rules for determining these base layer adjacent blocks. In particular, FIG. 19c shows the densely determined sample positions of the coexisting blocks (108), namely the top-left, bottom-left, and top-right sample positions (601). Based on the sample locations, the four adjacent blocks of block (108) are determined in the same manner as described for the enhancement layer adjacent blocks (92a,b and 94a,b) with respect to the predetermined sample locations (581, 583 and 585) of the current block (28): the four base layer adjacent blocks (603a, 603b, 605a and 605b) are found in this manner, (603a) clearly corresponds to the enhancement layer adjacent block (92a), the base layer block (603b) corresponds to the enhancement layer adjacent block (92b), the base layer block (605a) corresponds to the enhancement layer adjacent block (94a), and the base layer block (605b) corresponds to the enhancement layer adjacent block (94b). In the same manner as previously described, the distances of base layer motion parameters M1 to M4 of base layer blocks (903a,b and 905a,b) and base layer motion parameter m of coexisting base layer block (108) are used to sort motion parameter candidates within a list (528 / 512) formed from motion parameters M1 to M4 of enhancement layer blocks (92a,b and 94a,b).
[0229] According to the possibility of FIG. 20, the formation (562) of the final motion parameter candidate list for the enhancement layer signal for part (28) is performed only optionally as described with respect to FIG. 18 and / or 19. It may be the same as (528) or (512) or (528 / 512') and the reference numeral (564) is used in FIG. 20. According to FIG. 20, the index (566) pointing to the motion parameter candidate list (564) is determined, for example, by relying on the index (567) for the motion parameter candidate list (568) used to code / decode the base layer signal for the coexisting block (108). For example, in restoring a base layer signal in block (108), a list of motion parameter candidates (568) may be determined based on the motion parameters (548) of adjacent blocks (548) of block (108) having the same adjacency relationship (left adjacency, top adjacency, A1, A2, B1, B2, B0, or refer to FIG. 36 to 38 for further examples) as the adjacency relationship between the current block (28) and predetermined adjacent enhancement layer blocks (92, 94) and the determination (572) of list (567) using potentially the same configuration rule as used in a formation (562) such as alignment among the list members of lists (568 and 564). More generally, the index (566) for an enhancement layer can be determined in such a way that adjacent enhancement layer blocks (92, 94) are indicated for the index that coexists with the base layer block (548) related to the indexed base layer candidates, i.e., what the index (567) points to. The index (567) can function as a meaningful prediction of the index (566). The enhancement layer motion parameter is subsequently determined using the index (566) to the motion parameter candidate list (564), and the motion compensation prediction of the block (28) is performed using the determined motion parameter.
[0231] Regarding the motion parameters mentioned in FIG. 20, they apply to FIG. 18 and 19 in the same way as mentioned above.
[0233] With respect to the following drawings, it is explained how a scalable video encoder or decoder can be implemented to form an embodiment of the present application according to a further aspect of the application, as mentioned above with respect to FIGS. 1 through 4. Detailed embodiments of the aspects described herein are described below by reference to aspect V.
[0235] This aspect relates to residual coding within an enhancement layer. In particular, FIG. 21 exemplarily shows a picture (22a) of a base layer signal (200) and a picture (22b) of a temporal registration method of an enhancement layer signal (400). FIG. 21 illustrates a restoration within a scalable video decoder, or an encoding method within an enhancement layer signal and a scalable video encoder, focusing on a predetermined portion (404) and a predetermined conversion factor block of conversion factors (402) representing the enhancement layer signal (400). In other words, the conversion factor block (402) represents a spatial decomposition of a portion (404) of the enhancement layer signal (400). As already described, according to the coding / decoding order, the corresponding portion (406) of the base layer signal (200) may already be decoded / coded at the time the conversion factor block (402) is decoded / coded. As far as the base layer signal (200) is concerned, predictive coding / decoding may be used therein and includes signaling of a base layer residual signal within a coded data stream, such as a base layer substream (6a).
[0237] According to the view described with respect to FIG. 21, the scalable video decoder / encoder utilizes the fact that the measurement (evaluation, 408) of the base layer residual signal or base layer signal in the part (406) coexisting with the part (404) can yield an advantageous choice of subdivision of the transform factor block (402) into subblocks (412). In particular, several possible subblock subdivisions for subdividing the transform factor block (402) into subblocks can be supported by the scalable video decoder / encoder. These possible subblock subdivisions can regularly subdivide the transform factor block (402) into rectangular subblocks (412). It is possible that the transformation coefficients (414) of the transformation coefficient block (402) can be arranged in rows and columns, and according to possible sub-block subdivisions, these transformation coefficients (414) are regularly clustered into sub-blocks (412) such that the sub-blocks (412) themselves are arranged in columns and rows. Evaluation (408) makes it possible to set the ratio between the number of columns and the number of rows of the sub-blocks (412), i.e., their width and height, in the most efficient way for coding the transformation coefficient block (402) using the selected sub-block subdivision. For example, if the measurement (408) reveals the restored base layer signal (200) within the coexisting portion (406), or at least if the base layer residual signal within the corresponding portion (406) mainly constitutes a horizontal edge in the spatial region, it will likely be present with quantized transform coefficients, i.e., non-zero transform coefficient levels, i.e., near the zero horizontal frequency side of the transform coefficient block (402), which are significant to the transform coefficient block (402).In the case of a vertical edge, the conversion factor block (402) will likely be present with non-zero conversion factor levels at a location near the zero vertical frequency side of the conversion factor block (402). Therefore, in the first instance, the subblocks (412) should be selected to be smaller along the horizontal direction and longer along the vertical direction, and in the second instance, the subblocks should be selected to be smaller in the vertical direction and longer in the horizontal direction.
[0239] It is that the scalable video decoder / encoder will select a subblock subdivision from a set of possible subblock subdivisions based on the base layer signal or base layer residual signal. The coding (414) or decoding of the transform factor block (402) will be performed by applying the selected subblock subdivision. In particular, the positions of the transform factors (414) will be traversed in units of the subblock (412), and all positions within one subblock will be traversed in a continuous manner as soon as they proceed to the next subblock in the defined order of subblocks. The currently visited subblock is such as the subblock (412) for the reference numeral (412) shown exemplarily in 22 of FIG. 40, and a syntax element is signaled within a data stream such as an enhancement layer substream (6b) indicating whether or not the currently visited subblock has any significant transform factors. In FIG. 21, syntax elements (416) are illustrated for two exemplary subblocks. If each syntactic element of an individual subblock represents a non-important transformation factor, nothing else needs to be transmitted within the enhancement layer substream (6b) or data stream. Rather, the scalable video decoder can set the transformation factors within that subblock to zero. However, if the syntactic element (416) of each subblock indicates that this subblock has some important transformation factor, additional information regarding the transformation factor within that subblock is signaled within the substream (6b) or data stream. In terms of decoding, the scalable video decoder decodes the syntactic elements (418) representing the levels of the transformation factors within each subblock from the substream (6b) or data stream.The syntax elements (418) can signal the location of important transformation coefficients within each subblock according to the scan order of these transformation coefficients within each subblock, and optionally according to the scan order of the transformation coefficients within each subblock.
[0241] FIG. 22 illustrates different possibilities for performing a selection among the possible sub-block subdivisions in the measurement (408). FIG. 22 again illustrates a portion (404) of an enhancement layer signal, in which the transform factor block (402) represents the spectral decomposition of the latter portion (404). For example, the transform factor block (402) represents the spectral decomposition of an enhancement layer residual signal with a scalable video decoder / encoder that predictively codes / decodes the enhancement layer signal. In particular, transform coding / decoding is performed in a block-by-block manner, i.e., into blocks into which the pictures (22b) of the enhancement layer signal are subdivided, and the transform coding / decoding is used by the scalable video decoder / encoder to encode the enhancement layer residual signal. FIG. 22 illustrates a corresponding or coexisting portion (406) of a base layer signal, wherein the scalable video decoder / encoder also applies predictive encoding / decoding to the base layer signal while utilizing transform coding / decoding with respect to the predictive residue of the base layer, i.e., with respect to the base layer residue signal. In particular, block-wise transformation is used with respect to the base layer residue signal, that is, the base layer residue signal is transformed block-wise with individually transformed blocks indicated by dashed lines in FIG. 22. As illustrated in FIG. 22, the block boundaries of the transformed blocks of the base layer do not necessarily have to coincide with the contour of the coexisting portion (406).
[0243] Nevertheless, to perform the measurement (408), one or a combination of the following options A to C may be used.
[0245] In particular, a scalable video decoder / encoder may perform a transformation (422) on a restored base layer signal or base layer residual signal within a portion (406) to obtain a transformation factor block (424) of transformation factors that matches the size of the transformation factor block (402) to be coded / decoded. An inspection of the distribution of values of the transformation factors within the transformation factor blocks (424, 426) may be used to appropriately set the size of the sub-blocks (412) according to the direction of horizontal frequencies, i.e., 428, and the size of the sub-blocks (412) according to the direction of vertical frequencies, i.e., 432.
[0247] Additionally or alternatively, a scalable video decoder / encoder may examine all transform factor blocks of base layer transform blocks (434), which are differently hatched and illustrated in FIG. 22, that at least partially overlap the coexisting portion (406). In the exemplary case of FIG. 22, there are four base layer transform blocks, and their transform factor blocks will be examined. In particular, all of these base layer transform blocks may be of different sizes from one another and additionally differ in size with respect to the transform factor block (412), and scaling (436) may be performed on the transform factor blocks of the overlap of base layer transform blocks (434) to derive an approximation of the transform factor block (438) of the spectral decomposition of the base layer residual signal within the portion (406). The distribution of values of the transformation coefficients within the transformation coefficient block (438), i.e., 442, can be used within the measurement to appropriately select the sub-block size (428) and (432), thereby selecting the sub-block subdivision of the transformation coefficient block (402).
[0249] Additional alternatives that may be used to perform the measurement (408) include, for example, determining the main gradient direction or using edge detection (444) to determine the gradient determined within the coexisting part (406) or the extension direction of the detected edges to properly set the subblock size (428 and 432), and examining the restored base layer signal or base layer residual signal within the spatial area.
[0251] Although not specifically described above, in traversing the locations of the units and transformation coefficients of the subblocks (412), it may be desirable to traverse the subblocks (412) in order from the zero-frequency corner of the transformation coefficient block to the highest-frequency corner (402) of the block, that is, from the top-left corner of FIG. 21 to the bottom-right corner of FIG. 21. Furthermore, entropy coding may be used to signal syntactic elements within the data stream (6b): that the syntactic elements (416 and 418) may be coded using some other forms of entropy coding or entropy coding such as arithmetic or variable-length coding. Also, the order of the traversed subblocks (412) may depend on the subblock shape selected according to 408: for subblocks selected to be wider than their height, the order of traversal may be to traverse the subblocks row-direction first and then proceed to the next row, etc. Beyond this, it should be noted again that the base layer information used to select the subblock size may be a base layer signal that is restored itself or a base layer residual signal.
[0253] In the following, different embodiments that can be combined with the aspects described above are described. The embodiments described below relate to many different aspects or methods for rendering much more efficient scalable video coding. Partially, the above aspects are described in more detail below to illustrate their derived embodiments, while retaining only the general concept. These descriptions presented below may be used to obtain extensions or alternatives to the above embodiments / perspectives. Most of the embodiments described below, however, relate to sub-perspectives and can be combined with the aspects already described above, which can optionally, but not necessarily, be executed simultaneously with the above embodiments within a single scalable video decoder / encoder.
[0255] To facilitate a better understanding, a more detailed description of an embodiment for implementing a scalable video encoder / decoder suitable for the integration of embodiments and combinations thereof is presented below. The different aspects described below are enumerated using alphabetic symbols. Descriptions of some of these aspects refer to elements in the drawings described below, and according to one embodiment, these aspects may be commonly implemented. However, it should be noted that regarding each aspect, the presence of all elements in the implementation of the scalable video decoder / encoder is not essential as long as all aspects are involved. Based on the aspect at issue, some elements and some interconnections may be excluded from the drawings described below. The elements cited for each aspect must exist only to perform the task or function mentioned in the description of each aspect; however, when some elements are mentioned in relation to a single function, alternatives may sometimes exist specifically.
[0257] However, to provide an overview of the functionality of a scalable video decoder / encoder in which the perspectives described below can be implemented, the elements shown in the figure below are now briefly explained.
[0259] FIG. 23 shows a scalable video decoder for decoding the coded data stream (6) into a video that is coded in such a way that an appropriate subpart of the coded data stream (6), namely 6a, represents the video at a first resolution or quality level, whereas an additional part (6b) of the coded data stream corresponds to a representation of the video at an increased resolution or quality level. To keep the amount of data in the coded data stream (6) low, inter-layer redundancies between the substreams (6a and 6b) are used to form the substream (6b). Some aspects described below aim at inter-layer prediction from the base layer to which the substream (6a) is associated to the enhancement layer to which the substream (6b) is associated.
[0261] The scalable video decoder includes two block-based prediction decoders (80, 60) that operate in parallel and receive substreams (6a) and (6b), respectively. As shown in the drawing, a demultiplexer (40) can provide their corresponding substreams (6a) and (6b) and decoding stages (80) and (60) separately.
[0263] The internal configuration of the block-based predictive coding stages (80) and (60) may be similar as shown in the drawing. From the input of each decoding stage (80, 60), the layer signal (600) and the restored enhancement layer signal (360) at the end of the serial connection are each induceable by connecting an entropy decoding module (100; 320), an inverse transformer (560; 580), an adder (adder, 180; 340), and optional filters (120; 300 and 140; 280) in series. On the other hand, the outputs of the adder (180, 340) and filters (120, 140, 300 and 280) provide different versions of base layer and enhancement layer signal recovery, respectively, and a particular prediction provider (160; 260) is provided, based thereon, to provide a prediction signal for the residual input of the adder (180; 340) and to receive all or a subset of these versions. Entropy decoding stages (100; 320) decode coding parameters and transformation coefficient blocks entering the inverse converter (560; 580), each containing prediction parameters for the prediction provider (160; 260), respectively, from the particular input signals (6a) and (6b).
[0265] In this way, the prediction providers (160) and (260) predict blocks of video frames at each resolution / quality level, and for this reason, the same can be selected from specific prediction modes such as spatial intra-prediction mode and temporal inter-prediction mode and intra-layer prediction mode, that is, the prediction modes depend only on the data of the substream entering each level.
[0267] However, to utilize the aforementioned inter-layer redundancy, the enhancement layer decoding stage (60) additionally includes a prediction provider (260) that additionally / alternatively supports inter-layer prediction modes capable of providing an enhancement layer prediction signal (420) based on data derived from the internal states of the base layer decoding stage (80) in comparison with a coding parameter inter-layer predictor (240), a resolution / quality improver (220), and / or a prediction provider (150). A resolution / quality refiner (220) is subject to either the base layer residual signal (480) or the restored base layer signals (200a, 200b, and 200c) for resolution or quality refinement to obtain an inter-layer prediction signal (380), and a coding parameter inter-layer predictor (240) is intended to somehow predict coding parameters such as prediction parameters and motion parameters, respectively. A prediction provider (260) may additionally support inter-layer prediction modes in which the restored portions of the base layer residual signal (640), which are potentially improved to an increased resolution / quality level, or the restored portions of the base layer signals, such as 200a, 200b, and 200c, are used as a reference / base.
[0269] As described above, the decoding stages (60) and (80) may operate in a block-based manner. This means that frames of the video may be subdivided into parts like blocks. Different granularity levels may be used to assign prediction modes performed by the prediction providers (160) and (260), local transformations by the inverse transformations (560) and (580), filter coefficient selections by the filters (120) and (140), and prediction parameters set for the prediction modes by the prediction providers (160) and (260). This means that the sub-partitioning of frames into prediction blocks may be a series of sub-partitionings of frames into blocks in which prediction modes are selected, for example, by coding units or prediction units. The sub-partitioning of frames into blocks for transformation coding can be called transformation units, and can vary from partitioning to prediction units. Some of the inter-layer prediction modes used by the prediction provider (260) are described below with respect to the aspects. The same applies to some intra-layer prediction modes, that is, for each prediction mode that internally derives each prediction signal input to the adder (180) and (340), that is, each based only on the states included in the current level coding stage (60) and (80).
[0271] Some further details regarding the blocks shown in the drawings will become clear from the descriptions of individual perspectives below. It should be noted that, unless such descriptions are specifically related to the perspectives provided, they may generally be translated as equivalent to the descriptions in the drawings and other perspectives.
[0273] In particular, the embodiment of the scalable video decoder in FIG. 23 represents a possible embodiment of the scalable video decoders according to FIG. 2 and 4. While the scalable video decoder according to FIG. 23 is described above, FIG. 23 shows a corresponding scalable video encoder, and the same reference numerals are used for internal elements of the predictive coding / decoding designs in FIG. 23 and 24. The reason for this is, as explained above: in order to maintain a common prediction basis between the encoder and the decoder, reconstructible versions of the base and enhancement layer signals are also used in the encoder, which, for this reason, reconstructs already coded parts to obtain a reconstructible version of the scalable video. Thus, the difference from the description in FIG. 23 is that the prediction provider (160) and the prediction provider (260), as well as the coding parameter inter-layer predictor (240), determine the prediction parameters within the process of some rate / distortion optimization rather than receiving the same from the data stream. Rather, the providers send the determined prediction parameters to entropy decoders (19a and 19b), which send the enhancement layer substream (6b) and each base layer substream (6a) in turn through the multiplexer (16) to be included in the data stream (6). In the same way, these entropy encoders (19a and 19b) receive the prediction residue between the enhancement layer versions (4a and 4b) and the original base layer and the restored enhancement layer signal (400) and the restored base layer signal (200), as obtained through a subtracter (720 and 722) followed by a conversion module (724, 726), rather than outputting the entropy decoding result of such residue.However, in addition to this, the configuration of the scalable video encoder in FIG. 24 is consistent with the configuration of the scalable video decoder in FIG. 23, and therefore regarding these issues, the above description of FIG. 23 is referenced, and as just described, parts referring to any induction from any data stream must be converted into each determination of each element with the next input to each data stream.
[0275] The intra-coding technique for enhancement layer signals used in the embodiments described below includes multiple methods for generating intra-predicted signals for enhancement layer blocks (using base layer data). These methods are additionally provided for generating intra-predicted signals based solely on restored enhancement layer samples.
[0277] Intra prediction is part of the restoration process for intra-coded blocks. The final restored block is obtained by adding a transform-coded residual signal (which can be zero) to the intra prediction signal. The residual signal is generated by the inverse quantization (scaling) of the transform coefficient levels transmitted in the bitstream followed by the inverse transform.
[0279] The following description applies to scalable coding with quality enhancement layers (an enhancement layer represents an input video having the same resolution as the base layer but with higher quality or fidelity) and scalable coding with spatial enhancement layers (an enhancement layer has a higher resolution than the base layer, i.e., larger samples). For quality enhancement layers, as in block (220), upsampling of base layer signals is not required, but filtering, such as (500), of the restored base layer samples may be applied. For spatial enhancement layers, as in block (220), upsampling of base layer signals is generally required.
[0281] The perspective described below supports different methods for using restored base layer samples (see (200)) or base layer residual samples (see (640)) for intra-prediction of enhancement layer blocks. In addition to intra-layer intra-coding, it is possible to support one or more of the methods described below. (Here, only restored enhancement layer samples (see (400)) are used for intra-prediction.) The use of a particular method may be signaled at the level of the largest supported block size (such as macroblocks in H.264 / AVC or coding tree blocks / maximum coding units in HEVC), or at all supported block sizes, or at a subset of supported block sizes.
[0283] For all methods described below, the prediction signal can be used directly as a restoration signal for the block, i.e., no residual is transmitted. Alternatively, the method selected for inter-layer intra-prediction can be combined with residual coding. In certain embodiments, the residual signal is transmitted via transform coding, that is, the quantized transform coefficients (transform coefficient levels) are transmitted using an entropy coding technique (e.g., variable length coding or arithmetic coding (see (19b))), and the residual is obtained by inversely quantizing (scaling) the transmitted transform coefficient levels and applying an inverse transform (see (580)). In a specific version, the complete residual block corresponding to the block where the inter-layer intra-prediction signal is generated is transformed using a single transform (see (726)) (i.e., the entire block is transformed using a single transform of the same size as the prediction block). In another embodiment, the prediction block may be further subdivided into smaller blocks (e.g., using hierarchical decomposition), and an individual transform is applied to each of the smaller blocks (which may have different block sizes). In a further embodiment, the coding unit may be subdivided into smaller prediction blocks, and a prediction signal is generated for zero or more of the prediction blocks using one of the methods for inter-layer intra-prediction. Then, the residual of the entire coding unit is transformed using a single transform (see (726)), or the coding unit is subdivided into different transform units, wherein the subdivision to form the transform units (blocks to which the single transform is applied) is different from the subdivision to decompose the coding unit into prediction blocks.
[0285] In certain embodiments, the (upsampled / filtered) reconstructed base layer signal (see 380) is used directly as a prediction signal. Various methods for using the base layer to intra-predict an enhancement layer include the following: The (upsampled / filtered) reconstructed base layer signal (see 380) is used directly as an enhancement layer prediction signal. This method is similar to the known H.264 / SVC inter-layer intra-prediction mode. In this method, the prediction block for the enhancement layer is formed by coexisting samples of the base layer reconstructed signal (see 220) that may have been upsampled to match corresponding sample locations of the enhancement layer and may have been optionally filtered before or after upsampling. In contrast to the SVC inter-layer intra-prediction mode, this mode is not supported only at the macro level (or at the maximum supported block size), but may be supported at any block size. This means that the above mode cannot be signaled only for the maximum supported block size, and blocks of the maximum supported block size (macroblocks in MPEG-4, coding tree blocks / maximum coding units in H.264 and HEVC) can be hierarchically subdivided into smaller blocks / coding units, and the use of the inter-layer intra-prediction mode can be signaled at any supported block size (for the corresponding block). In certain embodiments, this mode is supported only for selected block sizes. Subsequently, the syntax element signaling the use of this mode can be transmitted only for the corresponding block sizes, and the values of the syntax element signaling the use of this mode (among other coding parameters) can be correspondingly restricted for the other block sizes. H.Another difference regarding the inter-layer intra prediction mode in the 264 / AVC SVC extension is that the inter-layer intra prediction mode is supported not only when the coexisting regions of the base layer are intra-coded, but also when the coexisting regions of the base layer are inter-coded or partially inter-coded.
[0287] In a specific embodiment, spatial intra-prediction of the difference signal (see Aspect A) is performed. A number of methods include the following: a (potentially upsampled / filtered) restored base layer signal (see 380) is combined with a spatial intra-prediction signal, wherein the spatial intra-prediction (see 420) is derived based on different samples for adjacent blocks (see 260). The different samples represent the difference of the restored enhancement layer signal (see 400) and the (potentially upsampled / filtered) restored base layer signal (see 380).
[0289] FIG. 25 is an inter-layer prediction method that uses a difference signal (734) (EH Diff) of already coded adjacent blocks (736) and a sum of a (upsampled / filtered) base layer signal (380) (BL Reco), wherein the difference signal (EH Diff) for already coded blocks (736) is generated by subtracting (738) the (upsampled / filtered) base layer restoration signal (380) (BL Reco) from the restored enhancement layer signal (EH Reco) (see 400) in which the already coded / decoded parts are hatched and shown, and the currently coded / decoded block / region / part is (28). This means that the inter-layer intra prediction method illustrated in FIG. 25 uses two superimposed input signals to generate a prediction block. For this method, a difference signal (734) is required, which is the difference between a coexisting restored base layer signal (200) and a restored enhancement-layer signal (400) that can be optionally filtered before or after upsampling and upsampled to match corresponding sample locations of the enhancement layer (220) (it may also be filtered when upsampling is not applied, as in the case of quality scalable coding). In particular, for spatial scalable coding, the difference signal (734) generally contains mainly high-frequency components. The difference signal (734) is available for all already restored blocks (i.e., already coded / decoded for all enhancement layer blocks). The difference signal (734) for adjacent samples (742) of the already coded / decoded blocks (736) is used as input to a spatial intra prediction technique (such as spatial intra prediction modes specified in H.264 / AVC or HEVC). A prediction signal (746) for the difference component of the block (28) predicted by spatial intra prediction illustrated by the arrow (744) is generated. In a specific embodiment, (H.Any clipping function of the spatial intra prediction process (as known from 264 / AVC or HEVC) is modified or disabled to match the dynamic range of the difference signal (734). The intra prediction method actually used (which may be one of a number of provided methods and may include planar intra prediction, DC prediction, or directional intra prediction (744) with any specific angle) is signaled within the bitstream (6b). It is possible to use spatial intra prediction techniques different from the methods provided in H.264 / AVC and HEVC (a method of generating a prediction signal using samples of already coded adjacent blocks). The prediction block (746) obtained (using the difference samples of adjacent blocks) is the first portion of the final prediction block (420).
[0291] A second portion of the prediction signal is generated using a region (28) coexisting with the restored signal (200) of the base layer. For quality enhancement layers, the coexisting base layer samples can be used directly or they can be optionally filtered, for example, by a low-pass filter or a filter (500) that attenuates high-frequency components. For spatial enhancement layers, the coexisting base layer samples are unsampled. For upsampling (220), an FIR filter or a set of FIR filters can be used. It is also possible to use IIR filters. Optionally, the restored base layer samples (200) can be filtered before upsampling, or the base layer prediction signal (the signal obtained after upsampling the base layer) can be filtered after the upsampling stage. The base layer restoration process may include one or more additional filters, such as a deblocking filter (see 120) and an adaptive loop filter (see 140). The base layer recovery (200) used for upsampling may be a recovery signal before any of the loop filters (see 200c), or a recovery signal after a deblocking filter but before any additional filter (see 200b), or a recovery signal after a specific filter or a recovery signal after applying all filters used in the base layer decoding process (see 200a).
[0293] The two generated parts of the prediction signal (spatially predicted difference signal (746) and potentially filtered / upsampled base layer restoration (380)) are added sample by sample to form the final prediction signal (420) (732).
[0295] To convey the perspective just described with respect to the embodiments of FIGS. 6 through 10 is that the just described possibility of predicting the current block of the enhancement layer signal by each scalable video decoder / encoder as an alternative to the prediction design described with respect to FIGS. 6 through 10 may be supported. The mode used therein is signaled in the enhancement layer substream (6b) through each prediction mode identifier not shown in FIG. 9.
[0297] In a specific embodiment, intra prediction is followed by inter-layer residual prediction (see Aspect B). A number of methods for generating an intra prediction signal using base layer data include the following: a conventional spatial intra prediction signal (derived using adjacent restored enhancement layer samples) is combined with a base layer residual signal (upsampled / filtered) (the difference between base layer restoration and base layer prediction or the inverse transformation of base layer transformation coefficients).
[0299] FIG. 26 illustrates such occurrence of an inter-layer intra prediction signal (420) by the sum (752) of a (upsampled / filtered) base layer residual signal (754) (BL Resi) and a spatial intra prediction (756) using restored enhancement layer samples (758) (EH Reco) of already coded adjacent blocks, illustrated by dotted lines (762).
[0301] The concept shown in FIG. 26 is to superimpose two prediction signals to form a prediction block (420), one prediction signal (764) is generated from already restored enhancement layer samples (758) and the other prediction signal (754) is generated from base layer residual samples (480). The first part (764) of the prediction signal (420) is derived by applying spatial intra prediction (756) using the restored enhancement layer samples (758). Spatial intra prediction (756) may be one of the methods specified in H.264 / AVC or one of the methods specified in HEVC, or it may be another spatial intra prediction technique in which the prediction signal (764) generated for the current block (18) forms samples (758) of adjacent blocks (762). An intra prediction method (756) actually used (which may be one of a number of provided methods and may include planar intra prediction, DC intra prediction, or directional intra prediction with any specific angle) is signaled within the bitstream (6b). It is possible to use spatial intra prediction techniques different from the methods provided in H.264 / AVC and HEVC (methods that generate a prediction signal using samples of already coded adjacent blocks). A second part (754) of the prediction signal (420) is generated using the coexisting residual signal (480) of the base layer. For quality enhancement layers, the residual signal may be used as restored from the base layer and may be additionally filtered. For the spatial enhancement layer (480), the residual signal is upsampled (220) before being used as the second part of the prediction signal (to map base layer sample locations to enhancement layer sample locations). The base layer residual signal (480) may be filtered before or after the upsampling stage. For upsampling (220), residual signals and FIR filters may be applied.The upsampling process can be configured in such a way that filtering across transform block boundaries in the base layer is not applied for upsampling purposes.
[0303] The base layer residual signal (480) used for inter-layer prediction may be a residual signal obtained by inversely transforming (560) and scaling the transform coefficient levels of the base layer, or it may be the difference between the prediction signal (660) used in the base layer and the base layer signal (200) restored (before or after any filtering operations or additional filtering and deblocking).
[0305] Two generated signal components (spatial intra-prediction signal (764) and inter-layer residual prediction signal (754)) are added together to form a final enhancement layer intra-prediction signal (752).
[0307] This means that the prediction mode just described with respect to FIG. 26 can be used or supported by any scalable video decoder / encoder according to FIG. 6 to 10 to form an alternative prediction mode described above with respect to FIG. 6 to 10 with respect to the currently coded / decoded part (28).
[0309] In certain embodiments, weighted prediction of base layer restoration and spatial intra prediction (see Aspect C) is used. This substantially represents the mentioned details of specific implementations of the embodiments described above with respect to FIGS. 6 through 10, and thus the description of such weighted prediction is interpreted not only as an alternative to the above embodiments, but also as a description of the possibility of how to implement the embodiments described with respect to FIGS. 6 through 10 differently from certain aspects.
[0311] Various methods for generating an intra prediction signal using base layer data include the following: the (upsampled / filtered) reconstructed base layer signal is combined with a spatial intra prediction signal, and the spatial intra prediction is derived based on the reconstructed enhancement layer samples of adjacent blocks. The final prediction signal is obtained by weighting the base layer prediction signal and the spatial prediction signal (see 41) in such a way that different frequency components use different weights. This can be realized, for example, by filtering the base layer prediction signal with a low-pass filter (see 38) (see 62) and filtering the spatial intra prediction signal (see 34) with a high-pass filter, and adding the obtained filtered signals (see 66). Alternatively, frequency-based weighting can be realized by transforming (see 72, 74) the base layer prediction signal (see 38) and the enhancement layer prediction signal (see 34) and superimposing the obtained transform blocks (see 76, 78) in which different weighting factors (see 82, 84) are used for different frequency positions. The obtained transform block (see 42 in FIG. 10) can be inversely transformed (see 83) and used as the enhancement layer prediction signal (see 54), or the obtained transform factors are added to scaled transmitted transform factor levels (see 59) (see 52) and inversely transformed (see 84) to obtain a restored block (see 54) before deblocking and in-loop processing.
[0313] FIG. 27 shows such generation of an inter-layer intra prediction signal by a frequency-weighted sum of a (upsampled / filtered) base layer reconstruction signal (BL Reco) and a spatial intra prediction using restored enhancement layer samples (EH Reco) of already coded adjacent blocks.
[0315] The concept of FIG. 27 utilizes two superimposed signals (772, 774) to form a prediction block (420). The first part (774) of the signal (420) is utilized by applying a spatial intra-prediction (776) corresponding to (30) in FIG. 6, utilizing restored samples (778) of an adjacent block already configured in the enhancement layer. The second part (772) of the prediction signal (420) is generated using the coexisting restored signal (200) of the base layer. For quality enhancement layers, the coexisting base layer samples (200) can be used directly and can be optionally filtered, for example, by a filter or low-pass filter that attenuates high-frequency components. For spatial enhancement layers, the coexisting base layer samples are upsampled (220). For upsampling, an FIR filter or a set of FIR filters may be used. It is also possible to use IIR filters. Optionally, the reconstructed base layer samples may be filtered before upsampling, and the base layer prediction signal (the signal obtained before upsampling the base layer) may be filtered after the upsampling stage. The base layer reconstruction process may include one or more additional filters, such as a deblocking filter (120) and an adaptive loop filter (140). The base layer reconstruction (200) used for upsampling may be a reconstructed signal (200c) before any loop filters (120, 140), or a reconstructed signal (200b) after the deblocking filter (120) but before any additional filter, or a reconstructed signal (200a) after applying all filters (120, 140) used in the base layer decoding process or a reconstructed signal (200a) after a specific filter.
[0317] When comparing the reference numerals used in FIGS. 23 and 24 with those used in connection with FIGS. 6 to 10, block (220) corresponds to the reference numeral (38) used in FIG. 6, (39) corresponds to the part of (380) corresponding to the current part (28), (420) corresponds to (42), and spatial prediction (776) corresponds to (32) as long as the part coexisting with the current part (28) is involved.
[0319] Two prediction signals (potentially upsampled / filtered base layer restoration (386) and enhancement layer intra prediction (782)) are combined to form a final prediction signal (420). The method for combining these signals may have different weighting factors used for different frequency components. In a particular embodiment, the upsampled base layer restoration is filtered by a low-pass filter (see (62)) (it is also possible to filter the base layer restoration before upsampling (220)) and the intra prediction signal (obtained by (30) see (34)) is filtered by a high-pass filter (see (64)), and both filtered signals are added (784) (see (66)) to form the final prediction signal (420). The pair of low-pass and high-pass filters may represent a quadrature mirror filter pair, but this is not required.
[0321] In another specific embodiment (see FIG. 10), the process of combining two prediction signals (380) and (782) is realized through spatial transformation. Both (potentially upsampled / filtered) base layer restoration (380) and intra prediction signal (782) are transformed using spatial transformation (see (72), (74)). The transformation coefficients of both signals (see (76), (78)) are scaled by appropriate weighting factors (see (82), (84)) and then added to form a transformation coefficient block (see (42)) of the final prediction signal (see (90)). In one version, the weighting factors (see (82, 84)) are selected such that for each transformation coefficient position, the sum of the weighting factors for both signal components is equal to 1. In another version, the sum of the weighting factors may not be 1 for some or all transformation coefficient positions. In a specific version, weighting factors are selected such that, for transform factors representing low-frequency components, the weighting factor for base layer restoration is greater than the weighting factor for enhancement layer intra-prediction signal, and for transform factors representing high-frequency components, the weighting factor for base layer restoration is smaller than the weighting factor for enhancement layer intra-prediction signal.
[0323] In one embodiment, the obtained transformation coefficient block (see (42)), obtained by summing the weighted transformation signals for both components, is inversely transformed (see 84) to form the final predicted signal (420) (see 54). In another embodiment, the prediction is performed directly in the transformation area. This is so that the coded transformation coefficient levels (see 59) are added (see 52) to the transformation coefficients (see 42) of the predicted signal (obtained by summing the weighted transformation signals for both components), scaled (i.e., inversely quantized), and the resulting block of transformation coefficients (not shown in FIG. 10) is inversely transformed (see 84) to obtain the restored signal (420) for the current block (before additional in-loop filtering steps (140) and potential de-blocking (120). In other words, in the first embodiment, the transformation block obtained by summing the transformation signals weighted for both components can be used as enhancement layer prediction signals and can be inversely transformed, and in the second embodiment, the obtained transformation coefficients can be scaled and added to the transmitted transformation coefficient levels and can be inversely transformed to obtain the restored block before deblocking and in-loop processing.
[0325] Base layer restoration and selection of residual signals (See Perspective D) may also be used. Regarding the method of using the restored base layer signal (as described above), the following versions may be used:
[0327] · Restored base layer samples (200c) before deblocking (120) and additional in-loop processing (140) (like a sample adaptive offset filter or adaptive loop filter).
[0328] · Restored base layer samples (200b) after deblocking (120) but before additional in-loop processing (140) (like a sample adaptive offset filter or adaptive loop filter).
[0329] · Restored base layer samples (200a) after deblocking (120) and additional in-loop processing (140) (such as a sample adaptive offset filter or adaptive loop filter) or between multiple in-loop processing steps.
[0331] The selection of the corresponding base layer signals (200a, b, c) may be fixed for a specific decoder (and encoder) execution, or signaled into the bitstream (6). Different versions may be used for the following cases. The use of a specific version of the base layer signal may be signaled at the sequence level, or at the picture level, or at the slice level, or at the maximum coding unit level, or at the coding unit level, or at the prediction block level, or at the transformation block level, or at any other block level. In another version, the selection may depend on other coding parameters (like coding modes) or characteristics of the base layer signal.
[0333] In another embodiment, multiple versions of methods for utilizing the (upsampled / filtered) base layer signal (200) may be utilized. For example, two different modes may be provided that directly utilize the upsampled base layer signal, i.e., (200a), and said two modes may utilize different interpolation filters, or one mode may utilize additional filtering (500) of the (upsampled) base layer restored signal. Similarly, multiple different versions may be provided for the other modes described above. The upsampled / filtered base layer signal (380) utilized for different versions of the mode may differ in the interpolation filters utilized (including interpolation filters that also filter integer-sample locations), or the upsampled / filtered base layer signal (380) for the second version may be obtained by filtering (500) the upsampled / filtered base layer signal for the first version. The selection of one of the different versions can be signaled at the sequence, picture, slide, maximum coding unit, coding unit level, prediction block level, or transformation block level, or can be inferred from the characteristics of the corresponding restored base layer signal or the transmitted coding parameters.
[0335] The same thing as above can be applied to modes using the base layer residual signal restored through (480). Here, different versions that differ in interpolation filters or additional filtering steps may also be used.
[0337] Different filters can be used to upsample / filter the base layer residual signal and the reconstructed base layer signal. This means that a different approach is used for upsampling the base layer residual signal compared to upsampling the base layer reconstructed signal.
[0339] For base layer blocks where the residual signal is zero (i.e., when transform factor levels are not transmitted for the block), the corresponding base layer residual signal may be replaced with another signal derived from the base layer. This may be, for example, a high-pass filtered version of the base layer block being restored, any other difference-like signal derived from the restored base layer samples, or the restored base layer residual samples of adjacent blocks.
[0341] As long as samples are used for spatial intra prediction in the enhancement layer (see Aspect H), the following special treatments may be provided. For modes using spatial intra prediction, (since adjacent samples may be unavailable, and adjacent blocks may be coded after the current block), unavailable adjacent samples in the enhancement layer may be replaced with corresponding samples of the upsampled / filtered base layer signal.
[0343] Insofar as the coding of intra prediction modes (see Perspective X) is relevant, the following special modes and functionalities (functions) may be provided. (30a) For modes using spatial intra prediction, the coding of the intra prediction mode may be modified in such a way that information about the intra prediction mode in the base layer (if available) is used more efficiently to code the intra prediction mode in the enhancement layer. This may be, for example, parameters (56). If a coexisting region (see 36) is intra-coded using a specific spatial intra prediction mode in the base layer, a similar intra prediction mode may also be used in the enhancement layer block (see 28). Intra prediction modes are generally signaled in such a way that one or more modes among a set of possible intra prediction modes, which can be signaled with shorter code words (or with fewer arithmetic code binary decision results in fewer bits), are classified as the most common modes. In HEVC intra-prediction, the intra-prediction mode of the block to the top (if available) and the intra-prediction mode of the block to the left (if available) are included in the set of most likely modes. In addition to these modes, one or more additional modes (often used) are included in the list of most likely modes, and the actual additional modes depend on the availability of the intra-prediction modes of the block to the left of the current block and the block above the current block. In H.264 / AVC, one mode is classified as the most likely mode, and this mode is derived based on the intra-prediction modes available for the block to the left of the current block and the block above the current block. Any other concept (different from H.264 / AVC and HEVC) is possible to classify intra-prediction modes and can be used for the following extension.
[0345] To utilize base layer data for the efficient coding of intra-prediction modes in the enhancement layer, the concept of utilizing one or more most probable modes is modified in such a way that the most probable mode includes the intra-prediction mode utilized in the base layer block where the most probable mode coexists. In a specific embodiment, the following approach is used: for a given current enhancement layer block, a coexisting base layer block is determined. In a specific version, the coexisting base layer block is a base layer block covering the coexisting location of the top-left sample of the enhancement block. In another version, the coexisting base layer block is a base layer block covering the coexisting location of the sample in the middle of the enhancement block. In other versions, other samples within the enhancement layer block may be used to determine the coexisting base layer block. If the determined coexisting base layer block is intra-coded and the base layer intra-predicted mode specifies each intra-predicted mode, and the intra-predicted mode derived from the enhancement layer block to the left of the current enhancement layer block does not utilize each intra-predicted mode, the intra-predicted mode derived from the left enhancement layer block is replaced with the corresponding base layer intra-predicted mode. Otherwise, if the determined coexisting base layer block is intra-coded and the base layer intra-predicted mode specifies each intra-predicted mode, and the intra-predicted mode derived from the enhancement layer block above the current enhancement layer block does not utilize each intra-predicted mode, the intra-predicted mode derived from above the enhancement layer block is replaced with the corresponding base layer intra-predicted mode. In other versions, a different approach is used to modify the list of most probable modes (which can be composed of a single element) using the base layer intra-predicted mode.
[0347] Intercoding techniques for spatial and quality enhancement layers are proposed below.
[0349] In modern hybrid video coding standards (such as H.264 / AVC or the emerging HEVC), pictures in a video sequence are divided into blocks of samples. The block size can be fixed, or the coding approach may provide a hierarchical structure that allows the blocks to be further subdivided into blocks with smaller block sizes. Reconstruction of a block is generally achieved by adding the transmitted residual signal and generating a prediction signal for the block. The residual signal is typically transmitted using transform coding, which means that quantization indices for transform coefficients (also referred to as transform coefficient levels) are transmitted using entropy coding techniques; from the decoder's perspective, these transmitted transform coefficient levels are scaled and inversely transformed to obtain the residual signal that is added to the prediction signal. The residual signal is generated by intra prediction (using only data already transmitted for the current time instance) or inter prediction (using data already transmitted for different time instances), respectively.
[0351] In inter-prediction, prediction blocks are derived by motion-compensated prediction using samples of already restored frames. This can be performed on unidirectional prediction (using a single set of motion parameters and a single reference picture), or the prediction signal can be generated by multi-hypothesis prediction. In the following cases, two or more prediction signals overlap, that is, for each sample, a weighted average is configured to form the final prediction signal. The multiple prediction signals (overlapping) can be generated using different motion parameters for different hypotheses (e.g., different reference pictures or motion vectors). For unidirectional prediction, it is also possible to multiply samples of motion-compensated prediction signals with a constant factor and add a constant offset to form the final prediction signal. Such scaling and offset correction may also be used for all or selected hypotheses in multi-hypothesis prediction.
[0353] In scalable video coding, base layer information can also be utilized to support the inter-prediction process for the enhancement layer. In the latest video coding standards for scalable coding, the SVC extension of H.264 / AVC includes an additional mode to improve the coding efficiency of the inter-prediction process in the enhancement layer. This mode is signaled at the macroblock level (blocks of 16x16 luminance samples). In this mode, remnants restored from the lower layer are used to enhance the motion-compensated prediction signal in the enhancement layer. This mode is also referred to as inter-layer remnants prediction. If this mode is selected for a macroblock in the quality enhancement layer, the inter-layer prediction signal is constructed by coexisting samples of the restored lower-layer remnants. If the inter-layer remnants prediction mode is selected in the spatial enhancement layer, the prediction signal is generated by upsampling the coexisting restored base layer remnants. For upsampling, FIR filters are used, but no filtering is applied beyond the transform block boundaries. The prediction signal generated from the restored base layer residual samples is added to the conventional motion-compensated prediction signal to form the final prediction signal for the enhancement layer block. Generally, for the inter-layer residual prediction mode, the additional residual signal is transmitted via transform coding. The transmission of the residual signal may be omitted if it is correspondingly signaled within the bitstream (referred to as being equivalent to 0). The final restored signal is obtained by adding the restored residual signal (obtained by applying an inverse spatial transformation and scaling the transmitted transform coefficient levels) to the prediction signal (obtained by adding the inter-layer residual prediction signal to the motion-compensated prediction signal).
[0355] Next, techniques for inter-coding enhancement layer signals are described. This section describes methods for using base layer signals in addition to already restored enhancement layer signals to inter-predict enhancement layer signals to be coded in scalable video coding scenarios. By using base layer signals to inter-predict coded enhancement layer signals, prediction errors can be significantly reduced, which leads to overall bitrate savings in coding the enhancement layer. The main focus of this section is to discuss blocks based on motion compensation of enhancement layer samples using already coded enhancement layer samples that have additional signals from the base layer. The following description provides the possibility of using various signals from the coded base layer. Although quad-tree block partitioning is generally used in preferred embodiments, the proposed examples are applicable to general block-based hybrid coding approaches without assuming any specific block partitioning. For the inter-prediction of coded enhancement layer blocks, the use of base layer restoration of already coded pictures, or base layer residuals of the current time index, and base layer restoration of the current time index are described. Furthermore, it is explained how base layer signals can be combined with already coded enhancement layer signals to obtain better predictions for the current enhancement layer. One of the key modern techniques is inter-layer residual prediction in H.264 / SVC. In H.264 / SVC, inter-layer residual prediction can be utilized for all inter-coded macroblocks, regardless of whether they are coded using any of the conventional macroblock types or SVC macroblock types signaled by the base mode flag. The flag is added to the macroblock syntax for spatial and quality enhancement layers, signaling the utilization of inter-layer residual prediction.When the residual prediction flag is equal to 1, the residual signal of the corresponding region in the reference layer is upsampled block-by-block using a bilinear filter and used as a prediction for the residual signal of the enhancement layer macroblock, so that only the corresponding difference signal needs to be coded in the enhancement layer. For the description in this section, the following symbols are used:
[0356] t0 := time index of the current picture
[0357] t1 := time index of an already reconstructed picture
[0358] EL := enhancement layer
[0359] BL := base layer
[0360] EL(t0) := current enhancement layer picture to be coded
[0361] EL_reco := enhancement layer reconstruction
[0362] BL_reco := base layer reconstruction
[0363] BL_resi := base layer residual signal (inverse transform of base layer transform coefficients or difference between base layer restoration and base layer prediction)
[0364] EL_diff := Difference between restoration of enhancement layer and restoration of upsampled / filtered base layer
[0365] The difference base layer and enhancement layer signals used in the description are illustrated in FIG. 28.
[0367] For the explanation, the following characteristics of the filters are used:
[0368] · Linearity: Most of the filters mentioned in the above description are linear, but non-linear filters can also be used.
[0369] · Number of output samples: In a sampling operation, the number of output samples is greater than the number of input samples. Here, filtering of the input data generates more samples than the input values. In conventional filtering, the number of output samples is equal to the number of input samples. Such filtering operations can be used, for example, in quality scalable coding.
[0370] · Phase delay: For filtering samples at integer positions, the phase delay is typically zero (or is an integer-value delay in the samples). For samples occurring at fractional positions (e.g., half-pel or quarter-pel positions), filters that generally have a fractional delay (in a unit of samples) can be applied to samples in an integer grid.
[0372] Conventional motion-compensated prediction (i.e., MPEG-2, H.264 / AVC, or the upcoming HEVC standard), as used in all hybrid video coding standards, is illustrated in FIG. 29. To predict the signal of the current block, an area of the already restored picture is used as the prediction signal and is replaced. To signal the displacement, the motion vector is typically coded within the bitstream. However, it is also possible to transmit fractional-sample precision motion vectors. In this case, the prediction signal is obtained by filtering the reference signal with a filter having a fractional sample delay. The reference picture used can generally be identified by including the reference picture index in the bitstream syntax. Generally, it is also possible to superimpose two or more prediction signals to form the final prediction signal. This concept is supported, for example, in B-slices with two motion hypotheses. In this case, multiple prediction signals are generated using different motion parameters (e.g., different reference pictures or motion vectors) for different hypotheses. For unidirectional prediction, it is also possible to multiply samples of the motion-compensated prediction signal by a constant factor and add a constant offset to form the final prediction signal. Such scaling and offset correction may also be used for all or selected hypotheses in multi-hypothesis prediction.
[0374] The following description applies to scalable coding with quality enhancement layers (where the enhancement layer represents input audio with the same resolution as the base layer but with higher quality or fidelity) and scalable coding with spatial enhancement layers (where the enhancement layer has a higher resolution than the base layer, i.e., has a larger number of samples). For quality enhancement layers, upsampling of the base layer signals is not required, but filtering of the reconstructed base layer samples may be applied. In the case of spatial enhancement layers, upsampling of the base layer signals is generally required.
[0376] The embodiments support different methods for using base layer residual samples or restored base layer samples for inter-prediction of enhancement layer blocks. It is possible to support one or more of the methods described below in additional conventional inter-prediction and intra-prediction. The utilization of a specific method (such as a coding tree block / maximum coding unit in HEVC or a macroblock in H.264 / AVC) may be signaled at the level of the maximum supported block size, or at all supported block sizes, or at a subset of supported block sizes.
[0378] For all methods described below, the prediction signal can be used directly as a restoration signal for the block. Alternatively, the selected method for inter-layer inter-prediction can be combined with residual coding. In certain embodiments, the residual signal is transmitted via transform coding, that is, quantized transform coefficients (transform coefficient levels) are transmitted using entropy coding techniques (e.g., variable-length coding or arithmetic coding), and the residual is obtained by inversely quantizing (scaling) the transmitted transform coefficient levels and applying an inverse transform. In certain versions, the complete residual block corresponding to the block where the inter-layer inter-prediction signal is generated is transformed using a single transform (i.e., the entire block is transformed using a single transform of the same size as the prediction block). In another embodiment, the prediction block may be further subdivided into smaller blocks (e.g., using hierarchical decomposition), and individual transforms are applied to each of the smaller blocks (which may have different block sizes). In addition, the coding unit may be divided into smaller prediction blocks, and for zero or more of the prediction blocks, the prediction signal is generated using one of the methods for inter-layer inter-prediction. Subsequently, the remainder of the entire coding unit is transformed using a single transformation, or the coding unit is subdivided into different transformation units, and the subdivision to form the transformation units (blocks to which a single transformation is applied) is different from the subdivision for the decomposition of the coding unit into prediction blocks.
[0380] In the following, possibilities for performing prediction using base layer residuals and enhancement layer restoration are described. Many methods include the following: a conventional inter-prediction signal (derived by motion-compensated interpolation of already restored enhancement layer pictures) is combined with an upsampled / filtered base layer residual signal (the difference between base layer prediction and base layer restoration or the inverse transformation of base layer transform coefficients). This method is also referred to as the BL_resi mode (see Fig. 30).
[0382] Simply put, the prediction for the enhancement layer samples can be written as follows:
[0383] EL prediction = filter( BL_resi(t 0 ) ) + MCP_filter( EL_reco(t 1 ) ).
[0384] It is possible to use two or more hypotheses of enhancement layer restoration signals, for example as follows:
[0385] EL prediction = filter( BL_resi(t 0 ) ) + MCP_filter1( EL_reco(t 1 ) ) + MCP_filter2( EL_reco(t2) ).
[0386] Motion-compensated prediction (MCP) filters used in enhancement layer (EL) reference pictures can have integer or fractional sample accuracy. The MCP filters used in EL reference pictures may be the same or different from the MCP filters used in BL reference pictures during the BL decoding process.
[0387] The motion vector MV(x,y,t) can be defined to indicate a specific location within the EL reference picture. The parameters x and y indicate spatial locations within the picture, while the parameter t is used to handle the temporal index of the reference pictures, also known as the reference index. Often, the term motion vector is used to refer only to the two spatial components (x,y). The integer part of MV is used to fetch a set of samples from the reference picture, and the fractional part of MV is used to select an MCP filter from a set of filters. The fetched reference samples are filtered to generate filtered reference samples. Motion vectors are typically coded using differential prediction. This means that the motion vector predictor is derived based on already coded motion vectors (and syntactic elements potentially indicating the use of one of the set of potential motion vector predictors), and the difference vector is included in the bitstream. The final motion vector is obtained by adding the motion vector differences transmitted to the motion vector predictor. Generally, it is also possible to fully derive motion parameters into a block. Therefore, a list of potential motion parameter candidates is typically constructed based on already coded data. This list may include motion parameters derived based on the motion parameters of blocks coexisting in the reference frame, as well as motion parameters of spatially adjacent blocks. The base layer (BL) residual signal can be defined according to one of the following:
[0388] · Inverse transformation of BL transformation coefficients, or
[0389] · The difference between BL restoration and BL prediction, or
[0390] · For BL blocks where the inverse transform of the BL transform coefficients is 0, they can be replaced with another signal derived from the BL, for example, a high-pass filtered version of the restored BL block, or
[0391] · Combination of the above methods.
[0392] To compute EL prediction components from the current BL residue, regions of the BL picture coexisting with the considered regions in the EL picture are identified, and residual signals are taken from the identified BL regions. The definition of the coexisting regions can be made to describe the generation of the same EL resolution according to an integer scaling factor of the BL resolution (e.g., 2x scalability), a fractional scaling factor of the BL resolution (e.g., 1.5x scalability), or a BL resolution (e.g., quality scalability). In the case of quality scalability, the blocks coexisting in the BL picture have the same coordinates as the predicted EL blocks.
[0393] Coexisting BL residues can be upsampled / filtered to generate filtered BL residue samples.
[0394] The final EL prediction is obtained by adding the filtered BL residual samples and the filtered EL reconstructed samples.
[0396] A number of methods related to prediction using base layer restoration and enhancement layer difference signals include the following methods (see Aspect J). The (upsampled / filtered) restored base layer signal is combined with the motion-compensated prediction signal, and the motion-compensated prediction signal is obtained by motion-compensated difference pictures. The difference pictures represent the difference between the (upsampled / filtered) restored base layer signal and the restored enhancement layer signal with respect to reference pictures. This method is also referred to as the BL_reco mode.
[0398] This concept is illustrated in Fig. 31. Simply put, the prediction for the EL samples can be described as follows:
[0399] EL prediction = filter( BL_reco(t 0 ) ) + MCP_filter( EL_diff(t 1 ) ).
[0401] It is also possible to use two or more hypotheses of EL difference signals, for example,
[0402] EL prediction = filter( BL_resi(t 0 ) ) + MCP_filter1( EL_diff(t 1 ) ) + It is MCP_filter2( EL_diff(t2) ) .
[0404] For EL difference signals, the following versions can be used:
[0405] · The difference between EL restoration and upsampled / filtered BL restoration, or
[0406] · Difference between previous EL restorations or loop filtering stages (such as deblocking, SAO, ALF) and upsampled / filtered BL restorations.
[0408] The use of a specific version can be fixed at the decoder or signaled at the sequence level, picture level, slice level, max coding unit level, coding unit level, or other partitioning levels. Alternatively, it can be created depending on other coding parameters.
[0410] When the EL difference signal is defined using the difference between the upsampled / filtered BL reconstruction and the EL reconstruction, it can be compromised to simply save the BL reconstruction and EL reconstruction and calculate the EL difference signal on-the-fly for the blocks using a prediction mode, thereby saving the memory required to store the EL difference signal. However, this introduces some computational complexity overhead.
[0412] The MCP filters used in EL difference pictures can be integer or fractional sample accuracy.
[0413] · For the MCP of the difference pictures, interpolation filters different from the MCP of the restored pictures may be used.
[0414] · For the MCP of difference pictures, interpolation filters can be selected based on the characteristics of corresponding regions in the difference pictures (or based on coding parameters or based on information transmitted in the bitstream).
[0416] The motion vector MV(x,y,t) is defined to point to a specific location in the EL difference picture. The parameters x and y point to spatial locations within the picture, and the parameter t is used to handle the time index of the difference picture.
[0418] The integer part of MV is used to fetch a set of sampled difference pictures, and the fractional part of MV is used to select an MCP filter from a set of filters. The fetched difference samples are filtered to generate filtered difference samples.
[0420] The dynamic range of the difference pictures can theoretically exceed the dynamic range of the original picture. Assuming an 8-bit representation of the images in range
[0255] , the difference images can have a range of [-255 255]. However, in practice, most amplitudes are distributed near positive or negative values near 0. In a preferred embodiment for storing the difference images, a constant offset of 128 is added, and the result is clipped to the range
[0255] and stored as regular 8-bit images. Later, during the encoding and decoding process, the offset of 128 is subtracted from the difference amplitude loaded from the difference pictures.
[0422] Regarding the method of using the restored BL signal, the following versions may be used. It can be fixed or signaled at the sequence level, picture level, slice level, max coding unit level, coding unit level, or other partitioning levels. Or it can be created depending on other coding parameters.
[0423] · Restored base layer samples before deblocking and additional in-loop processing (such as sample adaptive offset filters or adaptive loop filters)
[0424] · Restored base layer samples after deblocking, but before additional in-loop processing (such as sample adaptive offset filters or adaptive loop filters)
[0425] · Restored base layer samples between multiple in-loop processing steps or after additional in-loop processing and deblocking (such as sample adaptive offset filters or adaptive loop filters)
[0427] To calculate the EL prediction component from the current BL reconstruction, regions of the BL picture coexisting with the considered region of the EL picture are identified, and the reconstruction signal is taken from the identified BL region. The definition of the coexisting region can be made to describe an integer scaling factor of the BL resolution (e.g., 2x scalability), a fractional scaling factor of the BL resolution (e.g., 1.5 scalability), or an EL resolution equal to the BL resolution (e.g., SNR scalability). In the case of SNR scalability, the blocks coexisting in the BL picture have the same coordinates as the EL block to be predicted.
[0429] The final EL prediction is obtained by adding the filtered EL difference samples and the filtered BL restoration samples.
[0431] Several possible variations of the mode for combining the motion-compensation enhancement layer difference signal and the (upsampled / filtered) base layer restoration signal are listed below:
[0432] Various versions of methods for using (upsampled / filtered) BL signals may be used. The upsampled / filtered BL signals used for the three versions may differ in the interpolation filters used (including interpolation filters that also filter integer-sample locations), or the BL signal upsampled / filtered for the second version may be obtained by filtering the BL signal upsampled / filtered for the first version. The selection of one of the different versions may be signaled at the sequence, picture, slide, maximum coding unit, coding unit level, prediction block level, or transformation block level, or may be inferred from the characteristics of the corresponding restored base layer signal or the transmitted coding parameters.
[0433] Different filters can be used to upsample / filter the BL residual signal in the case of BL-resi mode and the BL restored signal in the case of BL_reco mode.
[0434] · It is also possible to combine two or more hypotheses of motion-compensated difference signals with an upsampled / filtered BL signal. This is illustrated in FIG. 32.
[0436] Considering the above, prediction can be performed using a combination of base layer restoration and enhancement layer restoration (see Aspect C). One major difference from the above description regarding FIGS. 11, 12 and 13 is the coding mode for obtaining intra-layer prediction (34) performed temporally rather than spatially. This means that temporal prediction (32) is used instead of spatial prediction (30) to form the intra-layer prediction signal (34). Thus, some aspects described below are easily convertible to the above embodiments of FIGS. 6 through 10 and 11 through 13, respectively. A number of methods include the following: A restored base layer signal (upsampled / filtered) is combined with an inter-prediction signal, and the inter-prediction is derived by motion-compensated prediction using the restored enhancement layer pictures. The final prediction signal is obtained by weighting the base layer prediction signal and the inter-prediction signal in such a way that the difference frequency components use different weights. This can be realized, for example, by the following:
[0437] · Add the obtained filtering signals, filter the inter-predicted signal with a high-pass filter, and filter the base layer predicted signal with a low-pass filter.
[0438] · The obtained transform blocks are superimposed to transform the inter-predicted signal and the base-layer predicted signal, wherein different weighting factors are used for different frequency positions. The obtained transform block can be used as an enhancement layer predicted signal and inversely transformed, or the obtained transform coefficients can be scaled and added to the transmitted transform coefficient levels and then inversely transformed to obtain the restored block before in-loop processing and deblocking.
[0440] This mode may also be referred to as the BL_comb mode illustrated in Fig. 33. Simply put, EL prediction can be expressed as follows.
[0441] EL prediction = BL_weighting( BL_reco(t 0 ) ) + EL_weighting(MCP_filter( EL_reco(t1) ) ).
[0443] In a preferred embodiment, the weights are made based on the ratio of EL resolution to BL resolution. For example, when BL is scaled up by a factor in the range [1, 1.25), a specific set of weights for EL and BL restoration may be used. When BL is scaled up by a factor in the range [1.25 1.75], a different set of weights may be used. When BL is scaled up by a factor of 1.75 or higher, a more different set of weights may be used, and so on.
[0445] Rendering individual weights that depend on a scaling factor separating the base and enhancement layers is also feasible in other embodiments related to spatial intra-layer prediction.
[0447] In another preferred embodiment, weights are created depending on the predicted EL block size. For example, for a 4x4 block in EL, a weighting matrix may be defined such that another weighting matrix specifies weights for BL restoration transform coefficients and a weighting matrix specifies weights for EL restoration transform coefficients. The weighting matrix for BL restoration transform coefficients may be, for example, as follows, and
[0448] 64, 63, 61, 49,
[0449] 63, 62, 57, 40,
[0450] 61, 56, 44, 28,
[0451] 49, 46, 32, 15,
[0453] And the weight matrix for the EL restoration transform coefficients is, for example,
[0454] 0, 2, 8, 24,
[0455] 3, 7, 16, 32,
[0456] 9, 18, 20, 26,
[0457] 22, 31, 30, 23,
[0458] It could be.
[0460] Individual weighting matrices can be defined for block sizes such as 8x8, 16x16, and 32x32, similarly and differently. The actual transformation used for frequency domain weights may be different from or the same as the transformation used to code the predicted residuals. For example, an integer approximation for the DCT can calculate the transformation coefficients of the predicted residuals coded in the frequency domain and be used for both frequency domain weights.
[0462] In another preferred embodiment, the maximum transform size is defined for frequency domain weighting to limit computational complexity. When the EL block size under consideration is larger than the maximum transform size, EL restoration and BL restoration are spatially divided into a series of adjacent sub-blocks, frequency domain weighting is performed on the sub-blocks, and the final predicted signal is formed by summing the weighted results.
[0464] In addition, the above weighting may be performed on a selected subset of color components or on luminosity and color difference components.
[0466] In the following, different possibilities for deriving enhancement layer coding parameters are described. Coding (or prediction) parameters used to reconstruct enhancement layer blocks can be derived from coding parameters coexisting in the base layer by multiple methods. The base and enhancement layers may have different spatial resolutions or the same spatial resolution.
[0468] In the Scalable Video Extension of H.264 / AVC, inter-layer motion prediction is performed on macroblock types, which is signaled by the syntactic element base mode flag. If the base mode flag is equal to 1, the corresponding reference macroblock in the base layer is inter-coded and the enhancement layer macroblock is also inter-coded, and all motion parameters are inferred from the coexisting base layer block(s). Otherwise (when the base mode flag is equal to 0), for each motion vector, the so-called motion prediction flag ( motion prediction flag ) syntax elements are transmitted, and it is determined whether base layer motion vectors are used as motion vector predictors. If the motion prediction flag is equal to 1, the motion vector predictors of the base layer's coexisting reference blocks are used as motion vector predictors and scaled according to the resolution ratio. If the motion prediction flag is equal to 0, the motion vector predictors are calculated while specified in H.264 / AVC.
[0470] In the following, methods for deriving enhancement layer coding parameters are described. Sample batches associated with a base layer picture are decomposed into blocks, and each block has associated coding (or prediction) parameters. In other words, all sample locations within a specific block have the same associated coding (or prediction) parameters. Coding parameters may include parameters for motion compensation prediction consisting of motion hypotheses, reference indices, motion vectors, motion vector predictor identifiers, and numbers of merge identifiers. Coding parameters may also include intra-prediction parameters, such as intra-prediction directions.
[0472] Using coexisting information from the base layer, the coding of blocks in the enhancement layer can be signaled within the bitstream.
[0474] For example, the derivation of enhancement layer coding parameters (see Aspect T) can be constructed as follows. For an NxM block in the enhancement layer, this is signaled using coexisting layer information, and the coding parameters related to sample locations within the block can be derived based on the coding parameters assigned to coexisting sample locations in the base layer sample batch.
[0476] In a specific embodiment, this process is performed by the following steps:
[0477] 1. Derive coding parameters for each sample position in the NxM enhancement layer block based on the base layer coding parameters.
[0478] 2. Inducing partitioning of NxM enhancement layer blocks into sub-blocks such that all sample locations within a specific sub-block have the same related coding parameters.
[0480] The second step can also be omitted.
[0482] Step 1 can be performed using a function of the enhancement layer sample location that provides coding parameters, i.e.
[0483] c = f c (p el )am.
[0485] For example, to ensure the minimum block size mxn of the enhancement layer, the above function f c Is
[0486] f p, m x m (p el )=p bl
[0487] x bl =floor * n
[0488] y bl =floor * m
[0489] p bl = (x bl , y bl )
[0490] p el = (x el , y el )
[0491] A function f given together p,m x n in p bl The related coding parameters c can be returned.
[0493] The distance between two horizontally or vertically adjacent base layer sample positions is equal to 1, and both the top-left base layer sample and the top-left enhancement layer sample have the position p=(0, 0).
[0495] In another example, the above function f c (p el ) is the base layer sample position p el It is possible to return coding parameters related to the base layer sample position closest to it. The above function f c (p el) may also interpolate coding parameters when the given enhancement layer sample positions have fractional components in units of the distance between base layer sample positions.
[0497] Before returning the motion parameters, the above function f c It rounds the spatial replacement components of the motion parameters to the nearest available value in the enhancement layer sampling grid.
[0499] After Step 1, each enhancement layer sample can be predicted because each sample location has relevant prediction parameters after Step 1. Nevertheless, in Step 2, block partitioning can be induced for the purpose of performing prediction tasks on larger blocks of samples, or for the purpose of transform coding prediction residuals within the blocks of the induced partitioning.
[0501] Step 2 can be performed by grouping the enhancement layer sample locations into square or rectangular blocks, each of which is decomposed into one of the sets of decompositions allowed as sub-blocks. The square or rectangular blocks correspond to leaves in a quad tree structure that may exist at different levels, as illustrated in FIG. 34.
[0503] The level and decomposition of each square or rectangular block can be determined by performing the following sequential steps:
[0504] a) Set the highest level to the level corresponding to the NxM size block. Set the current level to the lowest level, which is the level containing a single block of the minimum block size that is a square or rectangular block. Proceed to step b).
[0505] b) For each square or rectangular block at the current level, if an acceptable decomposition of the square or rectangular block exists, and all sample locations within each sub-block are associated with the same coding parameters or with coding parameters that differ slightly (depending on some difference measurements), the decomposition is a candidate decomposition. Among all candidate decompositions, select one that decomposes the square or rectangular block into the minimum number of sub-blocks. If the current level is the highest level, proceed to step c). Otherwise, set the current level to the next higher level and proceed to step b).
[0506] c) Termination
[0508] The above function f c is selected in such a way that at least one candidate decomposition always exists at the same level of step b).
[0510] Grouping with identical coding parameters is not limited to square blocks, and said blocks may be summarized into rectangular blocks.
[0512] Furthermore, the above grouping is not limited to a quadtree structure, and it is also possible to use a decomposition structure in which the block is decomposed into two rectangular blocks of the same or different sizes. It is also possible to use a decomposition structure that utilizes quadtree decomposition up to a certain level and then decomposes into two rectangular blocks. Additionally, any other block decomposition is possible.
[0514] In contrast to the SVC interlayer motion parameter prediction mode, the aforementioned mode is not supported at the macroblock level (or any block size that supports the largest block size). That is, the mode cannot signal the maximum supported block size, but blocks of the maximum supported block size (macroblocks in MPEG-4, H.264 coding tree blocks / maximum coding units in HEVC) can be hierarchically partitioned, and the use of the interlayer motion mode with small blocks / coding units can signal the supported block size (for the corresponding block). In certain embodiments, this mode is supported for selected block sizes. Furthermore, the use of this mode may be limited to transmitting only the size of the corresponding block (among other coding parameters) or the value of the signal syntax element, or the use of this mode may be restricted to other corresponding block sizes. Another difference from the interlayer motion parameter prediction mode of the SVC extension of H.264 / AVC is that blocks coded in this mode are not fully intercoded. The block may include intra-coded sub-blocks depending on the coexisting base layer signal.
[0516] One of several methods for reconstructing M × M enhancement layer blocks of samples using coding parameters derived by the method described above is that they can be signaled within a bitstream. Such a method for predicting enhancement layer blocks using derived coding parameters may include the following:
[0517] · Deriving a prediction signal for the enhancement layer block using the restored enhancement layer reference pictures and derived motion parameters for motion compensation.
[0518] Calculate the motion compensation signal (b) and the (upsampled / filtered) base layer restoration (a) for the current picture using the enhancement layer reference picture and derived motion parameters resulting from subtracting the (upsampled / filtered) base layer restoration from the restored enhancement layer picture.
[0519] · Combine the base layer residue (the difference between the predicted or inverse transform and the restored signal of the coded transform coefficient values) (a) for the current picture and the motion compensation signal (b) using the restored enhancement layer reference pictures and derived motion parameters.
[0521] The process of deriving coding parameters for the subblocks and partitioning the current block into smaller blocks allows some of the subblocks to be classified by intra-coding, while others are classified by inter-coding. For the inter-coded subblocks, motion parameters are derived from the coexisting base layer blocks. However, if the coexisting base layer blocks are intra-coded, the corresponding subblocks of the enhancement layer may be classified when intra-coded. For samples of such intra-coded subblocks, the enhancement layer signal can be predicted using information from the base layer according to the following example:
[0522] · The (upsampled / filtered) version of the corresponding base layer restoration is used as the intra-prediction signal.
[0523] · The derived intra prediction parameters are used for spatial intra prediction in the enhancement layer.
[0525] The following embodiments for predicting an enhancement layer block using a weighted combination of prediction signals include a method for generating a prediction signal for an enhancement layer block by combining (a) an enhancement layer internal prediction signal obtained by spatial or temporal (i.e., motion compensation) prediction using restored enhancement layer samples and (b) a base layer prediction signal which is a base layer restoration (upsampled / filtered) for the current picture. The final prediction signal is obtained by weighting the base layer prediction signal and the enhancement layer internal prediction signal in such a way that weights according to a weighting function are used for each sample.
[0527] The weighting function can be realized, for example, by the following method. The low-pass filtered version of the base layer restoration is compared with the low-pass filtered version of the original enhancement layer internal prediction signal. Weights for each sample position to be used to combine the (upsampled / filtered) base layer restoration and the original inter-prediction signal are derived from the comparison. The said weights can be derived by mapping the difference u - v to the weight w using the transfer function t, i.e.
[0528] t(uv)=w.
[0530] Different weighting functions can be used for different block sizes of the predicted current block. Additionally, the weighting function can be modified according to the temporal distance between reference pictures, and inter-prediction hypotheses are obtained therefrom.
[0532] When the prediction signal inside the enhancement layer is a prediction signal, for example, the weighting function may be realized using different weights depending on the position within the current block to be predicted. In a preferred embodiment, the method for deriving enhancement layer coding parameters is used, and step 2 of the method uses a set of allowable decompositions of square blocks as described in FIG. 35.
[0534] In a preferred embodiment, the function fc (p el ) is the function f described above with m = 4, n = 4 p,m x n (p el Returns the coding parameters related to the base layer sample position given by ).
[0536] In the example, the function f c (p el ) returns the following coding parameters:
[0537] ·First, the base layer sample position is p bl =f p,4 x 4 (p el It is derived according to ).
[0538] ·p bl If it has related inter-predicted parameters obtained by merging with this previously coded base layer block (or has identical motion parameters), c is identical to the motion parameters of the enhancement layer block corresponding to the base layer block used for merging in the base layer (i.e., the motion parameters are duplicated from the corresponding enhancement layer block).
[0539] · Otherwise, c is p bl It is identical to the coding parameters related to.
[0541] In addition, a combination of the above embodiments is also possible.
[0543] In another embodiment, for an enhancement layer block signaled using coexisting base layer information, a default set of motion parameters is associated with derived intra-predicted parameters and enhancement layer sample locations, so that the block can be merged with a block containing these samples. The default set of motion parameters consists of an indicator using one or two hypotheses, and reference indices refer to a first picture in a reference picture list, motion vectors having a zero (0) spatial arrangement.
[0545] In another embodiment, for an enhancement layer block signaled using coexisting base layer information, enhancement layer samples having derived motion parameters are first restored and predicted in a certain order. Then, samples having derived intra prediction parameters are predicted in the order of intra restoration. Accordingly, the intra prediction can utilize sample values already restored from (a) any adjacent inter prediction block and (b) adjacent intra prediction blocks that are predecessors in the order of intra restoration.
[0547] In another embodiment, for the enhancement layer blocks being merged (i.e., by taking motion parameters derived from other inter-prediction blocks), the list of merge candidates additionally includes a candidate from the corresponding base layer block, and if the enhancement layer has a higher spatial sampling rate than the base layer, includes up to four candidates derived from the base layer candidate by refining the spatial arrangement component for adjacent values available only in the enhancement layer.
[0549] In another embodiment, the difference measurement used in step 2 refers to b) a very small difference in the sub-block when there are no differences at all, that is, the sub-block can be formed only when all included sample locations have the same derived coding parameters.
[0551] In another embodiment, the difference measurement used in step 2 refers to only very small differences in the sub-blocks when (a) all included sample positions have derived motion parameters and no pairs of sample positions within the block have derived motion parameters different from a specific value according to the vector norm applied to the corresponding motion vectors, or b) all included sample positions have derived intra-prediction parameters and no pairs of identical positions within the block have derived intra-prediction parameters different from a specific angle of directional intra-prediction. The resulting parameters for the sub-blocks are calculated by mean or median operations. In another embodiment, the partitioning obtained by inferring coding parameters from the base layer can be further refined based on additional information signaled within the bitstream. In another embodiment, the residual coding for the block in which coding parameters are inferred from the base layer is independent of the partitioning into blocks inferred from the base layer. For example, the inference of coding parameters from the base layer may partition the blocks into several sub-blocks, each with an individual set of coding parameters, but a single transformation may be applied to the blocks. Alternatively, the partitioning and coding parameters for the sub-blocks may be divided for the purpose of transforming the blocks into smaller blocks, where the partitioning into transformed blocks is independent of the partitioning inferred into blocks with different coding parameters.
[0553] In another embodiment, residual coding for blocks inferred from the base layer depends on partitioning into blocks inferred from the base layer. This means that, for example, for transform coding, the partitioning of blocks in transform blocks depends on partitioning inferred from the base layer. In one version, a single transform may be applied to each of the sub-blocks having different coding parameters. In another version, the partitioning may be improved based on additional information including a bitstream. In another version, some of the sub-blocks may be summarized into larger blocks when signaled within the bitstream for the purpose of transform coding the residual signal.
[0555] Examples obtained by combining the embodiments described above are also possible.
[0557] Regarding enhancement layer motion vector coding, the following section describes a method to reduce motion information in scalable video coding applications by utilizing motion information coded in the base layer to efficiently code motion information in the enhancement layer and providing multiple enhancement layer predictors. This idea is applicable to scalable video coding that includes spatial, temporal, and quality scalability.
[0559] Inter-layer motion prediction in H.264 / AVC Scalable Video Extension is a syntactic element base mode flag It is performed on macroblock types signaled by. If base mode flag If is equal to 1, the corresponding reference macroblock in the base layer is inter-coded and the enhancement layer macroblock is also inter-coded, and all motion parameters are inferred from the coexisting base layer block(s). Otherwise, ( base mode flag is equal to 0), for each motion vector, the so-called motion prediction flagSyntax elements are transmitted, and it is determined whether base layer motion vectors are used as motion vector predictors. If motion prediction flag If is equal to 1, the motion vector predictor of the coexisting reference block in the base layer is scaled according to the resolution ratio and used as a motion vector predictor. If motion prediction flag If this is equal to 0, the motion vector predictor is calculated as specified in H.264 / AVC. In HEVC, motion parameters are predicted by applying advanced motion vector competition (AMVP). AMVP features two competing spatial and one temporal motion vector predictors. Spatial candidates are selected from the locations of adjacent prediction blocks assigned to the left or above the current prediction block. Temporal candidates are selected from coexisting locations of the previously coded picture. The locations of all spatial and temporal candidates are illustrated in Fig. 36.
[0561] After spatial and temporal candidates are inferred, a redundancy check is performed to introduce a 0 motion vector as a candidate for the list. An index handling the candidate list is transmitted to identify the motion vector predictor used with the motion vector difference for motion compensation prediction.
[0562] HEVC further utilizes a block merging algorithm aimed at reducing coding redundant motion parameters by deriving a quad-tree-based coding design. This is achieved by generating regions composed of multiple prediction blocks that share the same motion parameters. These motion parameters need to be coded once for the first prediction block of each region—which assigns new motion information. A block merging algorithm similar to AMVP constructs a list containing possible merge candidates for each d-side block. The number of candidates is NumMergeCands It is defined by, which is signaled in the slice header and ranges from 1 to 5. Candidates are inferred from spatially adjacent prediction blocks and from prediction blocks in coexisting temporal pictures. The possible sample locations for the prediction blocks considered as candidates are the same as those shown in Fig. 36. An example of a block merging algorithm with possible prediction block partitioning in HEVC is illustrated in Fig. 37. In Fig. (a), the bold lines define prediction blocks that all have the same motion data and are merged into a single region. This motion data is transmitted only in block S. The current prediction block to be coded is indicated by X. The blocks in the underlined region do not yet have the relevant prediction data, and these prediction blocks are successors to prediction block X in the block scanning order. The dots indicate the sample locations of adjacent blocks that are possible spatial merge candidates. Before the possible candidates are inserted into the predictor list, a redundancy check for spatial candidates is performed, as shown in Fig. 37 (b).
[0564] The number of spatial and temporal candidates NumMergeCands In smaller cases, additional candidates are provided by zero motion vector candidates or by combining existing candidates. When a candidate is added to the list, an index is provided that is used to identify the candidate. With the addition of a new candidate to the list, the index is (starting from 0) the list index NumMergeCands - It is incremented until it is completed with the final candidate identified by 1. A fixed-length codeword is used to code the merge candidate index to ensure independence from the parsing of the bitstream and the derivation of the candidate list.
[0566] The following section describes a method for utilizing multiple enhancement layer predictors, including predictors derived from the base layer, to code the motion parameters of the enhancement layer. Motion information already coded for the base layer can be used to significantly reduce the motion data rate while coding the enhancement layer. This method includes the possibility of directly deriving all motion data for the prediction block from the base layer when no additional motion data to be coded is required. In the following description, the term prediction block is referred to as a prediction unit in HEVC and an MxN block in H.264 / AVC, and can be understood as a general set of samples in a picture.
[0568] The first part of the current section concerns expanding the list of motion vector prediction candidates by the base layer motion vector predictor (see Aspect K). Base layer motion vectors are added to the list of motion vector predictors during enhancement layer coding. This is achieved by inferring one or more motion vector predictors of coexisting prediction blocks from the base layer and using them as candidates in the list of predictors for motion compensation prediction. The coexisting prediction blocks of the base layer are located in the center, left, top, right, and below the current block. The base layer prediction blocks at selected locations do not contain any motion-related data, are not located outside the current range, are not currently accessible, and alternative locations can be used to infer motion vector predictors. These alternative locations are illustrated in Fig. 38.
[0570] Motion vectors inferred from the base layer can be scaled according to the resolution ratio before being used as predictor candidates. An index covering the motion vector difference as well as a list of motion vector predictor candidates is transmitted for the prediction block, which specifies the final motion vector used for motion-compensated prediction. In contrast to the scalable extensions based on H.264 / AVC standards, the embodiments presented herein do not constitute the utilization of motion vector predictors of blocks coexisting in the reference picture—rather, they can be handled by the transmitted index and are available in the list among other predictors.
[0571] In an embodiment, a motion vector is derived from the center position C1 of the coexisting prediction block in the base layer and added to the top of the candidate list as the first input. The candidate list of motion vector predictors is expanded by a single item. If there is no motion data available for the sample positions C1 in the base layer, the list configuration is untouched. In another embodiment, any sequence of sample positions in the base layer can be identified for motion data. If motion data is found, the motion vector predictor at the corresponding position is available for motion compensation prediction in the enhancement layer and is inserted into the candidate list. Furthermore, a motion vector predictor derived from the base layer can be inserted into the candidate list at any other position in the list. In another embodiment, a base layer motion predictor can be inserted into the candidate list if certain constraints are satisfied. These constraints may include the value of the merge flag of the coexisting reference block, which must be equal to 0. Another constraint may be the size of the prediction block in the enhancement layer that makes the size of the coexisting prediction block in the base equal with respect to the resolution ratio. For example, in an application of K x spatial scalability—if the width of coexisting blocks in the base layer is N, the motion vector predictor can be inferred only when the width of the prediction block to be coded in the enhancement layer is K*N.
[0572] In another embodiment, one or more motion vector predictors from several sample locations of the base layer may be added to the candidate list of the enhancement layer. In another embodiment, a candidate having a motion vector predictor inferred from coexisting blocks may replace spatial or temporal candidates in the list rather than expanding the list. It may also include multiple motion vector predictors derived from base layer data into a motion vector prediction candidate list.
[0574] The second part concerns expanding the list of merge candidates based on base layer candidates (see Perspective K). Motion data from one or more coexisting blocks of the base layer is added to the list of merge candidates. This approach allows for the creation of merge regions that share the same motion parameters across the base and enhancement layers. Similar to the previous section, the base layer block covering the coexisting sample at the central location is not limited to this central location but can be derived from any location in the immediate vicinity, as depicted in Fig. 38. If motion data is not available or accessible for a specific location, alternative locations may be selected to infer possible merge candidates. The derived motion data may be scaled according to the resolution ratio before being inserted into the list of merge candidates. An index handling the list of merge candidates is transmitted to define the motion vector, which is used for motion compensation prediction. However, the above method may also compress possible motion predictor candidates based on the motion data of the prediction block in the base layer.
[0575] Sample location in FIG. 38 in the example C 1 The motion vector predictor of a coexisting block in the base layer covering it is considered as a possible merge candidate for coding the current prediction block in the enhancement layer. However, the motion vector predictor is if the reference block's merge_flagIf it is identical to 1 or if the coexisting reference block does not contain motion data, it is not inserted into the list. Motion vector predictors derived in any other case are added as a second input to merge the candidate list. Note that in this embodiment, the length of the merge candidate list is maintained and not expanded. In another embodiment, one or more motion vector predictors may be derived from prediction blocks covering any of the sample locations as depicted in FIG. 38 and may be added to the merge candidate list. In yet another embodiment, some motion vector predictors of the base layer may be added to the merge candidate list at any location. In yet another embodiment, one or more motion vector predictors may be added to the merge candidates only if certain constraints are satisfied. Such constraints include the size of the prediction blocks of the enhancement layer matching the size of the coexisting block of the base layer (regarding the resolution ratio as described in the previous embodiment section on motion vector prediction). Another constraint is in yet another embodiment merge_flag It may be that the value is equal to 1. In another embodiment, the length of the merge candidate list may be extended by the number of motion vector predictors inferred from the coexisting reference blocks of the base layer.
[0577] The third part of this specification concerns rearranging the motion parameter (or merge) candidate list using base layer data (see Aspect L) and describes the process of rearranging the merge candidate list according to information already coded in the base layer. When a coexisting base layer covering the samples of the current block predicts motion rewards along with candidates derived from a particular origin, (if any) corresponding enhancement layer candidates from an equal origin are placed as the first input at the top of the merge candidate list. This step is equivalent to handling this candidate with the lowest index and assigning the cheapest codeword to this candidate.
[0578] In the embodiment, the coexisting base layer block is motion-compensated with a candidate originating from the prediction block covering sample position A1, as depicted in FIG. 38. If the merge candidate list of the prediction block in the enhancement layer includes a candidate originating from the corresponding sample position A1 within that motion vector prediction enhancement layer, this candidate is placed as the first input to the list. Consequently, this candidate is indexed by index 0 and thus assigned the shortest fixed-length codeword. In this embodiment, this step is performed after the derivation of the motion vector predictor of the coexisting base layer block for the merge candidate list in the enhancement layer. For this reason, the reordering process assigns the lowest index to the candidate originating from the corresponding block as the motion vector predictor of the coexisting base layer block. The second lowest index is assigned to the candidate derived from the coexisting block in the base layer, as described in the second part of this section. Furthermore, the reordering process of the coexisting block in the base layer merge_flagIt occurs only when is identical to 1. In another embodiment, the rearrangement process of the prediction blocks coexisting in the base layer merge_flag It can be performed independently of the value. In another embodiment, a candidate having a motion vector predictor of a corresponding origin may be placed at any position in the merge candidate list. In another embodiment, the rearrangement process may remove all other candidates from the merge candidates. Here, candidates whose motion vector predictors have the same origin as the motion vector predictors used for motion compensation prediction of blocks coexisting in the base layer are retained in the list. In this case, a single candidate is available and the index is not transmitted.
[0580] The fourth part of this specification includes a process of rearranging motion vector predictor candidates using base layer data (see Perspective L) and rearranging the list of motion vector prediction candidates using the motion parameters of the base layer blocks. If coexisting base layer blocks covering the samples of the current prediction block use motion vectors from a specific source, the motion vector predictor from the corresponding source in the enhancement layer is used as the first input in the motion vector prediction list of the current prediction block. This assigns the cheapest codeword to this candidate.
[0581] In the embodiment, the coexisting base layer blocks are motion-compensated with candidates originating from the prediction block covering sample position A1, as depicted in FIG. 38. If the motion vector predictor candidates for the blocks in the enhancement layer include candidates originating from the corresponding sample position A1 within the enhancement layer, these candidates are placed as the first entry for the list. Consequently, these candidates are indexed by index 0 and thus assigned the shortest fixed-length codeword. For this reason, the rearrangement process assigns the lowest index to the candidate originating from the corresponding block as the motion vector predictor for the coexisting base layer blocks. The second lowest index is assigned to the candidate derived from the blocks coexisting in the base layer, as described in the first part of this section. Furthermore, the rearrangement process of the blocks coexisting in the base layer merge_flag It occurs only when is equal to 0. In another embodiment, the rearrangement process of the prediction blocks coexisting in the base layer merge_flag It can be performed independently of the value of. In another embodiment, a candidate having a motion vector predictor of a corresponding source can be placed at any position in the motion vector predictor candidate list.
[0583] The following concerns the enhancement layer coding of transformation coefficients.
[0585] In modern video and image coding, the residual of the predicted signal is forward-transformed, and the resulting quantized transform coefficients are signaled within the bitstream. This coefficient coding follows a fixed design:
[0587] Different scan directions are defined depending on the transform size (for luma residuals: 4x4, 8x8, 16x16, and 32x32). Given the first and last positions in the scan order, these scans uniquely determine which coefficient positions may be important, and thus need to be coded as such. In all scans, the first coefficient is set as the DC coefficient at position (0, 0), while the last position must be signaled within the bitstream, which is accomplished by coding its x (horizontal) and y (vertical) positions within the transform block. Starting from the last position, the signaling of the important coefficients is performed in reverse scan order until the DC position is reached.
[0589] For transformation sizes 16x16 and 32x32, only one scan is defined, namely a 'diagonal scan', whereas transformation blocks of sizes 2x2, 4x4, and 8x8 may additionally utilize 'vertical' and 'horizontal' scans. However, the utilization of vertical and horizontal scans is limited to the remnants of intra-predictive coding units, and the scans actually utilized are derived from the direction modes of intra-predict. Direction modes with indices in ranges 6 and 14 yield vertical scans, whereas direction modes with indices in ranges 22 and 30 yield horizontal scans. All remaining direction modes yield diagonal scans.
[0591] FIG. 39 shows diagonal, vertical, and horizontal scans as defined for a 4x4 transformation block. The coefficients of larger transformations are divided into subgroups of 16 coefficients. These subgroups allow for hierarchical coding of important coefficient locations. Subgroups signaled as non-important do not contain any important coefficients. Scans for 8x8 and 16x16 transformations are described in FIG. 40 and FIG. 41, respectively, along with their corresponding subgroup divisions. Large arrows indicate the scan order of the coefficient subgroups.
[0593] In a zigzag scan, for blocks larger than 4x4, the subgroup consists of 4x4 pixel blocks scanned in the zigzag scan. Fig. 42 shows a vertical scan for a 16x16 transformation as proposed in JCTVC-G703.
[0595] The following section describes extensions to transform coefficient coding. These include modified coding of important coefficient locations, methods for assigning scans to transform blocks, and the introduction of new scan modes. These extensions allow for better adaptation to different coefficient distributions within transform blocks, thereby achieving coding gain in the sense of rate-distortion.
[0597] New implementations for vertical and horizontal scan patterns are introduced for 16x16 and 32x32 transformation blocks. In contrast to previously proposed scan patterns, the size of the scan subgroups is 16x1 for the horizontal scan and 1x16 for the vertical scan, respectively. Subgroups of 8x2 and 2x8 sizes may also be selected, respectively. The subgroups themselves are scanned in the same manner.
[0599] Vertical scanning is effective for transformation factors located as column-wise spreads. This can be observed in images containing horizontal edges.
[0601] Horizontal scanning can be effective for transformation factors found in row-direction variance. This can be observed in images containing vertical edges.
[0603] Figure 43 shows the implementation of vertical and horizontal scans for 16x16 transformation blocks. Count subgroups are defined as a single column or a single row, respectively. The VerHor scan is an introduced scan pattern that allows for the coding of counts in columns by row-wise scanning. For 4x4 blocks, the first column is scanned, followed by the remainder of the first row, the remainder of the second column, and the remainder of the second row counts. Then the remainder of the third column is scanned, and finally, the fourth row and column are scanned.
[0605] For larger blocks, the block is divided into 4x4 subgroups. These 4x4 blocks are scanned in a VerHor scan, while the subgroups are the scanned VerHor scans themselves.
[0607] Verhor scans can be used when coefficients are located in the first column and row of a block. In this way, coefficients are scanned earlier than in cases where other scans are used, for example, for a diagonal scan. This can be observed for images that include both horizontal and vertical edges.
[0609] Figure 44 shows a VerHor scan for a 16x16 transformation block.
[0611] Other scans are also feasible. All combinations between subgroups and scans can be utilized. For example, a horizontal scan is used for 4x4 blocks with diagonal scans of subgroups. Adaptive selection of scans can be applied by selecting different scans for each subgroup.
[0613] It should be noted that different scans can be implemented in a manner where the transform coefficients are rearranged after quantization on the encoder side and conventional coding is utilized. On the decoder side, the transform coefficients are generally decoded and rearranged before scaling and inverse transformation (or before scaling and after inverse transformation).
[0615] Different parts of the base layer signal can be utilized to derive coding parameters from the base layer signal. Among such signals:
[0616] · Coexisting restoration base layer signals
[0617] · Coexisting residual base layer signal
[0618] · Estimated enhancement layer residual signal obtained by subtracting the enhancement layer prediction signal from the restored base layer signal
[0619] · Picture splitting of the base layer frame
[0621] Gradient parameters ( Gradient parameters ):
[0622] Gradient parameters can be derived according to the following:
[0623] For each pixel of the investigated block, a slope is calculated. From this slope, a size and an angle are calculated. The angle that occurs most frequently in the block is related to the block (block angle). The angle can be rounded so that only three directions are used: horizontal (0°), vertical (90°), and diagonal (45°).
[0625] Detecting corners ( Detecting edges ):
[0626] A corner detector can be applied to blocks being investigated according to the following:
[0627] First, the above block is an nxn smoothing filter (e.g., Gaussian).
[0628] A gradient matrix of size mxm is used to calculate the gradient for each pixel. The size and angle of all pixels are calculated. The angles can be rounded so that only three directions are used: horizontal (0°), vertical (90°), and diagonal (45°).
[0629] For all pixels with a size greater than a specific threshold 1, adjacent pixels are checked. If an adjacent pixel has a size greater than threshold 2 and has the same angle as the current pixel, the counter of this angle is incremented. The counter with the highest number for the entire block is selected based on the block angle.
[0631] Base layer coefficients obtained by forward transformation ( Obtaining base layer coefficients by forward transformation)
[0632] To derive coding parameters, for a specific TU, the coexisting signals investigated from the frequency domain of the base layer signal (restored base layer signal / residual base layer signal / measurement enhancement layer signal) can be converted to the frequency domain. Preferably, as used by a specific enhancement layer TU, this is performed using the same conversion.
[0633] The resulting base layer transform coefficients are quantized or not.
[0634] To obtain comparable coefficient distributions as in enhancement layer blocks, rate distortion quantization with modified lambda can be used.
[0636] Scan valid score of the given distribution and scan ( Scan effectiveness score of a given distribution and scan)
[0637] The scan effective score of a given important coefficient distribution can be defined according to the following:
[0638] Let each location within the investigated block be represented by its index according to the order of the scans. Subsequently, the sum of the index values of the critical coefficient locations is defined by the valid score of this scan. In this way, scans with smaller scores represent a specific distribution more efficiently.
[0640] Adaptive scan pattern selection for transform coefficient coding selection for transformation coefficient coding)
[0641] If several scans are available for a specific TU, a rule needs to be defined to uniquely select one of them.
[0643] Scan pattern selection method ( Methods for scan pattern selection )
[0644] The selected scan can be derived directly from already decoded signals (without any additional data transmission). This can be performed based on the characteristics of coexisting base layer signals, or by utilizing only enhancement layer signals.
[0645] The scan pattern can be derived from the EL signal by the following.
[0646] · As explained above, the latest induction rules
[0647] · Use of scan patterns for color difference residues selected for coexisting brightness residues (luminance residues)
[0648] · Defines a fixed mapping between coding modes and scan patterns used
[0649] · Deriving the scan order from the final critical coefficient location (related to the assumed fixed scan pattern)
[0651] In a preferred embodiment, the scan pattern is selected based on the final position already decoded according to the following:
[0653] The final position is represented by x and y coordinates within the transformation block and is already decoded (for scan-dependent final coding, a fixed scan pattern is assumed for the decoding process of the final position, which may be the latest scan pattern of that TU). Let T be a defined threshold, which may depend on a specific transformation size. If neither the x nor the y coordinates of the last significant position exceed T, a diagonal scan is selected.
[0655] Otherwise, x is compared with y. If x exceeds y, a horizontal scan is selected; otherwise, a vertical scan is selected. The preferred value of T for 4x4 TUs is 1. The preferred value of T for TUs larger than 4x4 is 4.
[0657] In a more preferred embodiment, as described in the previous embodiment, the derivation of the scan pattern is limited to be performed only on TUs of size 16x16 and 32x32. Furthermore, it may be limited only to luminance signals.
[0659] The scan pattern may also be derived from the BL signal. To derive a selected scan pattern from the base layer signal, any coding parameter described above may be used. In particular, the gradient of the coexisting base layer signal may be calculated and compared to predefined thresholds, and / or potentially discovered edges may be utilized.
[0661] In a preferred embodiment, the scan direction is derived according to the block inclination angle as follows: for an inclination quantized in the horizontal direction, a vertical scan is used. For an inclination quantized in the vertical direction, a horizontal scan is used. Otherwise, a diagonal scan is selected.
[0663] In a more preferred embodiment, the scan pattern is derived as described in the previous embodiment, but only for such transformation blocks, the number of occurrences of the block angle exceeds a threshold. The remaining transformation units are decoded using the latest scan pattern of the TU.
[0665] If base layer coefficients of coexisting blocks are available, they are explicitly signaled in the base layer data stream or calculated by forward transformation, and these can be utilized in the following ways.
[0666] · For each available scan, the cost of coding base layer coefficients can be measured. The scan with the minimum cost is used to decode enhancement layer coefficients.
[0667] The valid score for each available scan is calculated against the base layer coefficient distribution, and the scan with the minimum score is used to decode the enhancement layer coefficients.
[0668] The distribution of base layer coefficients within the transformation block is classified into one of a predefined set of distributions, which is associated with a specific scan pattern.
[0669] · The scan pattern is selected based on the final critical base layer coefficient.
[0671] If coexisting base layer blocks are predicted using intra prediction, the intra direction of the prediction can be used to induce an enhancement layer scan pattern.
[0673] Furthermore, the transformation size of the coexisting base layer blocks can be used to induce a scan pattern.
[0675] In a preferred embodiment, the scan pattern is derived only for TUs from the BL signal, which represent the remnants of INTRA_COPY mode prediction blocks, and their coexisting base layer blocks are intra-predicted. A modified latest scan selection is used for these blocks. In contrast to the latest scan selection, the intra-predicted direction of the coexisting base layer blocks is used to select the scan pattern.
[0677] Signaling of scan pattern index within bitstream ( Signaling of an scan pattern index within the bitstream) (refer to Perspective R)
[0678] The scan patterns of the conversion blocks can be selected by the encoder in terms of rate-distortion and subsequently signaled within the bitstream.
[0680] A specific scan pattern can be coded by signaling an index to a list of available scan pattern candidates. This list can be a fixed list of scan patterns defined for a specific transformation size, or it can be dynamically populated during the decoding process. Dynamically populating the list enables adaptive picking of such scan patterns, which can likely code a specific coefficient distribution most efficiently. By doing so, the number of available scan patterns for a specific TU can be reduced, and thus, signaling an index to that list is less expensive.
[0681] As described above, the process of selecting scan pattern candidates for a specific TU can utilize any coding parameters and / or follow specific rules, which leverages specific characteristics of the specific TU. Among these:
[0682] · TU represents the residue of the luminance / chrominance signal.
[0683] · TU has a specific size.
[0684] ·TU represents the residue of a specific prediction mode.
[0685] · The final critical location within the TU is known by the decoder and is located within a specific subdivision of the TU.
[0686] ·TU is a part of I / B / P-Slice.
[0687] · The coefficients of TU are quantized using specific quantization parameters.
[0689] In a preferred embodiment, the list of scan pattern candidates includes three scans: 'diagonal scan', 'vertical scan', and 'horizontal scan' for all TUs.
[0691] Additional embodiments can be obtained by making the candidate list include any combination of scan patterns.
[0693] In a specific preferred embodiment, the list of scan pattern candidates may include any of the scans: 'diagonal scan', 'vertical scan', and 'horizontal scan'.
[0695] On the other hand, the scan pattern selected by the latest scan induction (as described above) is set first in the list. If a specific TU has a size of 16x16 or 32x32, additional candidates are added to the list. The order of the remaining scan patterns depends on the final critical coefficient position.
[0697] (Note: Diagonal scans are always the first pattern in the list assuming 16x16 and 32x32 transformations.)
[0699] If its x-coordinate size exceeds its y-coordinate size, the horizontal scan is selected next, and the vertical scan is input to the final position. Otherwise, the vertical scan is to the second position (second position, 2 nd Input is entered into the position, followed by a horizontal scan.
[0701] Other preferred embodiments are obtained by further restricting the conditions for one or more candidates in the list.
[0703] In another embodiment, where the coefficients represent the residue of the luminance signal, the vertical and horizontal scans may be added only to the candidate lists of 16x16 and 32x32 conversion blocks.
[0705] In another embodiment, if both the x and y coordinates of the final critical location are greater than a specific threshold, the vertical and horizontal scans are added to the candidate lists of the transformation blocks. This threshold may be mode and / or TU size dependent. The preferred threshold value is 3 for all sizes greater than 4x4 and 1 for 4x4 TUs.
[0707] In another embodiment, if the x or y coordinates of the final critical location are greater than a specific threshold, the vertical and horizontal scans are added to the candidate list of the transformation block. This threshold may be mode and / or TU size dependent. The preferred threshold value is 3 for all sizes greater than 4x4 and 1 for 4x4 TUs.
[0709] In another embodiment, if both the x and y coordinates of the final critical location are greater than a specific threshold, vertical and horizontal scans are added only to the candidate lists of 16x16 and 32x32 transformation blocks. This threshold may be mode and / or TU size dependent. The preferred threshold value is 3 for all sizes greater than 4x4 and 1 for 4x4 TUs.
[0711] In another embodiment, if both the x and y coordinates of the final critical location are greater than a specific threshold, vertical and horizontal scans are added only to the candidate lists of 16x16 and 32x32 transformation blocks. This threshold may be mode and / or TU size dependent. The preferred threshold value is 3 for all sizes greater than 4x4 and 1 for 4x4 TUs.
[0713] For the described embodiments, specific scan patterns are signaled within a bitstream, and the signaling itself may be performed at different signaling levels. In particular, for each TU (corresponding to a subgroup of TUs having signaled scan patterns), the signaling may be performed at any node of the residual quad-tree (all sub-TUs of the node using the signaled scan use the same candidate list index), at the CU / LCU level, or at the slice level.
[0715] Indexes for candidate lists can be transmitted using fixed-length coding, variable-length coding, arithmetic coding (including context-adaptive binary arithmetic coding), or PIPE coding. If context-adaptive coding is used, the context can be derived based on parameters of adjacent blocks, the coding modes described above, and / or specific characteristics of the specific TU itself.
[0717] In a preferred embodiment, context-adaptive coding is used to signal indices for a list of scan pattern candidates of a TU, whereas a context model is derived based on the location and / or transformation size of the final important location within the TU. Each method described above for deriving scan patterns may also be used to derive a context model to signal explicit scan patterns for a specific TU.
[0719] To code the final critical scanning locations, the following modifications can be used in the enhancement layer:
[0720] Individual context models are used for a subset or the whole of coding modes that utilize base layer information. It is also possible to use different context models for different modes that possess base layer information.
[0721] · Context modeling can rely on data from coexisting base layer blocks (e.g., transformation coefficient distribution in the base layer, gradient information in the base layer, final scanning position in coexisting base layer blocks)
[0722] · The final scanning position can be coded based on the difference from the final base layer scanning position.
[0723] · When the final scanning position is coded by signaling x and y positions within the TU, the context modeling of the second signaled coordinate may depend on the value of the first.
[0724] Each method described above to derive scan patterns independent of the final critical location can also be used to derive context models for signaling the final critical location.
[0726] In certain versions, scan pattern derivation depends on the final critical location.
[0727] · When the final scanning position is coordinated by signaling the x and y positions within the TU, the context modeling of the second coordinate can rely on scan patterns that are still possible candidates when the first coordinate is already known.
[0728] · When the final scanning position is coded by signaling its x and y positions within the TU, the context modeling of the second coordinate can depend on whether the scan pattern has already been uniquely selected when the first coordinate is already known.
[0730] In another version, the scan pattern induction is independent of the final critical location.
[0731] · Context modeling can rely on scan patterns used in a specific TU.
[0732] Each of the methods described above for deriving scan patterns may also be used to derivate context models that signal the final important location.
[0734] Each, to code significance flags and significant locations within the TU (sub-group flags and / or significance flags for single transformation coefficients), the following modifications may be used in the enhancement layer:
[0735] Individual context models are used for a subset or the whole of coding modes that utilize base layer information. It is also possible to use different context models for different modes that possess base layer information.
[0736] · Context modeling can rely on data from coexisting base layer blocks (e.g., important transformation coefficients for specific frequency locations).
[0737] Each method described above to derive scan patterns may also be used to derive context models for signaling important locations and / or their levels.
[0738] A generalized template may be used that measures both the number of important transform coefficients in coexisting base layer signals at similar frequency locations and the number of already coded important transform coefficient levels in the spatially adjacent portions of the coefficients to be coded.
[0739] A generalized template may be used that measures both the levels of important transform coefficients in coexisting base layer signals at similar frequency locations and the number of already coded important transform coefficient levels in the spatially adjacent portions of the coefficients to be coded.
[0740] · Context modeling for sub-group flags may depend on the scan pattern used and / or specific transformation sizes.
[0742] Different context initialization tables for the base and enhancement layers may be utilized. The context model initialization for the enhancement layer can be modified in the following way:
[0743] · The enhancement layer uses individual sets of initialization values.
[0744] · The enhancement layer uses individual sets of initialization values for different operation modes (spatial / temporal or quality scalability)
[0745] Enhancement layer context models having counterparts in the base layer can use the state of their counterparts as an initialization state.
[0746] The algorithm for deriving the initial states of context models may be dependent on the base layer QP and / or delta QP.
[0748] Next, the feasibility of backward adaptive enhancement layer coding using base layer data is explained. The following section describes methods for generating enhancement layer prediction signals in scalable video coding systems. The method utilizes base layer decoded picture sample information to infer the values of prediction parameters; although this information is not transmitted in the coded video bitstream, it is used to form a prediction signal for the enhancement layer. In this way, the total bitrate required to code the enhancement layer signal is reduced.
[0750] Modern hybrid video encoders typically use a hierarchy to decompose the source image into blocks of different sizes. For each block, the video signal is predicted from spatially adjacent blocks (intra-prediction) or from temporally previously coded pictures (inter-prediction). The difference between the predicted and actual images is transformed and quantized. The resulting prediction parameters and transform coefficients are entropy-coded to form the encoded video bitstream. The matching decoder follows the steps in inverse order.
[0751] Scalable video coding of a bitstream consists of different layers. The base layer proposes complete decodable video and enhancement layers that can be additionally utilized for decoding. The enhancement layers can provide high spatial resolution (spatial scalability), temporal resolution (temporal scalability), or quality (SNR scalability).
[0752] In previous standards such as H.264 / AVC SVC, syntactic elements such as motion vectors, reference picture indices, or intra prediction modes are predicted directly from the corresponding syntactic elements in the coded base layer.
[0753] A mechanism exists at the block level to switch between using prediction signals derived from enhancement layer samples decoded in the enhancement layer or other enhancement layer syntactic elements, or from base layer syntactic elements.
[0755] In the following section, base layer data is used to derive enhancement layer parameters from the decoder side.
[0757] Method 1: Derivation of Motion Parameter Candidates derivation
[0758] For a block (a) of a spatial or quality enhancement layer picture, a corresponding block (b) of a base layer picture is determined, and it covers the same picture area. An inter-prediction signal for the enhancement layer block (a) is formed using the following method:
[0759] 1. Candidate sets of motion compensation parameters are determined, for example, from temporal or spatial adjacent enhancement layer blocks or their derivatives.
[0760] 2. Motion compensation is formed for each candidate motion compensation parameter to form an inter-prediction signal in the enhancement layer.
[0761] 3. An optimal set of motion compensation parameters is selected by minimizing the error measurement between the predicted signal for the enhancement layer block (a) and the restored signal of the base layer block (b). For spatial scalability, the base layer block (b) can be spatially unsampled using an interpolation filter.
[0763] Motion compensation parameters include specific combinations of motion compensation parameters. Motion compensation parameters may be motion vectors, reference picture indices, a selection between uni- and bi-prediction, and other parameters. In an alternative embodiment, candidate sets of motion compensation parameters from base layer blocks are used. Inter-prediction is also performed in the base layer (using base layer reference pictures). To apply error measurement, the base layer block (b) reconstructed signal can be used directly without upsampling.
[0765] The selected optimal set of motion compensation parameters is applied to the enhancement layer reference pictures to form the prediction signal of block (a). When applying motion vectors in the spatial enhancement layer, the motion vectors are scaled according to the change in resolution. Both the encoder and the decoder can perform the same prediction steps to select the optimal set of motion compensation parameters from among the available candidates and generate the same prediction signals. These parameters are not signaled in the coded video bitstream.
[0767] The selection of the prediction method can be signaled in the bitstream and coded using entropy coding. Within a layered block subdivision structure, this coding method can alternatively be selected at all sub-levels or only on subsets of the coding layer. In an alternative embodiment, the encoder can transmit a refinement motion parameter set prediction signal to the decoder. The refinement signal contains differentially coded values of motion parameters. The refinement signal can be entropy coded.
[0769] In an alternative embodiment, the decoder generates a list of optimal candidates. Indices of the set of motion parameters used are signaled in the coded video bitstream. The indices can be entropy-coded. In an exemplary embodiment, the list can be sorted by increasing the error measure.
[0771] One exemplary embodiment uses HEVE's Adaptive Motion Vector Prediction (AMVP) candidate list to generate motion compensation parameter set candidates.
[0772] Another exemplary embodiment uses the HEVC merge mode candidate list to generate motion compensation parameter set candidates.
[0774] Method 2: Motion vector derivation
[0775] For a block (a) of a spatial or quality enhancement layer picture, a corresponding block (b) of a base layer picture is determined, and it covers the same picture area.
[0777] An inter-prediction signal for block (a) of the enhancement layer is formed using the following method:
[0778] 1. A motion vector predictor is selected.
[0779] 2. Motion measurement of a defined set of search locations is performed on the enhancement layer reference pictures.
[0780] 3. For each search location, an error measurement is determined, and the motion vector with the minimum error is selected.
[0781] 4. The prediction signal for block (a) is formed using the selected motion vector.
[0783] In an alternative embodiment, the search is performed on the restored base layer signal. The motion vector selected for spatial scalability is scaled according to the change in spatial resolution before generating the prediction signal in step 4.
[0785] The search locations may be at full or sub-pixel resolutions. The search may be performed in multiple steps, for example, by first determining an optimal full-pixel location followed by another set of candidates based on a selected full-pixel location. The search may be terminated early, for example, when the error measurement is below a defined threshold.
[0787] Both the encoder and the decoder can perform the same prediction steps to select the optimal motion vector among the candidates and generate the same prediction signals. These vectors are not signaled in the coded video bitstream.
[0789] The selection of the prediction method can be signaled within the bitstream and coded using entropy coding. Within a hierarchical block subdivision structure, this coding method can be alternatively selected at all sub-levels or only on subsets of the coding hierarchy.
[0790] In an alternative embodiment, the encoder may transmit a refinement motion vector prediction signal to the decoder. The refinement signal may be entropy-coded.
[0792] An exemplary embodiment uses the algorithm described in Method 1 to select a motion vector predictor.
[0794] Another exemplary embodiment is motion from temporally or spatially adjacent blocks of an enhancement layer. The Adaptive Motion Vector Prediction (AMVP) method of HEVC is used to select a vector predictor.
[0796] Method 3: Intra prediction mode derivation
[0797] For each block (a) of the enhancement layer (n) picture, a corresponding block (b) covering the same area in the restored base layer (n-1) picture is determined.
[0799] In the scalable video decoder for each base layer block (b), the intra prediction signal is formed using an intra prediction mode (p) inferred by the following algorithm.
[0800] 1) The intra prediction signal is generated for each available intra prediction mode following the rules for intra prediction of the enhancement layer, but using sample values from the base layer.
[0801] 2) Optimal prediction mode (p best ) is determined by minimizing the error measurement between the intra-predicted signal and the decoded base layer block (b) (e.g., the sum of absolute difference values).
[0802] 3) Prediction selected from Step 2) (p best The ) mode is used to generate a prediction signal for an enhancement layer block (a) by following intra-prediction rules for the enhancement layer.
[0804] Both the encoder and decoder form a matching prediction signal and the optimal prediction mode (p best Identical steps can be formed to select ). Actual intra prediction mode (p best ) is not signaled in a video bitstream coded in this way.
[0806] The selection of the prediction method can be signaled in the bitstream and coded using entropy coding. Within a hierarchical block subdivision structure, this coding mode can be selected only on a subset of the coding layer or on all sub-levels. An alternative embodiment uses samples from the enhancement layer in step 2) to generate an intra-prediction signal. For a spatially scalable enhancement layer, the base layer can be unsampled using an interpolation filter to apply an error measure.
[0808] An alternative embodiment has a smaller block size (a i Divide the enhancement layer block into multiple blocks of ) (e.g., 16x16 block (a) into 16 4x4 blocks (a i It can be divided into ). The algorithm described above applies to each sub-block (a i ) and corresponding base layer block (b i Applies to ). Prediction block (a i Residual coding is applied before ) and the above result is the prediction block (a i+1 It is used for ).
[0810] An alternative embodiment is a predicted intra-prediction mode (p best To determine ) surrounding sample values (b) or (b iUses ). For example, a 4x4 block (a) of the spatial enhancement layer (n). i ) corresponds to a 2x2 base layer block (b i When having ), (b i The surrounding samples of ) are the predicted intra prediction mode (p beat A 4x4 block (c) used to determine ) i It is used to form ).
[0812] In an alternative embodiment, the encoder may transmit an improved intra-prediction direction signal to the decoder. For example, in video codecs such as HEVC, most intra-prediction modes correspond to the angle at which boundary pixels are used to form the prediction signal. The offset for the optimal mode is the prediction mode (p) (determined as described above). best It can be transmitted as a difference for ). The enhancement mode can be entropy coded.
[0814] Intra-predicted modes are generally coded based on their probability. In H.264 / AVC, a single most probable mode is determined based on the modes used in the (spatial) adjacent parts of the block. These most probable modes can be selected using fewer symbols in the bitstream than required by the total number of modes. An alternative embodiment is an intra-predicted mode (p) predicted for block (a) (determined as described in the algorithm above) as the most probable mode or a member of the list of most probable modes. best Uses ).
[0816] Method 4: Intra prediction using boundary regions border areas)
[0817] In a scalable video decoder for forming an intra-prediction signal for a block (a) of a scalable or quality enhancement layer (see FIG. 45), lines of samples (b) from the surrounding region of the sample layer are used to fill the block region. These samples are taken from already coded regions (typically, but not necessarily at the top and left boundaries).
[0819] The following alternative variations of this pixel selection can be used:
[0820] a) If a pixel in the surrounding area has not yet been coded, the pixel value is not used to predict the current block.
[0821] b) If a pixel in the surrounding area has not yet been coded, the pixel value is derived from adjacent pixels that have already been coded (e.g., by iteration).
[0822] c) If a pixel in the surrounding area has not yet been coded, the pixel value is derived from a pixel in the corresponding area of the decoded base layer picture.
[0824] To form the intra prediction of block (a), the adjacent lines of pixels (b) (derived as described above) are each line (a) of block (a). j It is used as a template to fill in ).
[0826] Lines of block (a) (a j ) is filled stepwise along the x-axis. To achieve the optimal possible prediction signal, the rows of template samples (b) are the relevant lines (a j For ), the predicted signal (b' j It is shifted along the y-axis to form ).
[0828] To find the optimal prediction for each line, shift offset (o j) are the sample values of the corresponding line in the base layer and the result prediction signal (a j It is determined to minimize error measurement between ).
[0830] If (o j If ) is a non-integer value, the interpolation filter is (a as shown in (b`7) j It can be used to map the values of (b) to the integer sample locations of ).
[0832] When spatial scalability is utilized, the interpolation filter can be used to generate matching numbers of sample values of corresponding lines in the base layer.
[0834] The fill direction (x-axis) can be horizontal (left to right or right to left), vertical (top to bottom or bottom to top), diagonal, or any other angle. The samples used for the template line (b) are samples from the immediate adjacent parts of the block along the x-axis. The template line (b) is shifted along the y-axis, forming a 90° with respect to the x-axis.
[0836] To find the optimal direction of the x-axis, a full intra prediction signal is generated for block (a). An angle having the minimum error measurement between the prediction signal and the corresponding base layer block is selected.
[0837] The number of possible angles may be limited.
[0839] Both the encoder and the decoder run the same algorithm to determine the optimal predicted angles and offsets. It is not necessary for specific angle or offset information to be signaled in the bitstream. In an alternative embodiment, only samples of the base layer picture have offsets (o i It is used to determine ).
[0841] In an alternative embodiment, prediction offsets (o iThe enhancement (e.g., difference value) of ) is signaled in the bitstream. Entropy coding can be used to code the enhancement offset value.
[0843] In an alternative embodiment, the improvement of the prediction direction (e.g., difference value) is signaled in the bitstream. Entropy coding can be used to code the improvement direction value. An alternative embodiment is a line (b' j It uses a threshold to select whether ) is used for prediction. Optimal offset (o j If the error measurement of ) is below the threshold, line(c i ) is a block line(a i It is used to determine the values of ). Optimal offset (o j If the error measurement for ) is above the threshold, the (upsampled) base layer signal is block line (a j It is used to determine the values of ).
[0845] Method 5: Other prediction parameters
[0846] Other predictive information is inferred similarly to methods 1-3, for example, the division of blocks into sub-blocks.
[0848] For a block (a) of a spatial or quality enhancement layer picture, a corresponding block (b) of a base layer picture is determined, which covers the same picture area.
[0850] The prediction signal for block (a) of the enhancement layer is formed using the following method:
[0851] 1) A prediction signal is generated for each possible value of the tested parameter.
[0852] 2) Optimal prediction mode (p best ) is determined by minimizing the error measurement between the predicted signal and the decoded base layer block (b) (e.g., the sum of absolute differences)
[0853] 3) The prediction selected in Step 2) (p bestThe ) mode is used to generate a prediction signal for the enhancement layer block (a).
[0855] Both the encoder and the decoder can generate identical prediction signals and form identical prediction steps to select the optimal prediction mode from among possible candidates. The actual prediction mode is not signaled in the coded video bitstream.
[0857] The selection of the prediction mode is signaled in the bitstream and can be coded using entropy coding. Within a hierarchical block subdivision structure, this coding method can be selected at all sub-levels or, alternatively, only on a subset of the coding layer.
[0859] The following description briefly summarizes the above embodiments.
[0861] Generating an intra-prediction signal using restored base layer samples Enhancement layer coding with a number of methods
[0862] Key Perspective: Enhancement layer coding having multiple methods for generating intra-predictive signals using restored base layer samples to code blocks is provided in addition to methods for generating predictive signals based solely on restored enhancement layer samples.
[0863] Sub-perspectives ( Sub-aspects) :
[0864] Many methods include the following: the restored base layer signal (upsampled / filtered) is directly used as the enhancement layer prediction signal.
[0865] Many methods include the following: the restored base layer signal (upsampled / filtered) is combined with a spatial intra-prediction signal, where the spatial intra-prediction is derived based on difference samples for adjacent blocks.
[0866] · Difference samples represent the differences in the reconstructed enhancement layer signal and the (upsampled / filtered) reconstructed base layer signal (see Perspective A).
[0867] · Many methods include the following: a conventional spatial intra-prediction signal (derived using adjacent restoration enhancement layer samples) is combined with a base layer residual signal (upsampled / filtered) (the difference between base layer prediction and base layer restoration or the inverse transformation of base layer transformation coefficients) (see Aspect B).
[0868] · Many methods include the following: the (upsampled / filtered) reconstructed base layer signal is combined with a spatial intra prediction signal, wherein the spatial intra prediction is derived based on the reconstructed enhancement layer samples of adjacent blocks. The final prediction signal is obtained by weighting the base layer prediction signal and the spatial prediction signal in such a way that different frequency components use different weights (see Aspect C1). This is realized, for example, by any of the following:
[0869] o Filter the base layer prediction signal with a low-pass filter and the spatial intra prediction signal with a high-pass filter, and add the obtained filtered signals (Perspective C2)
[0870] o Transform the base layer prediction signal and the enhancement layer signal and superimpose the resulting transformation blocks, wherein different weighting factors are used for different frequency positions (see Aspect C3). The resulting transformation block can be inversely transformed and used as the enhancement layer prediction signal (see Aspect C4). The resulting transformation coefficients are scaled and added to the transmitted transformation coefficient levels, and then inversely transformed to obtain the restored block before deblocking and in-loop processing.
[0871] · Regarding methods using the restored base layer signal, the following versions may be used. This can be fixed or signaled at the sequence level, picture level, slice level, maximum coding unit level, or coding unit level. Alternatively, it can be generated depending on other coding parameters.
[0872] o Restored base layer samples before additional in-loop processing and deblocking (such as sample adaptive offset filters or adaptive loop filters)
[0873] o Base layer samples restored after deblocking before additional in-loop processing (such as sample adaptive offset filters or adaptive loop filters)
[0874] o Restored base layer samples after additional in-loop processing and deblocking (such as a sample adaptive offset filter or adaptive loop filter) or between multiple in-loop processing steps (see Aspect D)
[0875] Multiple versions of methods utilizing (upsampled / filtered) base layer signals may be used. The upsampled / filtered base layer signal used for these versions may differ in the interpolation filters used (including interpolation filters that also filter integer-sample positions), or the upsampled / filtered base layer signal for a second version may be obtained by filtering the base layer signal upsampled / filtered for a first version. The selection of one of the different versions may be signaled at the sequence, picture, slice, maximum coding unit, or coding unit level, or may be inferred from the transmitted coding parameters or the characteristics of the corresponding reconstructed base layer signal (see Aspect E).
[0876] Difference filters can be used to upsample / filter the reconstructed base layer signal (see Perspective E) and the base layer residual signal (see Perspective F).
[0877] · For base layer blocks with a residual signal of 0, they can be replaced with another signal derived from the base layer, for example, a high-pass filtered version of the restored base layer block (see Aspect G).
[0878] · For modes using spatial intra-prediction, adjacent samples that are not available in the enhancement layer can be replaced with corresponding samples of the upsampled / filtered base layer signal (due to the given coding order) (see Aspect H).
[0879] For modes utilizing spatial intra-prediction, the coding of the intra-prediction mode may be modified. The list of most probable modes includes the intra-prediction modes of the coexisting base layer signal.
[0880] · In certain versions, enhancement layer pictures are decoded in a two-stage process. In the first stage, blocks that use only the base layer signal (but not adjacent blocks) or the inter-prediction signal for prediction are decoded and restored. In the second stage, the remaining blocks that use adjacent samples for prediction are restored. For the blocks restored in the second stage, the spatial intra-prediction concept can be extended (see Perspective I). Based on the availability of already restored blocks, adjacent samples to the top and left, as well as samples adjacent to the bottom and right of the current block, can be used for spatial intra-prediction.
[0882] Multiple sources that generate inter-prediction signals using restored base layer samples Enhancement layer coding using methods
[0883] Key Perspective: In addition to a method for generating a prediction signal based solely on restored enhancement layer samples to code blocks in the enhancement layer, various methods are provided for generating an inter-prediction signal using restored base layer samples.
[0885] Sub-perspectives:
[0886] · Many methods include the following: A conventional inter-prediction signal (derived by motion-compensated interpolation of already restored enhancement layer pictures) is combined with a base layer residual signal (upsampled / filtered) (inverse transform of base layer transform coefficients or the difference between base layer restoration and base layer prediction).
[0887] · Many methods include the following: The (upsampled / filtered) restored base layer signal is combined with a motion-compensated prediction signal, where the motion-compensated prediction signal is obtained by motion compensating the difference pictures. The difference pictures represent the difference of the restored enhancement layer signal and the (upsampled / filtered) restored base layer signal relative to the reference pictures (see Perspective J).
[0888] · Many methods include the following: A restored base layer signal (upsampled / filtered) is combined with an inter-prediction signal, wherein the inter-prediction is derived by motion-compensated prediction using restored enhancement layer pictures. A final prediction signal is obtained by weighting the base layer prediction signal and the inter-prediction signal in such a way that different frequency components use different weights (see Aspect C). This can be realized, for example, by any of the following.
[0889] o Filter the base layer prediction signal with a low-pass filter and the spatial intra prediction signal with a high-pass filter, and add the obtained filtered signals.
[0890] o The base layer prediction signal and the enhancement layer signal are transformed and the resulting transformed blocks are superimposed, wherein different weighting factors are used for different frequency positions. The resulting transformed blocks can be inversely transformed and used as enhancement layer prediction signals, and the resulting transformed coefficients are scaled and added to the transmitted transformed coefficient levels, and then inversely transformed to obtain the restored blocks before deblocking and in-loop processing.
[0891] · Regarding methods using the restored base layer signal, the following versions may be used. This can be fixed or signaled at the sequence level, picture level, slice level, maximum coding unit level, or coding unit level. Alternatively, it can be generated depending on other coding parameters.
[0892] o Restored base layer samples before additional in-loop processing and deblocking (such as sample adaptive offset filters or adaptive loop filters)
[0893] o Base layer samples restored after deblocking before additional in-loop processing (such as sample adaptive offset filters or adaptive loop filters)
[0894] o Restored base layer samples after additional in-loop processing and deblocking (such as a sample adaptive offset filter or adaptive loop filter) or between multiple in-loop processing steps (see Aspect D)
[0895] · For base layer blocks with a residual signal of 0, they can be replaced with another signal derived from the base layer, for example, a high-pass filtered version of the restored base layer block (see Aspect G).
[0896] Multiple versions of methods utilizing (upsampled / filtered) base layer signals may be used. The upsampled / filtered base layer signal used for these versions may differ in the interpolation filters used (including interpolation filters that also filter integer-sample positions), or the upsampled / filtered base layer signal for a second version may be obtained by filtering the base layer signal upsampled / filtered for a first version. The selection of one of the different versions may be signaled at the sequence, picture, slice, maximum coding unit, or coding unit level, or may be inferred from the transmitted coding parameters or the characteristics of the corresponding reconstructed base layer signal (see Aspect E).
[0897] Difference filters can be used to upsample / filter the reconstructed base layer signal (see Perspective E) and the base layer residual signal (see Perspective F).
[0898] · For motion-compensation prediction of difference pictures (the difference between the enhancement layer restoration and the upsampled / filtered base layer residual signal) (perspective J), difference interpolation filters are used rather than for motion-compensation prediction of restored pictures.
[0899] · For motion-compensated prediction of difference pictures (the difference between the enhancement layer restoration and the upsampled / filtered base layer signal) (see Perspective J), interpolation filters are selected based on the corresponding region characteristics of the difference pictures (or based on coding parameters or based on information transmitted in the bitstream).
[0901] Enhancement layer motion parameter coding coding)
[0902] Key Perspective: Utilizing at least one predictor derived from the base layer and multiple enhancement layer predictors for enhancement layer motion parameter coding.
[0904] Sub-perspective:
[0905] · Addition of (scaled) base layer motion vectors to the motion vector predictor list (Perspective K)
[0906] o Use a base layer block covering coexisting samples at the center position of the current block (other derivations possible)
[0907] o Scaled motion vectors based on resolution ratio
[0908] · Add motion data of coexisting base layer blocks to the merge candidate list (refer to Aspect K)
[0909] o Use a base layer block that covers the coexisting blocks at the center position of the current block (other derivations possible)
[0910] o Scaled motion vectors based on resolution ratio
[0911] o Do not add if merge_flag is equal to 1 in the base layer
[0912] · Rearrange the merge candidate list based on base layer merge information (refer to Perspective L)
[0913] If a coexisting base layer block is merged with a specific candidate, the corresponding enhancement layer candidate is used as the first entry (first entry) in the enhancement layer merge candidate list.
[0914] · Rearrange the list of motion predictor candidates based on base layer motion predictor information (refer to Aspect L)
[0915] o When coexisting base layer blocks use a specific motion vector predictor, the corresponding enhancement layer motion vector predictor is used as the first input in the enhancement layer motion vector predictor candidate list.
[0916] · Derivation of a merge index based on base layer information from coexisting blocks (i.e., candidates for merging the current block) (see Aspect M). For example, if a base layer block is merged into a specific adjacent block and this is signaled within the bitstream where an enhancement layer block is merged, the merge index is not transmitted; instead, the enhancement layer block is merged into the same adjacent block (but in the enhancement layer) based on the coexisting base layer block.
[0918] Enhancement layer partitioning and motion parameter inference (Enhancement layer partitioning and motion parameter inference)
[0919] Main Perspective: Enhancement layer partitioning and inference of motion parameters based on motion parameters and base layer partitioning (presumably required to combine this perspective with any of the sub-perspectives).
[0921] Sub-perspectives:
[0922] · Derivation of motion parameters for NxM sub-blocks of the enhancement layer based on coexisting base layer motion data; summarization of blocks with identical derived parameters (or parameters with small differences) into larger blocks; determination of prediction and coding units (refer to Perspective T).
[0923] · Motion parameters may include the following: multiple motion hypotheses, reference indices, motion vectors, motion vector predictor identifiers, merge identifiers.
[0924] · Signals one of a plurality of methods that generate an enhancement layer prediction signal; such methods may include the following:
[0925] o Motion compensation using derived motion parameters and restored enhancement layer reference pictures
[0926] o (a) restoration of the (upsampled / filtered) base layer for the current picture and (b) combination of the enhancement layer reference picture and the motion compensation signal using derived motion parameters resulting from subtracting the (upsampled / filtered) base layer restoration from the restored enhancement layer picture
[0927] o (a) base layer residue (upsampled / filtered) for the current picture (the difference between the restored signal and the prediction and the inverse transform of the coded transform coefficient values) and (b) motion compensation signal using the restored enhancement layer reference pictures and derived motion parameters
[0928] · When coexisting blocks in the base layer are intra-coded, the correspondence enhancement layer MxN blocks (or CUs) are also intra-coded, where the intra-predicted signal is derived using base layer information (see Perspective U), for example as follows:
[0929] The (upsampled / filtered) version of the corresponding base layer reconstruction is used as the intra-prediction signal (see Perspective U).
[0930] The intra prediction mode is derived based on the intra prediction mode used in the base layer, and the intra prediction mode is used for spatial intra prediction in the enhancement layer.
[0931] · For an MxN enhancement layer block (sub-block), if a coexisting base layer is merged with a previously coded base layer block (or has the same motion parameters), the MxN enhancement layer (sub-)block is also merged with the enhancement layer block corresponding to the base layer block used to be merged into the base layer (i.e., the motion parameters are duplicated from the corresponding enhancement layer block) (see Aspect M).
[0933] Coding of transformation coefficient levels / Context modeling coefficient levels / context modeling)
[0935] Key Perspectives: Transform factor coding using different scan patterns. Context modeling based on different initializations, coding modes, and / or base layer data for the enhancement layer and context models.
[0937] Sub-perspectives:
[0938] · Introduce one or more additional scan patterns, e.g., horizontal and vertical scan patterns. Instead of 4x4 subblocks, 16x1 or 1x16 subblocks may be used, or 8x2 and 8x2 subblocks may be used. Additional scan patterns may be introduced only for blocks of a specific size equal to or larger than a specific size, e.g., 8x8 or 16x16 (Perspective V).
[0939] · (If the coding block flag is equal to 1) the selected scan pattern is signaled within the bitstreams (see Aspect N). To signal corresponding syntax elements, a fixed context may be used. Alternatively, context derivation for corresponding syntax elements may depend on any of the following:
[0940] o Residual or coexisting slope of the restored base layer. Or edges detected in the base layer signal.
[0941] o Distribution of transformation coefficients in coexisting base layer blocks.
[0942] · The selected scan above can be derived directly from the base layer signal (without any additional data being transmitted) based on the characteristics of the coexisting base layer signal (see Aspect N).
[0943] o Gradient of the coexisting restored base layer signal or restored base layer residue. Or edges detected in the base layer signal.
[0944] o Distribution of transformation coefficients in coexisting base layer blocks
[0945] · Transform scans can be realized in a manner where transform coefficients are rearranged after quantization in the encoder and conventional coding is used. In terms of the decoder, transform coefficients are decoded as conventionally and rearranged before scaling and inverse transformation (or after scaling and before inverse transformation).
[0946] · To code importance flags (importance flags for single transformation coefficients and / or sub-group flags), the following modifications may be used in the enhancement layer:
[0947] Individual context models are used for all or a subset of coding modes that utilize base layer information. It is also possible to use different context models for modes that differ from the base layer information.
[0948] o Context modeling can rely on data in coexisting base layer blocks (e.g., the number of important transform coefficients for specific frequency positions) (see Aspect O).
[0949] o A generalized template can be used by measuring both the number of important transform coefficients in coexisting base layer signals at similar frequency locations and the number of already coded transform coefficient levels in the spatial proximity of the coefficients to be coded (see Aspect O).
[0950] · To code the final critical scanning locations, the following modifications can be used in the enhancement layer:
[0951] Individual context models are used for all or a subset of coding modes that utilize base layer information. It is also possible to use different context models for different modes that have base layer information (see Perspective P).
[0952] Context modeling can rely on data from coexisting base layer blocks (transform coefficient distribution in the base layer, gradient information in the base layer, final scanning position in coexisting base layer blocks).
[0953] The final scanning position can be coded according to the difference from the final base layer scanning position (see Perspective S).
[0954] · Utilization of different context initialization tables for base and enhancement layers
[0956] Backward adaptive enhancement layer coding using base layer data (Backward adaptive layer enhancement coding using base layer data)
[0957] Key Perspective: Utilization of base layer data to derive enhancement layer coding parameters.
[0959] Sub-perspective:
[0960] · Deriving merge candidates based on (potentially unsampled) base layer restoration. In the enhancement layer, only the use of merging is signaled, but the actual candidate used to merge the current block is derived based on the restored base layer signal. Thus, for all merge candidates, an error measurement between the corresponding predicted signals (derived using motion parameters for the merge candidates) and the (potentially upsampled) base layer signals for the current enhancement layer block is measured for all merge candidates (or a subset thereof), and the merge candidate associated with the minimum error measurement is selected. The calculation of the error measurement may also be performed in the base layer using base layer reference pictures and the restored base layer signal (see Aspect Q).
[0961] · Derivation of motion vectors based on the restoration of the (potentially upsampled) base layer. Motion vector differences are not coded but are inferred based on the restored base layer. Determine a motion vector predictor for the current block and measure a defined set of search locations around the motion vector predictor. For each search location, determine the error measurement between the replaced reference frame (the replacement is given by the search location) and the (potentially upsampled) base layer signal for the current enhancement layer block. Select the search location / motion vector that yields the minimum error measurement. The search can be divided into several steps. For example, a full-pixel search is performed first, followed by a half-pixel search around the optimal full-pixel vector, and then a quarter-pixel search around the optimal full / half-pixel vector. The search may be performed at the base layer using base layer reference pictures and the restored base layer signal, and the discovered motion vectors are subsequently scaled according to the resolution change between the base and enhancement layers. (Perspective Q)
[0962] Intra prediction modes are derived based on the (potentially upsampled) base layer reconstruction. The intra prediction modes are not coded but are inferred based on the reconstructed base layer. For each possible intra prediction mode (or a subset thereof), an error measurement is determined between the intra prediction signal and the (potentially upsampled) base layer signal for the current enhancement layer block (using the tested prediction mode). The prediction mode that yields the minimum error measurement is selected. The calculation of the error measurement can be performed at the base layer using the intra prediction signal and the reconstructed base layer signal. Furthermore, the intra block can be implicitly decomposed into 4x4 blocks (or other block sizes), and an individual prediction mode can be determined for each 4x4 block (see Aspect Q).
[0963] · The intra-predicted signal can be determined by line- or column-wise matching of boundary samples with the reconstructed base layer signal. To induce a shift between adjacent samples and the current line / column, an error measure is calculated between the reconstructed base layer signal and the shifted line / column of adjacent samples, and the shift that yields the minimum error measure is selected. Adjacent samples, (upsampled) base layer samples, or enhancement layer samples may be used. The error measure may also be calculated directly in the base layer (see Aspect W).
[0964] · Use of backward-adaptation methods for deriving other coding parameters, such as block partitioning, etc.
[0966] A further brief summary of the above embodiments is proposed below. In particular, the above embodiments are described.
[0968] A1) Scalable video decoder configured according to the following.
[0969] Recover base layer signals (200a, 200b, 200c) from the coded data stream (6) (80),
[0970] It restores (60) the enhancement layer signal (360), which is
[0971] The restored base layer signals (200a, 200b, 200c) are subjected to resolution or quality improvement to obtain an inter-layer prediction signal (380) (220), and
[0972] Calculate the difference signal between the inter-layer prediction signal (380) and the already restored portion (400a or 400b) of the enhancement layer signal (260);
[0973] To obtain a spatial intra-predicted signal, the difference signal is spatially predicted (260) from the first part (400, see FIG. 46) coexisting with the part of the enhancement layer signal (360) currently to be restored, from the second part (460) of the difference signal which is spatially adjacent to the first part belonging to the already restored part of the enhancement layer signal (360), and
[0974] It includes combining (260) spatial intra-prediction signals and inter-layer prediction signals (380) to obtain an enhancement layer prediction signal (420), and predictively restoring (320, 580, 340, 300, 280) an enhancement layer signal (360) using the enhancement layer prediction signal (420).
[0975] According to perspective A1, for example, as long as the base layer residual signal (640 / 480) is involved, the base layer signal can be restored by the base layer decoding stage (80) from the coded data stream (6) or substream (6a), respectively, in the block-based prediction method described above, but other restoration alternatives are also feasible.
[0976] As long as the restoration of the enhancement layer signal (360) by the enhancement layer decoding stage (60) is involved, the resolution or quality enhancement to which the base layer signal (200a, 200b, or 200c) is targeted may include, for example, tone-mapping from n bits to m bits having m > n in the case of bit depth enhancement or copying in the case of quality enhancement, or up-sampling in the case of resolution enhancement.
[0977] The calculation of the difference signal can be performed pixel by pixel, that is, on one hand, the coexisting pixels of the enhancement layer signal and on the other hand, the prediction signal (380) are subtracted from each other, and this is performed per pixel position.
[0978] Spatial prediction of the difference signal can be performed in any way, such as copying / interpolating already restored pixels that delineate the portion of the enhancement layer signal (360) to be restored, along the intra-prediction direction for the current portion of the enhancement layer signal, such as intra-prediction parameters transmitted within the substream (6b) or in the coded data stream (6). The combination may include much more complex combinations or weighted sums, sums, such as combinations that differently weight the distribution in the frequency domain.
[0979] Predictive restoration of an enhancement layer signal (360) using an enhancement layer prediction signal (420) may include, as shown in the drawing, entropy decoding and inverse transformation of an enhancement layer residual signal (540) and a combination (340) of the latter having the enhancement layer signal (420).
[0981] B1) The scalable video decoder is configured as follows.
[0982] The enhancement layer signal (360) is restored, and the base layer residual signal (480) is decoded (100) from the coded data stream (6), and this
[0983] To obtain an inter-layer residual prediction signal (380), the restored base layer residual signal (480) is subjected to (220) for resolution or quality improvement, and
[0984] To obtain an internal prediction signal of the enhancement layer, spatially predict (260) the portion of the enhancement layer signal (360) currently to be restored from the already restored portion of the enhancement layer signal (360);
[0985] To obtain an enhancement layer prediction signal (420), the enhancement layer internal prediction signal and the inter-layer residual prediction signal are combined (260); and
[0986] It includes predictively restoring (340) the enhancement layer signal (360) using the enhancement layer prediction signal (420).
[0987] As shown in the drawing, decoding of the base layer residual signal from the coded data stream can be performed using entropy decoding and inverse transformation. Furthermore, the scalable video decoder optionally performs the restoration of the base layer signal itself, that is, by deriving the base layer prediction signal (660), predictively decoding it, and combining the base layer residual signal (480) in the same way. As just mentioned, this is merely optional.
[0988] As far as the restoration of the enhancement layer signal is concerned, resolution or quality enhancement may be performed as indicated above in relation to A). As far as the spatial prediction of a portion of the enhancement layer signal is concerned, spatial prediction may be performed as exemplarily described in A) regarding the difference signal. Similar notes are valid insofar as combination and prediction restoration are involved.
[0989] However, it should be noted that in view B, the base layer residual signal (480) is not limited to being identical to the explicitly signaled version of the base layer residual signal (480). Rather, it may be possible for a scalable video decoder to subtract any restored base layer signal version (200) having the base layer prediction signal (660) and thereby obtain a base layer residual signal (480) that may show deviations from the explicitly signaled one due to deviations arising from filter functions (functions) such as filters (120) or (140). The latter reference is also valid in other views where the base layer residual signal is included in the inter-layer prediction.
[0991] C1) The scalable video decoder is configured as follows.
[0992] Recover (80) base layer signals (200a, 200b; 200c) from a coded data stream (6), and
[0993] It restores the enhancement layer signal (360), which is
[0994] To obtain an inter-layer prediction signal (380), the restored base layer signal (200) for resolution or quality improvement is targeted (220), and
[0995] To obtain an internal prediction signal of the enhancement layer, a portion of the enhancement layer signal (360) to be currently restored is predicted spatially or temporally (260) from the already restored portion of the enhancement layer signal (400a,b in the case of "spatial", 400a,b, c in the case of "temporal"), and
[0996] To obtain an enhancement layer prediction signal (420) in which the weights contributing to the enhancement layer prediction signal (420) by the inter-layer prediction signal and the enhancement layer internal prediction signal (380) vary across different spatial frequency components, the weighted average of the enhancement layer internal prediction signal (380) and the inter-layer prediction signal is formed (260) in the part currently to be restored; and
[0997] It includes predictively restoring (320, 340) the enhancement layer signal (360) using the enhancement layer prediction signal (420).
[0999] C2) In the current part to be restored, the formation of a weighted average (260) includes filtering the inter-layer prediction signal (380) with a low-pass filter to obtain the filtered signals in the current part to be restored (260), filtering the enhancement layer internal prediction signal with a high-pass filter (260), and adding the obtained filtered signals.
[1001] C3) In the part currently to be restored, the formation of the weighted average includes forming an inter-layer prediction signal and an enhancement layer internal prediction signal to obtain transformation coefficients (260); superimposing the transformation coefficients obtained using different weighting factors for different spatial frequency components to obtain superimposed transformation coefficients (260); and inversely transforming the superimposed transformation coefficients to obtain an enhancement layer prediction signal.
[1003] C4) Predictive restoration (320, 340) of an enhancement layer signal using an enhancement layer prediction signal (420) includes extracting (320) transformation coefficient levels for the enhancement layer signal from a coded data stream (6), performing a sum of the superimposed transformation coefficients and transformation coefficient levels to obtain a transformed version of the enhancement layer signal (340), and targeting the transformed version of the enhancement layer signal for an inverse transformation to obtain an enhancement layer signal (360) (i.e., inverse transformation T in the drawing -1 For the coding mode, it is located at least downstream of the adder (adder, 340).
[1004] With respect to the restoration of base layer signals, the above description is referenced as with respect to A) and B) and generally to drawings. The same applies not only to spatial prediction but also to the resolution or quality improvement mentioned in C.
[1005] The temporal prediction mentioned in C may each include a prediction provider (160) that derives motion prediction parameters from a coded data stream (6) and a substream (6a). The motion parameters may include: a motion vector, a reference frame index, or a combination of motion vectors per sub-block of the currently restored portion and motion subdivision information.
[1006] As previously described, the formation of the weighted average can be terminated in a spatial domain or a tra...
Claims
Claim 1 A video decoder comprising a processor for decoding a video represented by a base layer signal and an enhancement layer signal, wherein the video decoder comprises: a first decoding unit comprising a base layer decoder configured to restore a base layer signal based on a base layer residual signal of a coded data stream using the processor; and includes a second decoding unit comprising an enhancement layer decoder configured to use the processor to decode at least a predetermined block of the block and restore the enhancement layer signal in block units from the coded data stream based on syntactic elements of the coded data stream, wherein the decoding comprises: generating a set of possible subblock subdivisions including all possible subblock subdivisions for the predetermined block - each possible subblock subdivision corresponds to a possible method for subdividing the predetermined block of the enhancement layer signal into subblocks -, selecting a set of eligible subblock subdivisions from the set of possible subblock subdivisions for the predetermined block - at least one eligible subblock subdivision satisfies a similarity criterion -, for the predetermined block, selecting a subblock subdivision from the set of eligible subblock subdivisions - the predetermined block is subdivided into subblocks according to the selected subblock subdivision -, and using the selected subblock subdivision to the predetermined A decoder that includes predictive block restoration. Claim 2 A decoder according to claim 1, wherein the selection comprises detecting one or more edges within the base layer residual signal or the coexistence portion of the base layer signal. Claim 3 In claim 1, the predetermined block is a conversion factor block having a conversion factor representing the enhancement layer signal; and the decoding further comprises, for a current subblock being traversed, decoding in the coded data stream (a) a first syntax element indicating whether the current subblock contains any valid conversion factor, and (b) a second syntax element indicating a conversion factor level within the current subblock when the first syntax element indicates that the current subblock contains a valid conversion factor. Claim 4 In paragraph 3, the second decoding unit comprises: an inverse converter configured to perform an inverse conversion on the transform coefficients of the transform coefficient block to obtain an enhancement layer residual signal representing a predicted residual of a predicted signal for the enhancement layer signal; and a predictive decoder configured to spatially, temporally, and / or inter-layer predict the enhancement layer signal to obtain the predicted signal for the enhancement layer signal, and to apply the enhancement layer residual signal to the predicted signal for the enhancement layer signal to restore the enhancement layer signal. Claim 5 In paragraph 3, the second decoding unit is configured to: apply a transformation to the base layer residual signal or the base layer signal from the spatial domain to the frequency domain; and combine and scale a transformation coefficient block of the base layer residual signal to form a spectral decomposition of the coexisting portion of the base layer residual signal or the base layer signal. Claim 6 In claim 1, the base layer signal is restored using base layer coding parameters that vary spatially across the base layer signal; the selected sub-block subdivision is the coarsest sub-block subdivision among the set of possible sub-block subdivisions, which subdivides the base layer signal into regions when transmitted to the coexisting portion of the base layer signal so that the base layer coding parameters are sufficiently similar to each other within each region, a decoder. Claim 7 In claim 6, the decoder further comprises: for the predetermined block, predicting an enhancement layer coding parameter based on the coexistence of the base layer coding parameter with the predetermined block; and predictively restoring the predetermined block using the enhancement layer coding parameter. Claim 8 A non-transient computer-readable medium for storing data related to video, comprising a data stream stored in the non-transient computer-readable medium, wherein the data stream comprises information related to encoding represented by a base layer signal and an enhancement layer signal, and wherein the data stream comprises: an operation of restoring a base layer signal based on a base layer residual signal from the data stream using a processor; A medium comprising: decoding using a plurality of operations including decoding at least a predetermined block of the said block and restoring the enhancement layer signal from the said data stream in block units based on syntactic elements of the said data stream using the said processor, wherein the decoding comprises: generating a set of possible subblock subdivisions including all possible subblock subdivisions for the said predetermined block - each possible subblock subdivision corresponds to a possible method for subdividing the said predetermined block of the enhancement layer signal into subblocks - , selecting a set of eligible subblock subdivisions from the set of possible subblock subdivisions for the said predetermined block - at least one eligible subblock subdivision such that the coding parameters of the coexisting portion of the said base layer signal satisfy a similarity criterion - , for the said predetermined block, selecting a subblock subdivision from the set of eligible subblock subdivisions - the said predetermined block is subdivided into subblocks according to the selected subblock subdivision - , and predictively restoring the said predetermined block using the selected subblock subdivision. Claim 9 In paragraph 8, the medium comprises the selection of detecting one or more edges within the base layer residual signal or the coexistence portion of the base layer signal. Claim 10 In claim 8, the predetermined block is a transformation factor block having a transformation factor representing the enhancement layer signal; and the decoding further comprises, for a current subblock being traversed, decoding in the data stream (a) a first syntactic element indicating whether the current subblock contains any effective transformation factor, and (b) a second syntactic element indicating a transformation factor level within the current subblock if the first syntactic element indicates that the current subblock contains an effective transformation factor. Claim 11 A medium according to claim 10, wherein the plurality of operations further comprises: an operation of performing an inverse transformation on the transformation coefficients of the transformation coefficient block to obtain an enhancement layer residual signal representing the predicted residual of the prediction signal for the enhancement layer signal; and an operation of spatially, temporally, and / or inter-layer predicting the enhancement layer signal to obtain the prediction signal for the enhancement layer signal, and applying the enhancement layer residual signal to the prediction signal for the enhancement layer signal to restore the enhancement layer signal. Claim 12 A medium according to claim 10, wherein the plurality of operations further comprises: applying a transformation to the base layer residual signal or the base layer signal from the spatial domain to the frequency domain, combining and scaling the transformation coefficient blocks of the base layer residual signal to form a spectral decomposition of the coexisting portion of the base layer residual signal or the base layer signal. Claim 13 In paragraph 8, the base layer signal is restored using base layer coding parameters that vary spatially across the base layer signal; the selected sub-block subdivision is the coarsest sub-block subdivision among the set of possible sub-block subdivisions, and this divides the base layer signal into regions when transmitted to a co-extended portion of the base layer signal such that the base layer coding parameters are sufficiently similar to each other within each region, the medium. Claim 14 A medium according to claim 13, wherein the decoding further comprises: for the predetermined block, predicting an enhancement layer coding parameter based on the fact that the base layer coding parameter is co-matched with the predetermined block; and predictively restoring the predetermined block using the enhancement layer coding parameter. Claim 15 A video encoder comprising a processor for encoding a video, wherein the video encoder comprises: a first encoding unit comprising a base layer encoder configured to use the processor to determine a base layer residual signal for a base layer of the video and to encode a base layer signal into a data stream based on the base layer residual signal; The encoder comprises a second encoding unit including an enhancement layer encoder configured to use the processor to encode at least a predetermined block of the block to encode syntactic elements and enhancement layer signals into the data stream in blocks, wherein the encoding comprises: generating a set of possible sub-block subdivisions including all possible sub-block subdivisions for the predetermined block - each possible sub-block subdivision corresponds to a possible method for subdividing the predetermined block of the enhancement layer signal into sub-blocks - selecting a set of eligible sub-block subdivisions from the set of possible sub-block subdivisions for the predetermined block - at least one eligible sub-block subdivision ensures that the coding parameters of the coexisting portion of the base layer signal satisfy a similarity criterion - selecting a sub-block subdivision from the set of eligible sub-block subdivisions for the predetermined block - the predetermined block is subdivided into sub-blocks according to the selected sub-block subdivision - and constructing the predetermined block using the selected sub-block subdivision. Claim 16 In paragraph 15, the encoder comprises the selection of detecting one or more edges within the base layer residual signal or the coexistence portion of the base layer signal. Claim 17 In claim 15, the predetermined block is a transform factor block having a transform factor representing the enhancement layer signal; and the encoding further comprises the step of encoding, for a current subblock being traversed, (a) a first syntactic element indicating whether the current subblock contains any effective transform factor, and (b) a second syntactic element indicating a transform factor level within the current subblock when the first syntactic element indicates that the current subblock contains an effective transform factor, into the data stream. Claim 18 In claim 17, the second encoding unit comprises: a converter configured to perform a conversion on an enhancement layer residual signal to obtain the conversion coefficients of the conversion coefficient block, wherein the enhancement layer residual signal represents the predicted residual of a prediction signal for the enhancement layer signal; and a prediction encoder configured to encode the enhancement layer signal by spatially, temporally, and / or inter-layer predicting the enhancement layer signal to obtain the prediction signal for the enhancement layer signal. Claim 19 In claim 17, the second encoding unit is configured to: apply a transformation to the base layer residual signal or the base layer signal from the spatial domain to the frequency domain; and combine and scale a transformation coefficient block of the base layer residual signal to form a spectral decomposition of the coexisting portion of the base layer residual signal or the base layer signal. Claim 20 In paragraph 15, the base layer signal is encoded using base layer coding parameters that vary spatially across the base layer signal; the selected sub-block subdivision is the coarsest sub-block subdivision among the set of possible sub-block subdivisions, which subdivides the base layer signal into regions when transmitted to the coexisting portion of the base layer signal so that the base layer coding parameters are sufficiently similar to each other within each region, an encoder. Claim 21 In claim 20, the encoder further comprises: for the predetermined block, predicting an enhancement layer coding parameter based on the coexistence of the base layer coding parameter with the predetermined block; and predictively constructing the predetermined block using the enhancement layer coding parameter. Claim 22 A decoder according to claim 1, wherein the sub-block subdivision is selected from a set of possible sub-block subdivisions based on the spatial variation of a base layer coding parameter within at least one coexisting portion of the base layer residual signal or the base layer signal. Claim 23 In paragraph 15, the encoder, wherein the sub-block subdivision is selected from a set of possible sub-block subdivisions based on the spatial variation of a base layer coding parameter within at least one coexisting portion of the base layer residual signal or the base layer signal. Claim 24 A method for decoding video represented by a base layer signal and an enhancement layer signal, wherein the method comprises: a step of restoring the base layer signal based on a base layer residual signal of a coded data stream; and a step of decoding at least a predetermined block among the blocks to restore the enhancement layer signal in block units from the coded data stream based on syntactic elements of the coded data stream, wherein the decoding comprises: a step of generating a set of possible sub-block subdivisions including all possible sub-block subdivisions for the predetermined block - each possible sub-block subdivision corresponds to a possible method for subdividing the predetermined block of the enhancement layer signal into sub-blocks - a step of selecting a set of eligible sub-block subdivisions from the set of possible sub-block subdivisions for the predetermined block - at least one eligible sub-block subdivision such that the coding parameters of the coexisting portion of the base layer signal satisfy a similarity criterion - and, for the predetermined block, a step of selecting a sub-block subdivision from the set of eligible sub-block subdivisions - the predetermined block according to the selected sub-block subdivision A method comprising the step of subdividing into subblocks, and predictively restoring the predetermined block using the selected subblock subdivision. Claim 25 A method for encoding a video comprises: a step of determining a base layer residual signal for a base layer of the video; a step of encoding a base layer signal into a data stream based on the base layer residual signal; and a step of encoding at least a predetermined block of the block to encode syntactic elements and enhancement layer signals into a data stream in block units, wherein the encoding comprises: a step of generating a set of possible sub-block subdivisions including all possible sub-block subdivisions for the predetermined block - each possible sub-block subdivision corresponds to a possible method for subdividing the predetermined block of the enhancement layer signal into sub-blocks -, a step of selecting a set of eligible sub-block subdivisions from the set of possible sub-block subdivisions for the predetermined block - at least one eligible sub-block subdivision such that the coding parameters of the coexisting portion of the base layer signal satisfy a similarity criterion -, a step of selecting a sub-block subdivision from the set of eligible sub-block subdivisions for the predetermined block - the predetermined block is subdivided into sub-blocks according to the selected sub-block subdivision -, and the selected sub-block A method comprising the step of configuring the predetermined block using subdivision.
Citation Information
Patent Citations
Scalable encoding method and device, scalable decoding method and device and these program and their recording media
JP2007028034A
Video coding with fine granularity spatial scalability
KR1020080094041A
Methods and apparatus for artifact removal for bit depth scalability
KR1020100081348A
Scalable video coding using derivation of subblock subdivision for prediction from base layer
KR102447521B1