Low-latency concepts in multi-layer video coding

JP2024012407A5Pending Publication Date: 2026-03-13GE VIDEO COMPRESSION LLC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-10-27
Publication Date
2026-03-13

AI Technical Summary

Technical Problem

Current multi-view/layer coding technologies face challenges in achieving low latency without compromising the ability of decoders to handle the coding concept, particularly in scenarios where ultra-low latency is required.

Method used

The implementation of an interleaved multi-layer video data stream with additional timing control information that allows for interleaved or deinterleaved decoding units, enabling low latency operations while maintaining decoder compatibility.

Benefits of technology

This approach reduces end-to-end delay in multi-layer video encoding by allowing parallel processing of different layers, ensuring efficient decoding without requiring additional hardware or complex reordering mechanisms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

To provide a method for encoding and decoding into a multi-layer video data stream.SOLUTION: An encoder 720 spreads second timing information 802 into several timing control packets, each timing control packet indicating a second decoder search buffer time for a preceding decoding unit prior to the decoding unit with which each timing control packet is associated. The encoder 720 also reacts a change in spatial complexity in various layers of pictures 12 and 15 between coding to output second timing control information 802 while the current instantaneous layer is being encoded. The encoder 720 further evaluates first timing control information 800 prior to encoding the current moment and location layer and first timing control information 800 at the beginning or end of each access unit for each access unit.SELECTED DRAWING: Figure 17
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] This application relates to coding concepts that enable efficient multi-view / layer such as multi-view picture / video coding. [Background technology]

[0002] The Coded Picture Buffer (CPB) in Scalable Video Coding (SVC) works with complete Access Units (AUs). One AU for every Network Abstraction Layer Unit (NALU) is moved from the Coded Picture Buffer (CPB) at the same moment. One AU contains all layer packets (i.e., NALUs).

[0003] The concept of a Decoding Unit (DU) in the HEVC base specification [1] is expanded compared to H.264 / AVC. A DU is a group of NAL units at consecutive positions in the bitstream that belong to the same layer, i.e. the so-called base layer.

[0004] The HEVC base specification includes the necessary tools to enable decoding of bitstreams with ultra-low delay, i.e., by means of CPB operation at DU level as opposed to CPB operation at AU level as in H.264 / AVC, and CPB timing information with DU granularity. Thus, a device can operate on sub-portions of a picture to reduce the processing delay incurred. Just as ultra-low delay operates in multi-layer SHVC, MV-HEVC, and the HEVC extension 3D-HEVC, CPB operates at DU level across multiple layers and needs to be properly defined. In particular, bitstreams using several layers or views are required in which DUs of one AU are interleaved across multiple layers. That is, DUs of layer m of one given AU can follow DUs of layer (m+1) of the same AU in such an ultra-low delay enabled multi-layer bitstream, as long as they are independent of the following DUs in the bitstream order.

[0005] Ultra-low delay operation requires modifications of the CPB operation for multi-layer decoders compared to SVC and the MVC of H.264 / AVC extensions, which work on the basis of AUs. Ultra-low delay decoders may use additional timing information, for example provided by way of SEI messages.

[0006] Some implementations of one multi-layer decoder may prefer layer-wise decoding (and CPB operation at either DU or AU level), i.e., decoding of layer m before decoding of layer m+1, which would effectively prevent any multi-layer ultra-low latency applications using SHVC, MV-HEVC, and 3D-HEVC unless new mechanisms are provided.

[0007] Currently, the HEVC Base Specification includes two decoding operation modes. Access Unit (AU) based decoding: All decoding units of an access unit are moved out of the CPB at the same time. Decoding Unit (DU) based decoding: Each decoding unit has its own CPB removal time. [Prior art documents] [Non-patent literature]

[0008] [Non-Patent Document 1] B. Bross, W.-J. Han, J.-R. Ohm, GJ Sullivan, T. Wiegand (Eds.), "High Efficiency Video Coding (HEVC) text specification draft 10", JCTVC-L1003, Geneva, CH, Jan. 2013 [Non-Patent Document 2] G. Tech, K. Wegner, Y. Chen, M. Hannuksela, J.Boyce (Eds.), "MV-HEVC Draft Text 3 (ISO / IEC 23008-2 PDAM2)", JCT3V-C1004, Geneva, CH, Jan. 2013 [Non-Patent Document 3] G. Tech, K. Wegner, Y. Chen, S. Yea (Eds.), "3D-HEVC Test Model Description, draft specification", JCT3V-C1005, Geneva, CH, Jan. 2013 [Non-Patent Document 4] WILBURN, Bennett, et al. High performance imaging using large camera arrays. ACM Transactions on Graphics, 2005, 24. Jg., Nr. 3, S. 765-776. [Non-Patent Document 5] WILBURN, Bennett S., et al. Light field video camera. In:Electronic Imaging 2002. International Society for Optics and Photonics, 2001. S. 29-36. [Non-Patent Document 6] HORIMAI, Hideyoshi, et al. Full-color 3D display system with 360 degree horizontal viewing angle. In:Proc. Int. Symposium of 3D and Contents. 2010. S. 7-10. Summary of the Invention [Problem to be solved by the invention]

[0009] Nevertheless, it would be more advantageous to have a concept within reach that further improves the multi-view / layer coding concept.

[0010] Accordingly, it is an object of the present invention to provide a concept that further improves the multi-view / layer coding concept, in particular to provide the possibility of enabling low delay between terminals, but without abandoning at least one alternative where the decoder cannot handle or does not use the low delay concept.

[0011] This object is achieved by the subject matter of the pending independent claims. [Means for solving the problem]

[0012] The basic idea of ​​the present application is to provide an interleaved multi-layer video data stream with interleaved decoding units of different layers using further timing control information in addition to the timing control information reflecting the interleaved decoding unit arrangement. The additional timing control information relates either to an alternative according to which all decoding units of an access unit are handled in the unit-wise decoding buffer access, or to an alternative according to which an intermediate procedure is used and the interleaving of the DUs of the different layers is reversed according to additionally transmitted timing control information, thus allowing a DU-wise handling in the decoder buffer, but without interleaving of the decoding units for the different layers. Both alternatives may exist together. Various advantageous embodiments and alternatives are the subject of the various claims attached hereto.

[0013] Preferred embodiments of the present application are described below with reference to the drawings. [Brief description of the drawings]

[0014] [Figure 1] 1 illustrates a video encoder that serves as an example for implementing any of the multi-layer encoders outlined further with respect to the following figures. [Diagram 2] 2 shows a schematic block diagram of a video decoder suitable for the video encoder of FIG. 1; [Diagram 3] 1 shows a schematic diagram of a picture being subdivided into sub-streams for WPP processing. [Figure 4] 1 shows a schematic diagram illustrating a picture of several layers subdivided into blocks showing a further subdivision of the picture into spatial segments. [Diagram 5] 1 shows a schematic diagram of a picture of several layers, subdivided into blocks and tiles. [Figure 6] 1 shows a schematic diagram of a picture subdivided into blocks and sub-streams. [Figure 7] Here we show a schematic diagram of a multi-layered video data stream exemplarily comprising three layers, in which options 1 and 2 for arranging NAL units belonging to each time and each layer in the data stream are illustrated in the lower half of Figure 7. [Figure 8] A schematic diagram of a portion of one data stream is given by illustrating these two options in the exemplary case of two layers. [Figure 9] As a comparative embodiment, a schematic block diagram of a decoder configured to process a multi-layer video data stream according to FIGS. 7 and 8 of option 1 is shown. [Figure 10] 10 shows a simplified block diagram of an encoder suitable for the decoder of FIG. [Figure 11] 13 shows an example syntax of a portion of the VPS syntax extension that includes a flag indicating interleaved transmission of DUs of changing layers. [Figure 12a] 13 illustrates an exemplary syntax of an SEI message including timing control information enabling DU de-interleaving of DUs delivered interleaved from a buffer of a decoder according to an embodiment. [Figure 12b] 12a shows an exemplary syntax of an SEI message according to an alternative embodiment, which is to be interspersed at the beginning of interleaved DUs and which also carries timing control information of FIG. 12a. [Figure 12c] 13 illustrates an exemplary syntax of an SEI message that reveals timing control information that enables buffer search for a DU when maintaining interleaving DUs according to one embodiment. [Figure 13] 1 shows a schematic diagram of the bitstream order of DUs of three layers over time, with layer indices exemplified as registered numbers 0 to 2. [Figure 14] Compared with FIG. 13, a schematic diagram of the bitstream order of the DUs of the three layers in which the DUs are interleaved over time is shown. [Figure 15] 1 shows a diagram illustrating the distribution of multi-layer DUs for multiple CPBs according to one embodiment. [Figure 16] 1 shows a diagram illustrating memory address illustrations for multiple CPBs according to one embodiment; [Figure 17] 8 shows a simplified diagram of a multi-layer video data stream with an encoder modified in accordance with FIG. 7 to accommodate an embodiment of the present application; [Figure 18] 8 shows a simplified diagram of a multi-layer video data stream with an encoder modified in accordance with FIG. 7 to accommodate an embodiment of the present application; [Figure 19] 8 shows a simplified diagram of a multi-layer video data stream with an encoder modified in accordance with FIG. 7 to accommodate an embodiment of the present application; [Figure 20] 8 shows a simplified diagram of a multi-layer video data stream with an encoder modified in accordance with FIG. 7 to accommodate an embodiment of the present application; [Figure 21] 10 shows a schematic diagram of a multi-layer video data stream together with a modified decoder associated with the one exemplarily shown in FIG. 9 to correspond to an embodiment of the present application. [Figure 22] 10 shows a schematic diagram of a multi-layer video data stream together with a modified decoder associated with the one exemplarily shown in FIG. 9 to correspond to an embodiment of the present application. [Diagram 23]10 shows a schematic diagram of a multi-layer video data stream together with a modified decoder associated with the one exemplarily shown in FIG. 9 to correspond to an embodiment of the present application. [Figure 24] 10 shows a schematic diagram of a multi-layer video data stream together with a modified decoder associated with the one exemplarily shown in FIG. 9 to correspond to an embodiment of the present application. [Diagram 25] 1 shows a block diagram illustrating an intermediate network located upstream from a decoder buffer. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0015] First, as an overview, an example of an encoder / decoder structure is presented, suitable for the embodiments presented later, i.e., an encoder can be implemented to take advantage of the concepts outlined later, and the same applies for the decoder.

[0016] Fig. 1 shows the overall structure of an encoder according to an embodiment. The encoder 10 may be implemented in a multi-threaded manner or not, i.e., it can simply operate in a single thread. That is, the encoder 10 may be implemented, for example, using multiple CPU cores. In other words, the encoder 10 may support parallel processing, but it does not have to. The generated bitstream would also be producible / decodable by a single-threaded encoder / decoder. The encoding concept of the present application, however, allows applying parallel processing to a parallel processing encoder efficiently but without providing compression efficiency. With regard to parallel processing capabilities, a similar statement is valid for the decoder described later with reference to Fig. 2.

[0017] Encoder 10 is a video encoder. A picture 12 of a video 14 is shown entering encoder 10 at input 16. Picture 12 shows a particular scene, i.e., picture content. However, encoder 10 also receives at its input 16 another picture 15 that is related to the same moment as both pictures 12 and 15, belonging to a different layer. For illustrative purposes only, picture 12 is shown to belong to layer 0, while picture 15 is shown to belong to layer 1. FIG. 1 illustrates that layer 1 may include a higher spatial resolution with respect to layer 0, i.e., show the same scene with a higher number of picture samples. However, this is for illustrative purposes only, and picture 15 of layer 1 may instead have the same spatial resolution, but differ, for example, in view direction with respect to layer 0. That is, pictures 12 and 15 may be captured from different viewpoints. It is noted that the terms base and extended used in this text may refer to both a reference and a dependent layer in a hierarchy of layers.

[0018] The encoder 10 is a hybrid encoder, i.e. the pictures 12 and 15 are predicted by a predictor 18 of the encoder 10 and a prediction residual 20 obtained by a residual determiner 22 of the encoder 10 and are subjected to a transformation, such as a spectral decomposition, such as a DCT, and a quantization in a transform / quantization module 24 of the encoder 10. A transformed and quantized prediction residual 26 thus obtained is subjected to an entropy coding, for example an arithmetic coding or a variable length coding with context adaptivity, in an entropy coder 28. A reconstructable version of the residual is available to the decoder, i.e. a dequantized and retransformed residual signal 30 is recovered by a retransform / requantization module 31 and recombined with the prediction signal 32 of the predictor 18 by a combiner 33, thereby resulting in a reconstruction 34 of the pictures 12 and 15, respectively. However, the encoder 10 operates on a block-by-block basis. Accordingly, reconstructed signal 34 suffers from discontinuities at block boundaries, and accordingly, filter 36 may be applied to reconstructed signal 34 to produce reference pictures 38 based on which predictor 18 predicts the next different layer coded picture for pictures 12 and 15, respectively. As indicated by the dashed line in FIG. 1, predictor 18 may, however, also utilize reconstructed signal 34 directly, without filter 36 or an intermediate version, as in other prediction modes, such as spatial prediction mode.

[0019] The predictor 18 may select among different prediction modes to predict a particular block of the picture 12. One such block 39 of the picture 12 is representative of any block of the picture 12 into which the picture 12 is partitioned, exemplarily shown in FIG. 1, and there may be a temporal prediction mode according to which the block 39 is predicted based on a pre-encoded picture of the same layer, such as the picture 12′. A spatial prediction mode exists according to which the block 39 is predicted based on a pre-encoded portion of a neighboring block 39 of the same picture 12. A block 41 of the picture 15 is also exemplarily shown in FIG. 1, since it is representative of any of the other blocks into which the picture 15 is partitioned. For the block 41, the predictor 18 may support the prediction modes just discussed, i.e., the temporal and spatial prediction modes. In addition, the predictor 18 may provide for an inter-layer prediction mode according to which the block 41 is predicted based on a corresponding portion of the picture 12 of a lower layer. The "corresponding" in "corresponding portion" means spatial correspondence, i.e., a portion in picture 12 that shows the same part of the scene as predicted block 41 in picture 15.

[0020] The prediction of the predictor 18 may of course not be limited to picture samples. Prediction may be applied to any coding parameter, i.e. prediction mode, motion vectors for temporal prediction, disparity vectors for multiview prediction, etc. Simply, the residual may then be coded in the bitstream 40. That is, the coding parameters may be predictively coded / decoded using spatial and / or inter-layer prediction. Even here, disparity compensation may be used.

[0021] A particular syntax is used to compile the quantized residual data 26, i.e. to convert also the coefficient levels and other residual data, as well as the coding parameters, including, for example, the prediction modes and prediction parameters, for the individual blocks 39 and 41 of the pictures 12 and 15 as determined by the predictor 18. The syntax elements of this syntax are also subject to entropy coding by the entropy coder 28. The data stream thus obtained as output by the entropy coder 28 forms the bitstream 40 output by the coder 10.

[0022] Fig. 2 shows a decoder adapted to the encoder of Fig. 1, i.e. a decoder capable of decoding the bitstream 40. The decoder of Fig. 2 is generally indicated by a reference signal 50 and comprises an entropy decoder, a transform / inverse quantization module 54, a combiner 56, a filter 58 and a predictor 60. The entropy decoder 42 receives the bitstream and performs entropy decoding in order to receive residual data 62 and coding parameters 64. The retransform / inverse quantization module 54 inverse quantizes and retransforms the residual data 62 and transfers the residual signal thus obtained to the combiner 56. The combiner 56 also receives a prediction signal from the predictor 60. The predictor 60 in turn forms a prediction signal with the coding parameters 64 based on a reconstructed signal 68 determined by the combiner 56 by combining the prediction signal 66 and the residual signal 65. The prediction mirrors the prediction finally selected from the predictor 18. That is, the same prediction modes are available and these modes are selected for the individual blocks of pictures 12 and 15 and steered according to prediction parameters. As already explained above with respect to Fig. 1, predictor 60 may use a filtered version of reconstructed signal 68 or alternatively or in addition to the same intermediate versions thereof. The pictures of the different layers to be finally reproduced and output at output 70 of decoder 50 are also determined on the unfiltered version of combined signal 68 or the same filtered versions thereof.

[0023] The encoder 10 of Fig. 10 supports the tile concept. According to the tile concept, pictures 12 and 15 are subdivided into tiles 80 and 82, respectively, and within these tiles 80 and 82, predictions of at least blocks 39 and 41, respectively, are restricted to using only data relating to the same tile of the same picture as a basis for spatial prediction. This means that the spatial prediction of block 39 is restricted to using pre-coded parts of the same tile, but the temporal prediction mode is not restricted to relying on information of pre-coded pictures such as picture 12'. Similarly, the spatial prediction mode of block 41 is restricted to using pre-coded data of the same tile only, but the temporal and inter-layer prediction modes are not restricted. The predictor 60 of the decoder 50 is also configured to handle tile boundaries, in particular the selection and / or adaptation of the predictor and the entropy context is performed only within one tile without crossing any tile boundaries.

[0024] The subdivision of pictures 15 and 12 into six tiles, respectively, has been chosen merely for illustrative purposes. The subdivision into tiles can be selected and signaled in the bitstream 40 individually for pictures 12, 12' and 15, 15', respectively. The number of tiles for pictures 12 and 15, respectively, can be either 1, 2, 3, 4, 6, etc. The partitioning of tiles can be limited to a regular partitioning into rows and columns of tiles only. For completeness, it is noted that the scheme of coding tiles separately can be limited to intra- or spatial prediction, but can also encompass any prediction of coding parameters across tile boundaries, and context selection in entropy coding. That is, the latter can also be limited to rely only on data of the same tile. Thus, the decoder can perform the operations just mentioned in parallel, i.e. in units of tiles.

[0025] The encoder and decoder of Figures 1 and 2 may alternatively or additionally use / support the WPP (wavefront parallel processing) concept. With reference to Figure 3, the WPP substream 100 represents a spatial partitioning of the pictures 12, 15 into WPP substreams. In contrast to tiles and slices, the WPP substream does not impose restrictions on prediction and context selection across the WPP substream 100. The WPP substream 100 extends row-wise to cross LCUs (Largest Coding Units) 101, i.e. rows of blocks for which the predictive coding mode is most likely to be sent individually in the bitstream. Also, to allow parallel processing, only one compromise is made with respect to the entropy coding. In particular, instructions 102 are defined within WPP substream 100 that exemplarily lead from top to bottom, and for each of WPP substreams 100, except for the first WPP substream, in instruction 102 the probability estimates for the symbol alphabet, i.e., the entropy probabilities, are completely reset, but for each WPP substream, respectively, on the same side of pictures 12 and 15, such as on the left hand side, as indicated by arrow 106 in the LCU row direction, as indicated by line 104, starting with an LCU instruction, or substream decoder instruction, where the entropy estimate is taken from the subsequent resulting probability having its immediately preceding WPP substream coded / decoded, or equal to, as indicated by line 104, up to the second LCU. Accordingly, by following some encoding delay between sequences of WPP substreams of the same pictures 12 and 15, these WPP substreams 100 can be decoded / encoded in parallel, respectively, to form a kind of wavefront 108 that moves across the picture such that portions are encoded / decoded in parallel in each of the pictures 12, 15, i.e., tiled together from left to right.

[0026] Briefly note that instructions 102 and 104 also define a raster scan order among the LCUs leading from the top-left LCU 101 to the bottom-right LCU row by row from top to bottom. A WPP substream corresponds to one LCU row each. Referring briefly back to tiles, the latter may also be restricted to be aligned to LCU boundaries. A substream may be broken into one or more slices without being tied to an LCU boundary as far as boundaries between two slices inside the substream are concerned. Entropy probability is, however, employed in that case when passing from one slice of a substream to the next of the substream. In the case of tiles, the entire tile may be collapsed into one slice, or a tile may be broken into one or more slices that are not again tied to an LCU as far as boundaries between two slices inside the tile are concerned. In the case of tiles, the order among the LCUs is modified to first traverse the tile in the raster scan order in the tile order before proceeding to the next tile in the tile order.

[0027] As explained so far, picture 12 may be partitioned into tiles or WPP substreams, and likewise picture 15 may be partitioned into tiles or WPP substreams. In theory, the partitioning / concept of the WPP substream may be selected for one of pictures 12 and 15 while the partitioning / concept of the tile is selected for the two others. Alternatively, a restriction may be imposed on the bitstream according to which concept type, i.e., tiles or WPP substreams should be the same within a layer.

[0028] Another example of a spatial segment encompasses a slice. Slices are suitable for dividing the bitstream 40 for transmission purposes. Slices are packed into NAL units, which are the smallest entities for transmission. Each slice is independently encodable / decodable, i.e., any prediction across slice boundaries is forbidden, as is context selection, etc.

[0029] These are only three examples of spatial segments: slices, tiles, and WPP substreams. In addition, all three parallelization concepts, tiles, WPP substreams, and slices, may be used in combination. That is, picture 12 or picture 15 may be split into tiles, and each tile may be split into multiple WPP substreams. Slices may also be suitable for partitioning the bitstream into multiple NAL units, for example (but not limited to) at tile or WPP boundaries. If pictures 12, 15 are partitioned with tiles or WPP substreams and, in addition, with slices, and the partitioning of slices deviates from the partitioning of other WPP / tiles, then a spatial segment would be defined as the smallest independently decodable section of picture 12, 15. Alternatively, restrictions may be imposed on the bitstream that a combination of concepts may be used within a picture (12 or 15) and / or if boundaries should be aligned between the different used concepts.

[0030] Various prediction modes are supported by the encoder and decoder, as well as by restrictions imposed on the prediction modes and context derivation, to enable parallel processing concepts such as the tile and / or WPP concepts described above. It was also mentioned above that the encoder and decoder may operate on a block-by-block basis. For example, the prediction modes described above are selected on a block-by-block basis, i.e., at a finer granularity than the picture itself. Before proceeding with the described aspects of the present application, the relationship between slices, tiles, WPP sub-streams, and blocks just mentioned according to one embodiment will be explained.

[0031] FIG. 4 illustrates a picture, which may be a layer 0 picture, such as layer 12 or a layer 1 picture, such as picture 15. The picture is regularly subdivided into an array of blocks 90. Sometimes these blocks 90 are referred to as largest coding blocks (LCBs), largest coding units (LCUs), coding tree blocks (CTBs), etc. The subdivision of the picture into blocks 90 may form a type of basis or coarsest granularity on which the prediction and residual coding described above is performed. Also, this coarsest granularity, i.e., the size of the blocks 90, may be signaled and set by the encoder independently for layers 0 and 1. For example, a multi-tree, such as a 4-tree subdivision, may be used and signaled in the data stream to subdivide each of the blocks 90 into prediction blocks, residual blocks, and / or coding blocks, respectively. In particular, the coding block may be a leaf block of a recursive multi-tree subdivision of block 90, and some prediction relationship decisions are signaled at the granularity of the coding block, such as the prediction mode, and the prediction block is coded at the granularity of prediction parameters, such as motion vectors, in the case of inter-layer prediction, for example, temporal inter-prediction and disparity vectors, and the prediction residual is coded at the granularity of the residual block, which may be a leaf block of a multi-tree subdivision of a separate recursive code block.

[0032] Raster scan encoding / decoding instructions 92 may be defined within block 90. ​​The encoding / decoding instructions 92 limit the availability of neighboring portions for spatial prediction purposes: only the portion of the picture according to which the encoding / decoding instructions 92 precede the current portion, such as block 90 or some smaller block thereof, is available for spatial prediction within the current picture, since the syntax element being predicted in the current picture is related. Within each layer, the encoding / decoding instructions 92 traverse the entire block 90 of the picture to then continue with the traversal block of the next picture of each layer in the picture encoding / decoding instructions, which do not necessarily follow the temporal regeneration of the picture. Within each block 90, the encoding / decoding instructions 92 are refined to scan within smaller blocks, such as encoding blocks.

[0033] With respect to the just outlined block 90 and smaller blocks, each picture is further subdivided into one or more slices along the just mentioned encoding / decoding instructions 92. The slices 94a and 94b exemplarily shown in Fig. 4 cover each picture accordingly in a gapless manner. The boundary or interface 96 between consecutive slices 94a and 94b of one picture may or may not be aligned to the boundary of the adjacent block 90. ​​More precisely, consecutive slices 94a and 94b in one picture, illustrated on the right hand side of Fig. 4, may be adjacent to each other at the boundary of a smaller block, such as a coding block, i.e. a leaf block of one subdivision of the block 90.

[0034] Slices 94a and 94b of a picture may form the smallest unit in the portion of the data stream into which the picture is coded and may be packetized into packets, i.e., NAL units. Further possible properties of slices, i.e., restrictions on slices with respect to, for example, prediction and entropy context determination across slice boundaries, have been described above. Slices with such restrictions may be referred to as "standard" slices. In addition to standard slices, there may also be "dependent slices", as outlined in more detail below.

[0035] The encoding / decoding instructions 92 defined in the array of blocks 90 may change if a tile partitioning concept is used for the picture. This is shown exemplarily in FIG. 5, where the picture is shown partitioned into four tiles 82a-82d. As illustrated in FIG. 5, the tiles are themselves defined as regular divisions of the picture in units of blocks 90. That is, each tile 82a-82d is composed of an array of n×m blocks 90 with n set individually for each row of the tile and m set individually for each column of the tile. Following the encoding / decoding instructions 92, the blocks 90 in the first tile are first scanned in a raster scan order, before the procedure for the next tile 82b, etc., where the tiles 82a-82d are themselves scanned in a raster scan order.

[0036] According to the WPP stream partitioning concept, a picture is subdivided into WPP sub-streams 98a-98d in units of one or more rows of a block 90 along the encoding / decoding instructions 92. Each WPP sub-stream covers, for example, one complete row of a block 90 as illustrated in FIG.

[0037] The tile concept and the WPP substream concept can, however, also be mixed, in which case each WPP substream covers, for example, one row of blocks 90 within each tile.

[0038] Even slice partitioning of a picture may be used in common with tile partitioning and / or WPP substream partitioning. In terms of tiles, one or more slices of a picture may each be subdivided into either one complete tile or one or more complete tiles, or just composed of a sub-portion of one tile along the encoding / decoding instructions 92. Slices may also be used to form WPP substreams 98a-98d. For this purpose, slices forming the smallest unit for packetization may comprise, on the one hand, standard slices and, on the other hand, dependent slices: while standard slices impose the above-described restrictions on prediction and entropy context derivation, dependent slices do not impose such restrictions. Dependent slices that start at a picture boundary since the encoding / decoding instructions 92 are far enough away in terms of rows adopt the entropy context as resulting from the entropy decoding block 90 in the row immediately preceding block 90. Also, a dependent slice starting somewhere else may adopt the entropy coding context resulting from the entropy coding / decoding of the immediately preceding slice up to its end. In this manner, each of the WPP substreams 98a-98d may be composed of one or more dependent slices.

[0039] That is, the encoding / decoding instructions 92 defined in the block 90 lead linearly from the first side of each picture, exemplarily on the left side here, to the opposite side, exemplarily on the right side, and down to the next row of the block 90 in a downward / bottom direction. The available, i.e. already encoded / decoded portion of the current picture, as the current block 90, is accordingly located primarily to the left and top of the current encoding / decoding portion. Due to the disruption of prediction and entropy context derivation across tile boundaries, tiles of one picture can be processed in parallel. The encoding / decoding of tiles of one picture can even be started simultaneously. The limitations arising from the above-mentioned in-loop filtering in the case of the same are allowed to cross tile boundaries. In turn, the starting encoding / decoding of the WPP substreams is performed in a staggered manner from top to bottom. The intra-picture delay between successive WPP substreams is measured in a number of blocks 90, two blocks 90.

[0040] However, it may be preferable to even process the encoding / decoding of pictures 12 and 15 in parallel, i.e., the different layer instants. Obviously, the encoding / decoding of picture 15 of the dependent layer must be delayed compared to the encoding / decoding of the base layer to ensure that there is a "spatially corresponding" part of the base layer already available. These considerations are valid even without parallel processing of any of the encoding / decoding within any of pictures 12 and 15 individually. With non-tile and non-WPP substream processing, encoding / decoding of pictures 12 and 15, respectively, can be parallelized even when using one slice to cover the entire pictures 12 and 15. The signaling described next, i.e., the sixth aspect, makes it possible to express such encoding / decoding delays between multiple layers even in such cases, or regardless of whether tile or WPP processing is used for any of the pictures of the layers.

[0041] Before discussing the above-mentioned concepts of the present application, please refer again to Figures 1 and 2 and note that the block structures of the encoder and decoder in Figures 1 and 2 are for illustrative purposes only and the structures may also differ.

[0042] There are applications such as videoconferencing and industrial surveillance applications where the end-to-end delay would be as low as possible, however, with multi-layer (scalable) coding. The embodiments described in more detail below allow for lower end-to-end delay in multi-layer video coding. In this regard, it should also be noted that the embodiments described below are not limited to multi-view coding. The multiple layers described below may include different views, but may also represent the same view with varying degrees of spatial resolution, such as SNR accuracy. Possible scalable dimensions along the multiple layers discussed below increase the information content carried by the previous layer being multiple and with, for example, the number of views, spatial resolution, and SNR accuracy.

[0043] As explained above, NAL units are composed of slices. The tile and / or WPP concept can be freely selected individually for different layers of a multi-layer video data stream. Accordingly, each NAL unit having slices packetized therein can be spatially attributed to the area of ​​the picture to which each slice refers. Accordingly, in order to enable low-delay encoding in the case of intra-layer prediction, it would be preferable for the encoder and decoder to be able to interleave NAL units of different layers that are associated at the same instant in time, in order to enable parallel processing of these pictures of different layers, but also to be able to start encoding, transmitting, and decoding, respectively, slices that are packetized into these NAL units in a corresponding manner at the same instant in time. However, depending on the application, the encoder may be better off with the ability to use different encoding orders among pictures of different layers, such as the use of different GOP structures for different layers, beyond the ability to enable parallel processing in the layer dimension. The syntax of a data stream according to a comparative embodiment is described below with reference to FIG. 7.

[0044] 7 shows a multi-layer video material 201 composed of a sequence of pictures 204 for each of the different layers. Each of the layers may describe a different characteristic of the scene (video content) described by the multi-layer video material 201. That is, the meaning of a layer may be selected among, for example, color components, depth maps and / or viewpoints. Without loss of generality, let us assume that the video material 201 is a multi-view video, with the different layers corresponding to different viewpoints.

[0045] For applications that require low delay, the encoder may decide to signal long-term high-level syntax elements. In that case, the data stream generated by the encoder may look as shown in the center of FIG. 7, with a circle around it. In that case, the multi-layer video stream 200 is composed of a sequence of NAL units 202, such as NAL units 202 belonging to one access unit 206 for a picture of one temporal instant, and NAL units 202 of different access units for different instants. That is, the access unit 206 collects the NAL units 202 of one instant, i.e., the one associated with the access unit 206. Within each access unit 206, for each layer, at least some of the NAL units for each layer are grouped into one or more decoding units 208. This means that it follows that within the NAL units 202 there are different types of NAL units, such as VCL NAL units on the one hand and non-VCL NAL units on the other hand, as indicated above. More specifically, the NAL units 202 may be of different types, and these types may comprise:

[0046] 1) A NAL unit that carries syntax elements related to slices, tiles, WPP substreams, etc., i.e., prediction parameters and / or residual data that describe the picture content in terms of picture sample scale / granularity. There can be one or more such types. The VCL NAL unit is of such a type. Such a NAL unit is removable. 2) Parameter set NAL units may carry information that changes infrequently, such as, for example, long-term coding settings, as described above. Such NAL units may be interspersed in the data stream, for example, to some extent and repeatedly. 3) Supplemental Enhancement Information (SEI) NAL units may carry arbitrary data.

[0047] As an alternative to the term "NAL unit", "packet" is sometimes used subsequently to denote the first type of NAL unit, i.e., the VCL unit, "payload packet", while "packet" also encompasses non-VCL units, to which packets of types 2 and 3 in the above list belong.

[0048] A decoding unit may consist of the first of the above mentioned NAL units. More precisely, a decoding unit may consist of one or more VCL NAL units in an access unit and associated non-NAL units. A decoding unit therefore describes a particular area of ​​a picture, i.e., an area that is coded into one or more slices contained therein.

[0049] The decoding units 208 of NAL units related to different layers are interleaved because, for each decoding unit, the intra-layer prediction that previously coded each decoding unit is based on a portion of a picture of a layer other than the layer to which the decoding unit is related, which portion is coded to the decoding unit preceding each decoding unit in the access unit. See, for example, the decoding unit 208a in FIG. 7. Exemplarily, imagine an area 210 of each picture of the dependent layer 2 and this decoding unit for a particular instant. A collocated area in the base layer picture of the same instant is indicated by 212, and an area of ​​this base layer picture that slightly exceeds this area 212 may be required to fully decode the decoding unit 208a by utilizing intra-layer prediction. The slight excess may be, for example, the result of disparity compensated prediction. This means that the decoding unit(s) 208b preceding the decoding unit 208a in the access unit 206 should also completely cover the area required for intra-layer prediction. Reference is made to the above discussion regarding delay indications that may be used like boundaries for interleaving granularity.

[0050] However, if an application takes advantage of the freedom to select differently the decoding orders of pictures in different layers, the case depicted at the bottom of FIG. 7 with its two circles around it may be preferred. In this case, the multi-layer video data stream has individual access units for each picture belonging to one or more specific combinations of layer ID values ​​and a single temporal instant. As shown in FIG. 7, at the (i-1)th decoder instruction, i.e., at instant t(i-1), each layer may consist of access units AU1, AU2 (and so on), or none (cp instant t(i)), where all layers are contained in a single access unit AU1. However, interleaving is not possible in this case. The access units are arranged in the data stream 200 following the access units of decoding instruction index i, i.e., the decoding instruction i for each layer, followed by the access units related to the pictures of these layers corresponding to decoding instruction i+1, etc. Temporal intra-layer prediction signaling in a data stream signal for either equal coding orders or different picture coding orders for different layers may also be located even overlapping in one or more positions in the data stream, for example in a slice packetized into a NAL unit. In other words, case 2 subdivides the access unit scope: a separate access unit is opened for each combination of instants and layers.

[0051] Note that with respect to NAL unit types, ordering rules defined between them may enable a decoder to determine where boundaries between consecutive access units are located regardless of removable packet type NAL units that have been removed while being communicated or not communicated to the decoder. Removable packet type NAL units may comprise, for example, SEI NAL units, or overlapping picture data NAL units, or other specific NAL unit types. That is, the boundaries between access units remain stationary and the ordering rules are still followed within each access unit, but are broken at each boundary between any two access units.

[0052] For completeness, Figure 18 illustrates that case 1 of Figure 7 allows packets belonging to different layers but the same instant (i-1) to be distributed within one access unit, for example. Case 2 of Figure 16 is depicted similarly at 2 with a circle around it.

[0053] The fact that the NAL units contained in each access unit are actually interleaved or not with respect to their association with the layers of the data stream can be decided at the discretion of the encoder. To facilitate the handling of the data stream, a syntax element may signal the decoder to interleave or deinterleave the NAL units within an access unit that collects all NAL units of a particular time stamp, since the latter makes it easier to process the NAL units. For example, whenever interleaving is signaled to be switched on, the decoder may use one or more coded picture buffers as briefly illustrated with reference to FIG. 9.

[0054] FIG. 9 shows a decoder 700 that may be implemented as outlined above with reference to FIG. 2. Exemplarily, the multi-layer video data stream of FIG. 9, option 1 with a circle around it, is shown as input to the decoder 700. To more easily perform the deinterleaving of NAL units belonging to different layers, but common instant, access unit AU, the decoder 700 uses two buffers 702 and 704, with a multiplexer 706 forwarding for each access unit AU, the NAL units of the access unit AU, e.g., belonging to a first layer to buffer 702, and the NAL units of the access unit AU, e.g., belonging to a second layer to buffer 704. A decoding unit 708 then performs the decoding. For example, in FIG. 9, the NAL units belonging to the base / first layer are shown, e.g., not hatched, while the NAL units of the dependent / second layer are shown with hatching. If the interleaving signaling outlined above is present in the data stream, the decoder 700 may respond to this interleaving signaling in the following manner: If the interleaving signaling signals NAL unit interleaving to be switched on, i.e. NAL units of different layers are interleaved with each other within one access unit AU, and the decoder 700 uses buffers 702 and 704 with multiplexer 706 distributing the NAL units over these buffers as just outlined. Otherwise, however, the encoder 700 simply uses one of the buffers 702 and 704 for all NAL units comprised by any access unit, such as, for example, buffer 702.

[0055] To more easily understand the embodiment of FIG. 9, reference is made to FIG. 9 together with FIG. 10, which illustrates an encoder configured to generate a multi-layer video data stream as outlined above. The encoder of FIG. 10 is generally indicated with reference sign 720 and exemplarily encodes inbound pictures, here two layers, for ease of understanding, indicated as layer 12 forming the base layer, and layer 1 forming the dependent layer. These form different perspectives, as outlined previously. The overall encoding instructions together with the encoder 720 encoding the pictures of layers 12 and 15 scan the pictures of these layers substantially together with their temporal (presentation time) instructions, where the encoder instructions 722 may deviate from the presentation time instructions of pictures 12 and 15 in units of groups of pictures. At each temporal instant, the encoding instructions 722 pass the pictures of layers 12 and 15 and their dependencies, i.e., from layer 12 to layer 15.

[0056] The encoder 720 encodes the pictures of layers 12 and 15 into the data stream 40 in terms of the aforementioned NAL units, each of which is associated in a spatial sense with a portion of the respective picture. Hence, the NAL units belonging to a particular picture spatially subdivide or partition the respective picture, and as already explained, the intra-layer prediction renders the portion of the picture of layer 15 that depends on the portion of the temporally aligned picture of layer 12 that is substantially collocated with the portion of the picture of layer 15 by "substantially" surrounding the disparity displacement. In the example of FIG. 10, the encoder 720 has been chosen to exploit the interleaving possibilities when forming an access unit that collects all the NAL units that belong to a particular instant in time. The portions of the data stream 40 outside of the one illustrated in FIG. 10 are inbound to the decoder of FIG. 9. That is, in the example of FIG. 10, the decoder 720 uses intra-layer parallelism when encoding layers 12 and 15. As far as the instant t(i-1) is concerned, the encoder 720 starts encoding the layer 1 picture as soon as NAL unit 1 of the layer 12 picture has been encoded. Each NAL unit that has been completely encoded is output by the encoder 720 and is provided at an arrival time stamp that corresponds to the time that each NAL unit was output by the encoder 720. After encoding the first NAL unit of the layer 12 picture at the instant t(i-1), the encoder 720 starts encoding the content of the layer 12 picture and outputs the second NAL unit of the layer 12 picture, which is provided at an arrival time stamp that follows the arrival time stamp of the first NAL unit of the time-aligned picture of layer 15. That is, the encoder 720 outputs the NAL units of the pictures of layers 12 and 15, which all belong to the same instant, in an interleaved manner, and in this interleaved manner, the NAL units of the data stream 40 are actually transmitted. The circumstances in which the encoder 720 has chosen to utilize the interleaving possibilities may be indicated by the encoder 720 in the data stream 40 by respective method of interleaving signaling 724 .The end-to-end delay between the decoder of FIG. 9 and the encoder of FIG. 10 can be reduced so that the encoder 720 is able to output the first NAL unit of dependent layer 15 at an earlier instant t(i-1) compared to a deinterleaving scenario according to which the output of the first NAL unit of layer 15 would be delayed until the completion of encoding and output of all NAL units of the time-aligned base layer picture.

[0057] As already mentioned above, according to an alternative, in case of deinterleaving, i.e. in case of signaling 724 indicating a deinterleaving alternative, the definition of the access units may remain the same, i.e. an access unit AU may collect all NAL units belonging to a certain instant of time. In that case, signaling 724 simply indicates that within each access unit, NAL units belonging to different layers 12 and 15 are either interleaved or not.

[0058] As explained above, depending on the signaling 724, the decoder of FIG. 9 uses either one buffer or two buffers. If interleaving is switched on, the decoder 700 distributes the NAL units over two buffers 702 and 704, such that, for example, layer 12 NAL units are buffered in buffer 702, while layer 15 NAL units are buffered in buffer 704. Buffers 702 and 704 are emptied for access units. This is true in both cases of signaling 724 indicating interleaving or deinterleaving.

[0059] It is preferable if the encoder 720 sets the removal times within each NAL unit such that the decoder unit 708 takes advantage of the possibility of decoding layers 12 and 15 from data stream 40 using inter-layer parallelism. The end-to-end delay, however, is already reduced even if the decoder 700 does not apply inter-layer parallelism.

[0060] As already explained above, the NAL units can be of different NAL unit types. Each NAL unit can have a NAL unit type index indicating the type of each NAL unit among the possible types and within each access unit, and the NAL unit type of each access unit can follow the ordering rule among the NAL unit types, while only between two consecutive access units, the ordering rule is broken so that the decoder 700 can identify the access unit boundary by consulting this rule. For more detailed information, please refer to the H.264 standard.

[0061] With reference to Figures 9 and 10, a decoding unit, DU, is identifiable as consecutive NAL unit operations within one access unit, belonging to the same layer. The NAL units indicated with "3" and "4" in Figure 10 in access unit AU(i-1) form, for example, one DU. All other decoding units in access unit AU(i-1) comprise only one NAL unit. Together, access unit AU(i-1) in Figure 19 exemplarily comprises six decoding units DU, which are interleaved within access unit AU(i-1), i.e., they are composed of NAL unit operations of one layer with one layer alternating between layer 1 and layer 0.

[0062] 7-10 provide mechanisms for enabling and controlling CPB operation in a multi-layer video codec that meets the required ultra-low delay possible in current single layer video codecs such as HEVC. Based on the bitstream instructions described in the figures just mentioned, in the following we describe a video decoder that operates an incoming bitstream buffer, i.e. a picture buffer, that is coded at the decoding unit level, where in addition the video decoder operates multiple CPBs at the DU level. In particular, also in a manner applicable to HEVC extensions, an operation mode is described where additional timing information is provided for the operation of the multi-layer codec in a low-delay manner. This timing provides a mechanism for controlling CPBs for interleaved coding of different multiple layers in the stream.

[0063] In the embodiments described below, case 2 of figures 7 and 10 is not necessary, or in other words does not have to be realized: the access unit may maintain its function as a container that collects all payload packets (VCL NAL units) that carry information about pictures belonging to a particular time stamp or instant - regardless of layer. Nevertheless, the embodiments described below achieve compatibility with different types of encoders or with encoders that prefer different strategies when decoding an inbound multi-layer video data stream.

[0064] That is, the video encoders and decoders described below are still scalable, multi-view or 3D video encoders and decoders. The term layer follows from the above description to be used collectively for scalable video coding layers as well as for view and / or depth maps of multi-view coded video streams.

[0065] The DU-based decoding mode, i.e., DU CPB removal in a successive manner, may still be used by a single layer (basic specification) ultra-low delay decoder according to some of the embodiments outlined below. A multi-layer ultra-low delay decoder would use the interleaved DU-based mode of decoding to achieve low-delay operation with multiple layers as described for case 1 in Figures 7 and 10 and subsequent figures. On the other hand, a multi-layer decoder that does not decode interleaved DUs may fall back to an AU-based decoding process, or a DU-based decoding in a de-interleaved manner, according to various embodiments, which would provide a low-delay operation in between the interleaved and AU-based approaches. The resulting three operation modes are as follows: AU-based decoding: All DUs of an AU are removed from the CPB at the same time. Decoding based on consecutive DUs: Each DU of a multi-layer AU is attributed to the CPB removal time following the DU removal in consecutive order of layers, i.e., all DUs of layer m are removed from the CPB before the DUs of layer (m+1) are removed from the CPB. Decoding based on interleaved DUs: Each DU of a multi-layer AU is attributed to the CPB removal time following the DU removal in the interleaving order across multiple layers, i.e., the DUs of layer m can be removed from the CPB later than the DUs of layer (m+1) are removed from the CPB.

[0066] The additional timing information for interleaved operation enables a system layer device to determine the arrival times at which DUs must arrive at the CPB when the transmitter sends multi-layer data in an interleaved manner, regardless of the decoder operating mode, which is necessary for the correct operation of the decoder to prevent buffer overflow and underflow. How a system layer device (e.g., an MPEG-2 TS receiver) can determine the times at which data must arrive at the decoder CPB is exemplarily shown at the end of the following section, CPB Operation Alone.

[0067] The following table in Figure 11 gives an exemplary embodiment for signaling the presence in the bitstream of DU timing information for operation in interleaved mode.

[0068] Another embodiment would be a suggestion to accommodate an interleaved operating mode in which an indication of DU timing information is provided so that a device that cannot operate in interleaved DU mode can operate in AU mode and overrule the DU timing.

[0069] In addition, another operating mode featuring per-layer DU-based CPB removal, i.e., DU CPB removal in a de-interleaved manner across layers, is made to enable the same low-latency CPB operation based on DUs as in the interleaved mode for the base layer, but removes DUs from layer (m+1) only after completing removal of DUs of layer m. Thus, non-base layer DUs may remain in the CPB for a longer period than when removed in the interleaved CPB operating mode. The tables in Figures 12a and 12b provide an exemplary embodiment for signaling additional timing information in an additional SEI based on either the AU level or the DU level. Other possibilities include suggesting that timing provided by other means leads to CPB removal from a CPB that is interleaved across layers.

[0070] A further aspect is the possibility of applying the decoder operation modes described for the following two cases:

[0071] 1. A single CPB is used to accommodate data of all layers. NAL units of different layers may be interspersed within an access unit. This mode of operation is then referred to as single CPB operation. 2. One CPB per layer. The NAL units of each layer are located in consecutive positions. This mode of operation is then referred to as multiple CPB operation.

[0072] -Standalone CPB operation In Figure 13, the arrival of a decoding unit (1) at an access unit (2) is shown for a bitstream ordered with respect to layers. The numbers in the boxes refer to the layer IDs. As shown, first all DUs of layer 0 arrive, followed by the DUs of layer 1, then layer 2. In the example three layers are shown, however more layers may follow.

[0073] In Figure 14 the arrival of decoding unit (1) at access unit (2) is shown for an interleaved bitstream according to Figures 7 and 8, case 1. The numbers in the boxes refer to layer IDs. As shown, DUs of different layers can be mixed within an access unit.

[0074] A CPB removal time is associated with each decoding unit that is the start time of the decoding process. This decoding time can be lower than the final arrival time of the decoding unit, exemplarily shown as (3) for the first decoding unit. The final arrival time of the first decoding unit of the second layer, labeled (4), can be lowered by using an interleaved bitstream order as shown in FIG.

[0075] One embodiment is a video encoder that generates a decoder hint in a bitstream that indicates the lowest possible CPB removal (and therefore decoding time) in a bitstream that uses high-level syntax elements for an interleaved bitstream.

[0076] A decoder utilizing the decoder hints described for lower arrival times will remove decoding units from the CPB immediately or sooner after their arrival, so that parts of a picture can be fully coded (through all layers) sooner and therefore displayed sooner than they would be due to a deinterleaved bitstream.

[0077] A cheaper implementation of such a decoder can be achieved by constraining the signal timing in the following way for any DUn preceding a DUm in bitstream order: The CPB removal time for DUn will be lower or equal to the CPB removal time of DUm. As they arrive, packets are stored in consecutive memory addresses (typically a circular buffer) in the CPB, and this constraint avoids fragmentation of free memory in the CPB. Packets are removed in the same order that they are received. Instead of keeping a list of used and free memory blocks, the encoder can be implemented to keep only the beginning and ending addresses of used memory blocks. This ensures that newly arriving DUs do not need to be split into several memory locations since the used and free memory are consecutive blocks.

[0078] In the following, an embodiment based on the actual current HRD definition as used by the HEVC extension is described, where the timing information for interleaving is provided through an additional DU-level SEI message as mentioned above. The described embodiment allows DUs to be transmitted in order interleaved across layers, either consecutively or AU-wise, to be removed DU-wise from the CPB in an interleaved manner.

[0079] In a single CPB solution, the CPB removal times in Annex [1] C should be extended as follows (underlined):

[0080] Multiple tests may be required to verify the conformance of a bitstream, referred to as the bitstream under test. For each test, the following steps are applied in the listed order: For each access unit in BitstreamToDecode, starting with access unit 0, a buffer time SEI message (present in BitstreamToDecode or available through external means not specified in this specification) associated with the access unit and applied to the TargetOp is selected; a picture timing SEI message (present in BitstreamToDecode or available through external means not specified in this specification) associated with the access unit and applied to the TargetOp is selected; and when SubPicHrdFlag is 1 and sub_pic_cpb_params_in_pic_timing_sei_flag is 0, a decoding unit information SEI message (present in BitstreamToDecode or available through external means not specified in this specification) associated with a decoding unit in the access unit and applied to the TargetOp is selected; And when sub_pic_interleaved_hrd_params_present_flag is 1, a Decoding Unit Interleaving Information SEI message (present in BitstreamToDecode or available through external means not specified in this specification) associated with the decoding unit in the access unit and applied to the TargetOp is selected. When sub_pic_interleaved_hrd_params_present_flag is 1 in the selected syntax structure, the CPB will operate either at the AU level (in case the variable SubPicInterleavedHrdFlag is set to 0) or at the interleaved DU level (in case the variable SubPicInterleavedHrdFlag is set to 1). variable SubPicInterleavedHrdPreferredFlag is set to 0 either when specified by external means or when not specified by external means. When the value of the variable SubPicInterleavedHrdFlag was not set in step 9 in this subclause above, it is derived as follows. SubPicInterleavedHrdFlag = SubPicHrdPreferredFlag && SubPicInterleavedHrdPreferredFlag && sub_pic_interleaved_hrd_params_present_flag

[0081] SubPicHrdFlag, and SubPicInterleavedHrdFlag If it is 0, HRD operates at the access unit level and each decoding unit is an access unit, otherwise HRD operates at the sub-picture level and each decoding unit is a subset of an access unit.

[0082] For each bitstream conformance test, the CPB operation is specified in subclause C.2, the instantaneous decoder operation is specified in clauses 2 to 10, the DPB operation is specified in subclause C.3, and output cropping is specified in subclause C.3.3 and subclause C.5.2.2.

[0083] The HSS and HRD information for the number of enumerated delivery schedules and their associated bit rates and buffer sizes are specified in subclause E.1.2 and E.2.2. The HRD is initialized as specified by the buffer time SEI message specified in subclause D.2.2 and D.3.2. The removal times of the decoding units from the CPB and the output timing of the coded pictures from the DPB are specified in the decoding unit information SEI message (specified in subclause D.2.21 and D.3.21). or in a Decoding Unit Interleaving Information SEI message (as specified in subclauses D.2.XX and D.3.XX), The timing information for a particular decoding unit shall arrive prior to the CPB removal time of the decoding unit.

[0084] When SubPicHrdFlag is 1, the following applies: The variable duCpbRemovalDelayInc is derived as follows: - If SubPicInterleavedHrdFlag is set to 1, duCpbRemovalDelayInc is set equal to the value of du_spt_cpb_interleaved_removal_delay_increment in the Decoding Unit Interleaving Information SEI message, and subclause C1 is selected as specified in associated with decoding unit m . - Otherwise , sub_pic_cpb_params_in_pic_timing_sei_flag is If it is 0 and sub_pic_interleaved_hrd_params_present_flag is 0, duCpbRemovalDelayInc is set equal to the value of du_spt_cpb_removal_delay_increment in the decoding unit information SEI message, selected as specified in subclause C1, and associated with decoding unit m. - Otherwise, if sub_pic_cpb_params_in_pic_timing_sei_flag is 0 and sub_pic_interleaved_hrd_params_present_flag is 1, then duCpbRemovalDelayInc is set equal to the value of du_spt_cpb_removal_delay_increment in the decoding unit information SEI message and duCpbRemovalDelayIncInterleaved is set equal to the value of du_spt_cpb_interleaved_removal_delay_increment in the decoding unit interleaving information SEI message, selected as specified in subclause C1 and associated with decoding unit m. - Otherwise, du_common_cpb_removal_delay_flag is If it is 0 and sub_pic_interleaved_hrd_params_present_flag is 0, duCpbRemovalDelayInc is set equal to the value of du_cpb_removal_delay_increment_minus1[i]+1 for decoding unit m in the picture timing SEI message, selected as specified in subclause C.1, and associated with access unit n, where the first num_nalus_in_du_minus1[0]+1 NAL units with i value 0 in the access unit include decoding unit m, the next num_nalus_in_du_minus1[1]+1 NAL units with i value 1 in the same access unit, the next num_nalus_in_du_minus1[2]+1 NAL units with i value 2 in the same access unit, etc. - Otherwise, if du_common_cpb_removal_delay_flag is 0 and sub_pic_interleaved_hrd_params_present_flag is 1, then duCpbRemovalDelayInc is set equal to the value of du_cpb_removal_delay_increment_minus1[i]+1 in the Picture Timing SEI message for decoding unit m, and duCpbRemovalDelayIncInterleaved is set equal to the value of du_spt_cpb_interleaved_removal_delay_increment in the Decoding Unit Interleaving Information SEI message, selected as specified in subclause C.1 and associated with access unit n. There, the first num_nalus_in_du_minus1[0]+1 NAL units with i value 0, consecutive NAL units in an access unit include decoding unit m, the next num_nalus_in_du_minus1[1]+1 NAL units in the same access unit with i value 1, the next num_nalus_in_du_minus1[2]+1 NAL units in the same access unit with i value 2, etc. - Otherwise, duCpbRemovalDelayInc is set equal to the value of du_common_cpb_removal_delay_increment_minus1+1 in the picture timing SEI message, selected as specified in subclause C.1 and associated with access unit n. The nominal removal time of a decoding unit m from the CPB is specified as follows, where AuNominalRemovalTime[n] is the nominal removal time of access unit n: If decoding unit m is the last decoding unit in access unit n, then the nominal removal time of decoding unit m, DuNominalRemovalTime[m], is set equal to AuNominalRemovalTime[n]. Otherwise (decoding unit m is not the last decoding unit in access unit n), the nominal removal time DuNominalRemovalTime[m] for decoding unit m is derived as follows: if( sub_pic_cpb_params_in_pic_timing_sei_flag && !SubPicInterleavedHrdFlag) DuNominalRemovalTime[m] = DuNominalRemovalTime[m+1] - ClockSubTick*duCpbRemovalDelayInc (C13) else DuNominalRemovalTime[m] = AuNominalRemovalTime(n) - ClockSubTick*duCpbRemovalDelayInc

[0085] In which DU operation mode is used SubPicInterleavedHrdFlag determines either interleaved or non-interleaved operation mode, and DUNominalRemovalTime[m] is the removal time of the DU for the selected operation mode. In addition, the earliest arrival time of DUs is different than currently defined when sub_pic_interleaved_hrd_params_present_flag is 1, regardless of the operation mode. The earliest arrival time is therefore derived as follows: if(!SubPicInterleavedHrdFlag&& sub_pic_interleaved_hrd_params_present_flag) DuNominalRemovalTimeNonInterleaved[ m ] = AuNominalRemovalTime(n) - ClockSubTick*duCpbRemovalDelayIncInterleaved if( !subPicParamsFlag ) tmpNominalRemovalTime = AuNominalRemovalTime[m] (C6) else if(!sub_pic_interleaved_hrd_params_present_flag || SubPicInterleavedHrdFlag) tmpNominalRemovalTime = DuNominalRemovalTime[m] else tmpNominalRemovalTime = DuNominalRemovalTimeNonInterleaved[m] "

[0086] With regard to the above embodiment, it is worth noting that the operation of the CPB is a factor of the arrival time of the data packets at the CPB as well as the removal time of the data packets which is explicitly signaled. Such arrival time affects the behavior of intermediate devices forming buffers along the data packet transmission chain, for example elementary stream buffers in the receiver of an MPEG-2 Transport Stream, such elementary stream buffers acting as the CPB of the decoder. The HRD model based on the above embodiment derives the first arrival time based on the variable tmpNominalRemovalTime, whereby for the exact first arrival time of the data packets at the CPB (see C-6) either the removal time for DUs in case of interleaved DU operation (as if data were removed from the CPB in an interleaved manner) or the equivalent removal time "DuNominalRemovalTimeNonInterleaved" for the continuous DU operation mode is taken into account.

[0087] A further embodiment is layer-wise reordering of DUs for AU-based decoding operations. When a single CPB operation is used and data is received in an interleaved manner, it may be desirable for the decoder to operate on an AU basis. In such a case, data read from the CPB corresponding to several layers will be interleaved and immediately transmitted to the decoder. When an AU-based decoding operation is performed, the AUs are reordered / rearranged in such a way that all DUs from layer m precede DUs from layer m+1 before being transmitted for decoding, since a reference layer is always decoded before the enhancement layer that references it.

[0088] Multi-CPB Operation Instead, the decoder is described with one coded picture buffer for each DU of a layer.

[0089] Figure 15 shows the allocation of DUs to different CPBs. For each layer (number in the box), when its CPB is operated, the DUs are stored in different memory locations for each CPB. Exemplarily, the arrival times of interleaved bitstreams are shown. The allocation works in the same way as for non-interleaved bitstreams, based on layer identifiers.

[0090] Figure 16 shows memory usage in different CPBs. DUs of the same layer are stored in consecutive memory locations.

[0091] A multi-layer decoder can take advantage of such memory allocation because DUs belonging to the same layer can be accessed at consecutive memory addresses. The DUs arrive in the decoding order of each layer. Removal of DUs of different layers cannot create any "holes" in the used CPB memory area. The used memory block always covers consecutive blocks in each CPB. The multiple CPB concept also has the advantage of a bitstream that is split for layers at the transport layer. If different layers are transmitted using different channels, multiplexing of DUs into a single bitstream can be avoided. Thus, a multi-layer video decoder does not need to implement this additional step, and the implementation cost can be reduced.

[0092] When multi-CPB operation is used, in addition to the timing described for the single CPB case also applies:

[0093] A further aspect is the reordering of DUs from multiple CPBs when they share the same CPB removal time (DuNominalRemovalTime[m]). In both interleaved and non-interleaved modes of operation for DU removal, it may occur that DUs from different layers, and therefore different CPBs, share the same CPB removal time. In such cases, the DUs are ordered by increasing LayerId numbers before being transmitted to the decoder.

[0094] The above described embodiment also describes a mechanism for synchronizing with multiple CPBs below. In the present text [1], the reference time or anchor time is described as the first arrival time of the first decoding unit entering a (unique) CPB. In the case of multi-CPB, there is one master CPB and multiple slave CPBs, leading to dependencies between the multiple CPBs. A mechanism for the master CPB to synchronize with the slave CPBs is also described. This mechanism is advantageous because the CPBs receiving DUs remove their DUs at the unique time, i.e. with the same time reference. More specifically, the first DU that initializes the HRD synchronizes with the other CPBs, and the anchor time is set equal to the first arrival time of the DU for the extended CPB. In a particular embodiment, the master CPB is the CPB for the base layer DUs, while a random access point for the enhancement layer is enabled and it may be possible for the master CPB to correspond to the CPB that receives the enhancement layer data when initializing the HRD.

[0095] Thus, following the considerations outlined above following Figures 7-10, the comparative embodiments of these figures are modified in the manner outlined below with respect to the following figures. The encoder of Figure 17 operates similarly to the one discussed above with respect to Figure 10. Signaling 724 is, however, optional. Accordingly, the scope of the above description also applies to the following embodiment, and similar statements would apply to the decoder embodiment described subsequently.

[0096] In particular, the encoder 720 of FIG. 17 exemplarily encodes video content including video of layers 12 and 15 into a multi-layer video data stream 40, the same having video content encoded therein in units of sub-portions of pictures of the video content using intra-layer prediction for each of the plurality of layers 12 and 15. In the example, the sub-portions are denoted 1-4 for layer 0 and 1-3 for layer 1. Each of the sub-portions is encoded into one or more payload packets of a sequence of packets of the video data stream 40, each packet being associated with one of the plurality of layers, and the sequence of packets being divided into a sequence of access units AU, such that an access unit collects payload packets associated with a common instant. Two AUs are exemplarily denoted, one at instant i-1 and the other at instant i. The access units AU are subdivided into decoding units DU, such that each access unit is subdivided into two or more decoding units, with each decoding unit solely comprising a payload packet associated with one of the plurality of layers. Decoding units with packets associated to different layers are interleaved with each other. In simple terms, the encoder 720 controls the interleaving of decoding units within an access unit in order to reduce - or keep as low as possible - the end-to-end delay by traversing and encoding common instants in the traversal order of layers first and sub-parts later. So far, the encoder's modes of operation have already been provided above with respect to FIG. 10.

[0097] However, the encoder of Fig. 17 provides each access unit AU with two times of time control information: a first time control information 800 signals the respective decoder buffer search times for the access unit as a whole, and a second time control information 802 signals the decoder buffer search times for each decoding unit DU of the access unit AU which correspond to their sequential order in the multi-layer video data stream.

[0098] As illustrated in Fig. 17, the encoder 720 spreads the second timing information 802 into several timing control packets, each of which precedes the associated decoding unit DU and indicates a second decoder search buffer time for the preceding decoding unit DU. Fig. 12c shows an example of such a timing control packet. As can be seen, the timing control packet may form the beginning of the associated decoding unit and indicates an index associated with the respective DU, i.e., decoding_unit_idx, and a decoder search buffer time for the respective DU, i.e., du_spt_cpb_removal_delay_interleaved_increment, which indicates the search time or DPB removal time in a predefined time unit (increment). Thus, the second timing control information 802 may be output by the encoder 720 while the current instantaneous layer is being coded. For this, the encoder 720 reacts to the spatial complexity changes in the pictures 12 and 15 of the various layers during coding.

[0099] The encoder 720 may evaluate the first timing control information 800 at the decoding buffer search time for each access unit AU, i.e., prior to encoding the layer at the current instant and location, at the beginning of each AU, or - if possible according to a criterion - at the end of the AU.

[0100] In addition or instead of providing the timing control information 800, the encoder 720 provides, to each of the access units, third timing control information signaling a third decoder buffer search time for each of the decoding units of each of the access units, as shown in Fig. 18. As a result, according to the third decoder buffer search times for the decoding units DU of each of the access units, the decoding units DU in each of the access units AU are ordered according to a layer order defined among the layers, so that a non-decoding unit comprising a packet associated with a first layer follows any decoding unit in each of the access units, comprising a packet associated with a second layer that follows the first layer according to the layer order. That is, according to the buffer search times of the third timing control information 804, the DUs shown in Fig. 18 are repartitioned at the decoder side so that the DUs of parts 1, 2, 3 and 4 of picture 12 precede the DUs of parts 1, 2 and 3 of picture 15. The encoder 720 may estimate the decoder buffer search time according to the first timing control information 802 at the beginning of each AU, the third timing control information 80 for DUs prior to encoding the layer of the current moment and place. This possibility is exemplarily depicted in FIG. 12a in FIG. 8. FIG. 12a shows that ldu_spt_cpb_removal_delay_interleaved_increment_minus1 is transmitted for each DU and for each layer. Although the number of decoding units per layer may be equally restricted for all layers, i.e., one num_layer_decoding_units_minus1 would be used as illustrated in FIG. 12a, an alternative to FIG. 12a would be that the number of decoding units per layer could be set individually for each layer. In the latter case, the syntax element num_layer_decoding_units_minus1 may be read for each of the layers, in which case the reading would be moved from the location shown in FIG. 12a to, for example, between the two next for loops in FIG. 12a.As a result, num_layer_decoding_units_minus1 will be read for each layer in the next for loop using the counter variable j. If it is allowed to follow the criteria, the encoder 720 sets the timing control information at the end of the AU instead. Even instead, the encoder 720 may set the third timing control information at the beginning of each DUs just like the second timing control information. This is shown in Fig. 12b. Fig. 12b shows an example for a timing control packet that is set at the beginning of each DU (in their interleaved state). As can be seen from the fact that the timing control packet carrying the timing control information 804 for a particular DU may be set at the beginning of the associated decoding unit and indicates one index, i.e. layer_decoding_unit_idx, associated with each DU, it is layer specific, i.e. all DUs belonging to the same layer are attributed to the DU index of the same layer. Furthermore, ldu_spt_cpb_removal_delay_interleaved_increment, which indicates the decoder search buffer time for each DU, i.e., the search time or DPB removal time in a given temporal unit (increment), is signaled in such packet. According to these timings, the DUs are re-partitioned to follow the layer order, i.e., first the DUs of layer 0 are removed from the DPB, then the DUs of layer 1, etc. Accordingly, the timing control information 808 may be output by the encoder 720 during the encoding of the current instantaneous layer.

[0101] As explained, information 802 and 804 may exist in the data stream simultaneously. This is illustrated in FIG.

[0102] Finally, as illustrated in Figure 20, an interleaving flag 808 of the decoding unit can be inserted into the data stream by the encoder 720 for either the timing control information 802 or the timing control information 808 to be transmitted in addition to the timing control information 800 acting as 804. That is, if the encoder 720 decides to interleave the DUs of different layers as depicted in Figure 20 (and Figures 17-19), then the interleaving flag 808 of the decoding unit is set to indicate that the information 808 is equal to the information 802, and the above description of Figure 17 applies with respect to the remainder of the functionality of the encoder of Figure 20. However, if the encoder does not interleave packets of different layers within an access unit as depicted in FIG. 13, then the decoding unit interleaving flag 808 is set by the encoder 720 to indicate that information 808 is equal to information 804, with the difference from the depiction of FIG. 18 regarding the generation of 804 that the encoder therefore does not need to evaluate timing control information 804 in addition to information 802, but can determine the buffer search time of timing control information 804 on the fly using a procedure similar to that of generating timing control information 802 while encoding an access unit, thereby reacting to variations in layer specific coding complications within the layer sequence on the fly.

[0103] Figures 21, 22, 23 and 24 respectively show the data streams of Figures 17, 18, 19 and 20 as they enter the decoder 700. If the decoder 700 is configured as the one described with respect to Figure 9 above, then the decoder 700 can decode the data stream in the same way as described above with respect to Figure 9 using the timing control information 802. That is, the encoder and decoder both contribute to the minimum delay. In the case of Figure 24, the decoder receives the stream of Figure 20, which is obviously only possible in the case of DU interleaving used and indicated by flag 806.

[0104] For whatever reason, the decoder, however, decodes the multi-layer video data stream by emptying the decoder buffer for buffering the multi-layer data stream on an access unit basis using the first timing control information 800 in the cases of Figs. 21, 23 and 24 and regardless of the second timing control information 802. For example, the decoder may not be able to perform parallel processing. The decoder may not have, for example, one or more buffers. Since the decoder buffer operates on complete AUs rather than on a DU level, the delay is increased on both the encoder and decoder sides compared to utilizing the interleaved layer order of DUs according to the timing control information 802.

[0105] As already discussed above, the decoder of Figs. 21-24 does not need to have two buffers. One buffer like 702 is sufficient, rather than any alternative in the form of timing control information 800 and 804, especially if timing control information 802 is not utilized. On the other hand, the decoder buffer is composed of a partial buffer for each layer, which is useful when utilizing timing control information, since the decoder may buffer, for each layer, the decoding units with packets associated with each layer in the partial buffer for each layer. With 705 the possibility of having more than one buffer is illustrated. The decoder may empty the decoding units from different partial buffers to the decoder entities of different coders. Instead, the decoder uses a smaller number of partial buffers compared to the number of layers, i.e., each partial buffer for a subset of layers by transferring the Ds of a particular layer for the partial buffer associated with the set of layers to which the layer of each DU belongs. One partial buffer like 702 may be synchronized with other partial buffers like 704 and 705.

[0106] Whatever the reason, the decoder, in the cases of Figures 22 and 23, however, decodes the multi-layer video data stream by emptying the decoder buffer controlled via the timing control information 804, i.e., by removing the decoding units of the access units by deinterleaving according to the layer order. With this measure, the decoder - guided by the timing control information 804 - effectively recombines the decoding units associated with the same layer and belonging to an access unit, and reorders them following a certain rule, such as DUs of layer n before DUs of layer n+1.

[0107] As illustrated in FIG. 24, the decoding unit interleaving the flag 806 in the data stream may operate as if the timing control information 808 is either the timing control information 802 or 804. In this case, the decoder receiving the data stream may be configured to react to the decoding unit interleaving the flag 806 to empty the decoder buffer for buffering the multi-layer data stream in units of access units using the first timing control information 800 and regardless of the information 806 if the information 806 is the second timing control information according to 802, and to empty the decoder buffer for buffering the multi-layer data stream in units of decoding units using the information 806 if the information is the timing control information according to 804: that is, in this case, the end-to-end delay between the DUs that would otherwise be achieved by using 802 and the DU operation ordered using the timing control information 808 is not interleaved. Also, the maximum delay achievable by the timing control information 800 will result.

[0108] Whenever the timing control information 800 is used as an alternative, i.e. the encoder chooses to empty the encoder buffer on an access unit basis, the decoder may remove the decoded units of the access unit from the buffer 702 in a deinterleaving manner, or even fill the buffer 702 with DUs, so that the AUs with DUs are ordered according to the layer order. That is, the decoder may recombine the decoded units associated with the same layer and belonging to the access unit, and reorder them according to a certain rule, such as DUs of layer n before DUs of layer n+1, before the entire AU is removed from the buffer to be decoded accordingly. This deinterleaving is not necessary in the case of the decoding unit interleaving flag 806 of FIG. 24, which indicates that deinterleaving transmission has already been used, and the timing control information 808 works like the timing control information 804.

[0109] Although not specifically discussed above, the second timing control information 802 may be defined as an offset for the first timing control information 800 .

[0110] The multiplexer 706 shown in Figures 21-24 operates as an intermediate network device configured to forward the multi-layer video data stream for the coded picture buffer of the decoder. The intermediate network device is configured according to an embodiment to receive information identifying a decoder as being able to handle the second timing control information 802, and derives the earliest arrival time for scheduling the transfer from the timing control information 802 and 800 according to the first calculation rule, i.e., DuNominalRemovalTime, if the decoder can handle the second timing control information 802; and derives the earliest arrival time for scheduling the transfer from the timing control information 802 and 800 according to the second calculation rule, i.e., DuNominalRemovalTimeNonInterleaved, if the decoder cannot handle the second timing control information 802. To explain in more detail the problem just outlined, an intermediate network device is shown, also indicated with reference number 706, located between the inbound multi-layer video data stream and the output leading to a decoding buffer generally indicated with reference number 702, but reference is made to FIG. 25 outlined above. The decoding buffer may be composed of several partial buffers. As shown in FIG. 25, the intermediate network device 706 internally comprises a buffer 900 for buffering inbound DUs and thus forwards them to the decoding buffer 702. The above-mentioned embodiment was concerned with the decoding buffer removal time of inbound DUs, i.e. the time when these DUs need to be transferred from the decoding buffer 702 to the decoding unit 708 of the decoder. The storage capacity of the decoding buffer 702 is, however, guaranteed as much as possible in addition to the removal time, i.e. the amount of time for which the DUs buffered in the buffer 702 should be removed, and the earliest arrival time will also be managed.This is the purpose of the "earliest arrival time" mentioned before, and according to the embodiment outlined here, the intermediate network device 706 is configured to calculate these earliest arrival times based on the acquired timing control information according to different calculation rules, choosing the calculation rule according to information about the decoder's capabilities to operate on the inbound DUs in the interleaved format of the inbound DUs, i.e. depending on whether the decoder can operate in the same interleaved manner or in a different manner. In principle, the intermediate network device 706 may determine the earliest arrival time based on the DU removal time of the timing control information 802 by providing a constant time offset between the earliest arrival time and the removal time for each DU, in case of a decoder that can decode the inbound data stream using the DU interleaving concept. The intermediate network device 706 similarly provides a fixed time offset between AU removal times, as indicated by the timing control information 800, in order to derive the earliest arrival time of the access unit in case of a decoder that prefers a treatment with respect to the inbound data stream access unit, i.e., selects the first timing control information 800. Instead of using a fixed time offset, the intermediate network device 706 may also take into account the size of the individual DUs and AUs.

[0111] The problem of Fig. 25 will also be used as an opportunity to show possible improvements of the embodiments described above. In particular, the embodiments discussed above treated the timing control information 800, 802 and 804 as "decoder buffer removal time" signaling for the decoding unit and the access unit by directly signaling "removal time", i.e., the time when each DU and AU needs to be transferred from the buffer 702 to the decoding unit 708, respectively. However, as has become clear from the discussion of Fig. 25, the arrival time and removal time are interrelated with each other through the size of the decoder buffer 702, as their guaranteed minimum size, on the other hand, the size of the individual DUs in the case of the timing control information 802 and 804, and the size of the access unit in the case of the timing control information 800, respectively. Accordingly, all of the above embodiments may be interpreted such that the "decoder buffer removal time" signaled by the "timing control information" 800, 802 and 804, respectively, includes both alternatives, the earliest arrival time or the explicit signaling by way of the buffer removal time. All of the above discussions translate directly from the explanations discussed above with the explicit signaling of the buffer removal time as the decoder buffer search time in the alternative embodiment where the earliest arrival time is used as the decoder buffer removal time: the interleaved transmission DUs will be repartitioned according to the timing control information 804. The only difference: the repartitioning or deinterleaving will be done upstream, i.e. before the buffer 702, i.e. between the buffer 702 and the decoding unit 708, rather than downstream thereof.In the case of the intermediate network device 706 calculating the earliest arrival times from the inbound timing control information, the intermediate network device 706 will use these earliest arrival times to command a network entity located upstream with respect to the intermediate network device 706, such as the encoder itself or the same intermediate network entity, in the alternative case of deriving a buffer removal time from the inbound timing control information, to follow these earliest arrival times in feeding the buffer 900. The intermediate network device 706 - hence with explicit signaling of the earliest arrival times - will activate access units from the buffer 900 according to the removal times of DUs or, in the alternative case, the derived removal times.

[0112] Summarizing the very outline of the embodiment alternatives outlined above, this can be done by using the timing control information directly or indirectly to empty the decoder buffer: if the timing control information is implemented as direct signaling of decoder buffer removal times, then emptying the buffer can be done to be directly scheduled according to these decoder buffer removal times, and if the timing control information is implemented using decoder buffer arrival times, then a recalculation can be done to deduce the decoder buffer removal times from these decoder buffer arrival times according to which the removal of DUs or AUs takes place.

[0113] As noted in common in the above description of the various embodiments and figures illustrating "interleaved packet" transmission, it is presented that "interleaving" does not necessarily include combining packets belonging to DUs of different layers in a common channel. Rather, the transmission can be performed completely in parallel in separate channels (logically or physically separate channels): packets of different layers are output by the encoder in parallel, with the output time interleaved as discussed above, thus forming different DUs, and in addition to the DUs, the time control information mentioned above is also transmitted to the decoder. Among this timing control information, timing control information 800 indicates when DUs forming a complete AU need to be transferred from the decoder buffer to the decoder, timing control information 802 indicates for each individual DU when each DU needs to be transferred from the decoder buffer to the decoder, these removal times corresponding to the output time order of the DUs at the encoder, and timing control information 804 indicates for each individual DU when each DU needs to be transferred from the decoder buffer to the decoder, these search times deviating from the output time order of the DUs at the encoder and leading to a repartition: instead of being transferred from the decoder buffer to the decoder, in the output interleaving order, the DUs of layer i are transferred before the DUs of layer i+1 for all layers. As described, the DUs may be distributed to separate buffer portions according to the layer relationships.

[0114] The multi-layer video data streams of Figures 19 to 23 may be in accordance with AVC or HEVC or any extension thereof, but this is not exclusive of other possibilities.

[0115] Although some aspects have been described in the context of an apparatus, it will be apparent that these aspects also represent a description of a corresponding method, where a block or apparatus corresponds to a method step or a function of a method step. Analogously, aspects may also be described in the context of a method step, where they also represent a description of a corresponding block or item or feature of a corresponding apparatus. Some or all of the method steps may be performed by (or using) a hardware apparatus, such as, for example, a microprocessor, a programmable computer, or an electronic circuit. In some embodiments, one or more of the most significant method steps may be performed by such an apparatus.

[0116] The original encoded video signal may be stored on a digital storage medium or transmitted over a transmission medium, such as a wireless transmission medium or a wired transmission medium such as the Internet.

[0117] Depending on the particular implementation requirements, the embodiments of the present invention may be implemented in hardware or in software. The implementation may be performed using, for example, a digital storage medium such as a floppy disk, DVD, Blu-Ray, CD, ROM, PROM, EPROM, EEPROM or FLASH memory having electronically readable control signals stored thereon, which cooperates (or can cooperate) with a programmable computer system such that the respective methods are executed. Thus, the digital storage medium may be computer readable.

[0118] Some embodiments of the present invention comprise a data carrier having electronically readable control signals, which can cooperate with a programmable computer system to perform one of the methods described herein.

[0119] Generally, embodiments of the present invention may be implemented as a computer program product with program code that is operable to perform one of the methods when the computer program product runs on a computer. The program code may, for example, be stored on a device readable carrier.

[0120] Other embodiments comprise the computer program for performing one of the methods described herein, stored on a machine readable carrier.

[0121] In other words, an embodiment of the inventive method is, therefore, a computer program having a program code for performing one of the methods described herein, when the computer program runs on a computer.

[0122] A further embodiment of the inventive method is therefore a data carrier (or digital storage medium, or computer readable medium) comprising stored thereon a computer program for performing one of the methods described herein. The data carrier, digital storage medium or stored medium is typically tangible and / or non-transitory.

[0123] A further embodiment of the inventive method is therefore a data stream or a sequence of signals representing the computer program for performing one of the methods described herein. The data stream or the sequence of signals can for example be arranged to be transmitted via a data communications connection, for example via the Internet.

[0124] A further embodiment comprises a processing means, for example a computer, or a programmable logic device configured to or adapted to perform one of the methods described herein.

[0125] A further embodiment comprises a computer having installed thereon a computer program for performing one of the methods described herein.

[0126] Further embodiments according to the invention comprise an apparatus or a system configured to transmit (e.g. electronically or optically) a computer program for performing one of the methods described herein for a receiver. The receiver may be, for example, a computer, a mobile device, a memory device, etc. The apparatus or system may comprise, for example, a file server for transmitting the computer program for the receiver.

[0127] In some embodiments, a programmable logic device (e.g., a field programmable gate array) may be used to perform some or all of the functionality of the methods described herein. In some embodiments, a field programmable gate array may cooperate with a microprocessor to perform one of the methods described herein. In general, the methods are preferably performed by any hardware device.

[0128] The above-described embodiments are merely illustrative of the nature of the present invention. It is to be understood that improvements and variations of the arrangements and details described herein will be apparent to others skilled in the art. It is therefore intended to be limited only by the scope of the appended claims, and not by the specific details provided by the method of description and interpretation of the embodiments herein.

Claims

1. A decoder comprising a processor configured to decode a multilayer video data stream containing encoded video content represented by multiple layers, wherein the multilayer video data stream includes a sequence of network abstraction layer (NAL) units divided into a sequence of access units, Each access unit includes a plurality of decoding units, each of which includes at least one NAL unit corresponding to one of the plurality of layers, and the multilayer video data stream is First timing control information (800) that signals the first decoder buffer search time of a certain access unit from the sequence of access units, Second timing control information (802), separate from the first timing control information, signals the second decoder buffer search time of each decoding unit of a particular access unit according to a specific order of each decoding unit in the multilayer video data stream. A decoder that includes this.

2. The decoder according to claim 1, wherein the multilayer video data stream includes third timing control information (804) that signals the third decoder buffer lookup time of each decoding unit of a particular access unit according to a defined layer order for the plurality of layers, and according to the layer order, in the access unit, a first decoding unit associated with a first layer precedes a second decoding unit associated with a second layer that follows the first layer based on the layer order.

3. The decoder according to claim 2, wherein the decoding units associated with different layers among the plurality of layers are interleaved in the multilayer video data stream.

4. The decoder according to claim 3, wherein the processor is configured to deinterleave and remove the decoding units of the access units in order of the layers when emptying the decoder buffer in units of access units.

5. The decoder according to claim 1, further comprising a plurality of buffers, each of the plurality of buffers for storing a decoding unit for one of the plurality of layers, and the processor is configured to buffer the decoding unit for each layer, including the NAL unit associated with each layer in the corresponding buffer for each layer.

6. The decoder according to claim 1, further comprising a plurality of buffers, each of the plurality of buffers for storing decoding units of a subset of layers, and the processor is configured to buffer the decoding units for each layer, including NAL units associated with each layer in the corresponding buffer of the subset of layers to which each layer belongs.

7. The processor is The decoder buffer for buffering the multilayer video data stream is emptied on an access unit basis based on the first timing control information (800). In response to the second timing control information, the decoder buffer for buffering the multilayer video data stream is emptied in units of the decoding unit according to the specific order based on the second decoder buffer lookup time (802). In response to the third timing control information (804), the decoder buffer for buffering the multilayer video data stream is emptied in units of the decoding unit according to the order of the layers based on the third decoder buffer lookup time. The decoder according to claim 2, configured as follows.

8. The decoder according to claim 1, wherein the second timing control information (802) is decoded independently of the first timing control information.

9. An encoder comprising a processor configured to encode video content into multiple layers of a multilayer video data stream, wherein the multilayer video data stream includes a sequence of network abstraction layer (NAL) units, which are divided into a sequence of access units. Each access unit includes a plurality of decoding units, and each of the plurality of decoding units includes at least one NAL unit corresponding to one of the plurality of layers. The encoder is configured to encode the first timing control information and the second timing control information into the multilayer video data stream, where, The first timing control information (800) signals the first decoder buffer lookup time of a certain access unit from the sequence of access units, The second timing control information, separate from the first timing control information, signals the second decoder buffer search time of each decoding unit of a particular access unit according to a specific order of each decoding unit in the multilayer video data stream. encoder.

10. The encoder according to claim 9, wherein the multilayer video data stream includes third timing control information (804) that signals the third decoder buffer lookup time of each decoding unit of a particular access unit according to a defined layer order for the plurality of layers, and according to the layer order, in the access unit, a first decoding unit associated with a first layer precedes a second decoding unit associated with a second layer that follows the first layer based on the layer order.

11. The encoder according to claim 10, wherein decoding units associated with different layers among the plurality of layers are interleaved in the multilayer video data stream.

12. The encoder according to claim 9, wherein the second timing control information is encoded independently of the first timing control information.

13. A method for encoding video content into multiple layers of a multilayer video data stream, comprising the step of encoding first timing control information and second timing control information into a multilayer video data stream, wherein the multilayer video data stream includes a sequence of network abstraction layer (NAL) units, which are divided into a sequence of access units. Each access unit includes a plurality of decoding units, and each of the plurality of decoding units includes at least one NAL unit corresponding to one of the plurality of layers. The first timing control information signals the first decoder buffer search time of a particular access unit from the sequence of access units, The second timing control information, separate from the first timing control information, signals the second decoder buffer search time of each decoding unit of a particular access unit according to a specific order of each decoding unit in the multilayer video data stream. method.

14. The method according to claim 13, wherein the multilayer video data stream includes third timing control information (804) that signals the third decoder buffer lookup time of each decoding unit of a particular access unit according to a defined layer order for the plurality of layers, and according to the layer order, in the access unit, a first decoding unit associated with a first layer precedes a second decoding unit associated with a second layer that follows the first layer, based on the layer order.

15. The method according to claim 14, wherein decoding units associated with different layers among the plurality of layers are interleaved in the multilayer video data stream.

16. The method according to claim 13, wherein the second timing control information is encoded independently of the first timing control information.

17. A method for storing a data stream, comprising the step of storing a multilayer video data stream in a digital storage medium, wherein video content is encoded as multiple layers by the method of claim 13.