Low-Delay Concepts in Multi-Layer Video Coding
By employing interleaved multi-layer video data streams with additional timing control, the solution addresses low-delay coding challenges, enhancing decoder efficiency and reducing end-to-end delay in multi-view/layer coding applications.
Patent Information
- Application Number
- JP2021172080
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2013-07-15
- Filing Date
- 2021-10-21
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2034-07-15
AI Technical Summary
Existing multi-view/layer coding technologies face challenges in achieving low delay between terminals without compromising decoder capabilities, particularly in scenarios where decoders cannot handle or utilize low delay concepts.
The implementation of an interleaved multi-layer video data stream with additional timing control information that allows for either unit-wise decoding buffer access or an intermediate procedure to reverse the interleaving of decoding units across layers, enabling efficient low-delay decoding without requiring interleaving.
This approach enables low-delay multi-view/layer coding by allowing decoders to handle decoding units efficiently, supporting parallel processing and reducing end-to-end delay in applications like videoconferencing and industrial surveillance.
Smart Images

Figure 0007745417000001 
Figure 0007745417000002 
Figure 0007745417000003
Abstract
Description
[Technical Field]
[0001] This application relates to coding concepts that enable efficient multi-view / layer like multi-view picture / video coding. [Background technology]
[0002] The Coded Picture Buffer (CPB) in Scalable Video Coding (SVC) operates on complete Access Units (AUs). One AU for every Network Abstraction Layer Unit (NALU) is moved from the Coded Picture Buffer (CPB) at the same time. One AU contains all layer packets (i.e., NALUs).
[0003] The concept of a Decoding Unit (DU) in the HEVC base specification [1] is expanded compared to H.264 / AVC. A DU is a group of NAL units at consecutive positions in the bitstream. The NAL units belong to the same layer, i.e., the so-called base layer.
[0004] The HEVC base specification includes the tools necessary to enable bitstream decoding using ultra-low delay, i.e., by means of CPB operation at the DU level (as opposed to the AU level as in H.264 / AVC) and CPB timing information with DU granularity. Thus, a device can operate on sub-portions of a picture to reduce the processing delay incurred. Just as ultra-low delay operates in multi-layer SHVC, MV-HEVC, and the HEVC extension 3D-HEVC, CPB must operate at the DU level across multiple layers and be appropriately defined. In particular, bitstreams using several layers or views require DUs of one AU to be interleaved across multiple layers. That is, DUs of layer m of a given AU can follow DUs of layer (m+1) of the same AU in such an ultra-low delay-enabled multi-layer bitstream, as long as they are independent of the DUs that follow them in bitstream order.
[0005] Ultra-low delay operation requires modifications of CPB operation for multi-layer decoders compared to SVC and MVC in H.264 / AVC extensions, which work based on AUs. Ultra-low delay decoders may use additional timing information, for example provided by way of SEI messages.
[0006] Some implementations of a single multi-layer decoder may prefer layer-wise decoding (and CPB operation at either the DU or AU level), i.e., decoding of layer m before decoding of layer m+1, which would effectively hinder any multi-layer ultra-low latency applications using SHVC, MV-HEVC, and 3D-HEVC unless new mechanisms are provided.
[0007] Currently, the HEVC base specification includes two decoding modes of operation. Access Unit (AU) based decoding: All decoding units of an access unit are moved from the CPB at the same time. Decoding-based Decoding Unit (DU): Each decoding unit has its own CPB removal time. [Prior art documents] [Non-patent literature]
[0008] [Non-Patent Document 1] B. Bross, W.-J. Han, J.-R. Ohm, GJ Sullivan, T. Wiegand (Eds.), "High Efficiency Video Coding (HEVC) text specification draft 10", JCTVC-L1003, Geneva, CH, Jan. 2013 [Non-patent document 2] G. Tech, K. Wegner, Y. Chen, M. Hannuksela, J.Boyce (Eds.), "MV-HEVC Draft Text 3 (ISO / IEC 23008-2 PDAM2)", JCT3V-C1004, Geneva, CH, Jan. 2013 [Non-patent document 3] G. Tech, K. Wegner, Y. Chen, S. Yea (Eds.), "3D-HEVC Test Model Description, draft specification", JCT3V-C1005, Geneva, CH, Jan. 2013 [Non-patent document 4] WILBURN, Bennett, et al. High performance imaging using large camera arrays. ACM Transactions on Graphics, 2005, 24. Jg., Nr. 3, S. 765-776. [Non-patent document 5] WILBURN, Bennett S., et al. Light field video camera. In:Electronic Imaging 2002. International Society for Optics and Photonics, 2001. S. 29-36. [Non-patent document 6] HORIMAI, Hideyoshi, et al. Full-color 3D display system with 360 degree horizontal viewing angle. In:Proc. Int. Symposium of 3D and Contents. 2010. S. 7-10. Summary of the Invention [Problem to be solved by the invention]
[0009] Nevertheless, it would be more advantageous to have concepts within reach that further improve the multi-view / layer coding concept.
[0010] Accordingly, it is an object of the present invention to provide a concept that further improves the multi-view / layer coding concept, in particular to provide the possibility of enabling low delay between terminals, but without abandoning at least one alternative where the decoder cannot handle or does not use the low delay concept.
[0011] This object is achieved by the subject matter of the pending independent claims. [Means for solving the problem]
[0012] The basic idea of the present application is to provide an interleaved multi-layer video data stream with interleaved decoding units of different layers using additional timing control information in addition to timing control information reflecting the interleaved decoding unit arrangement. The additional timing control information relates to either an alternative according to which all decoding units of an access unit are handled in unit-wise decoding buffer access, or an alternative according to which an intermediate procedure is used and the interleaving of DUs of different layers is reversed according to additionally transmitted timing control information, thereby enabling DU-wise handling in the decoder buffer, but without interleaving of decoding units for different layers. Both alternatives may also exist. Various advantageous embodiments and alternatives are the subject of the various claims appended hereto.
[0013] Preferred embodiments of the present application are described below with reference to the drawings. [Brief explanation of the drawings]
[0014] [Figure 1] 1 shows a video encoder that serves as an example for implementing any of the multi-layer encoders outlined further with respect to the following figures. [Figure 2] 2 shows a simplified block diagram of a video decoder suitable for the video encoder of FIG. 1; [Figure 3] 1 shows a schematic diagram of a picture being subdivided into sub-streams for WPP processing. [Figure 4] 1 shows a diagram illustrating a picture of several layers subdivided into blocks showing a further subdivision of the picture into spatial segments. [Figure 5] 1 shows a schematic representation of a picture of several layers, subdivided into blocks and tiles. [Figure 6] 1 shows a schematic diagram of a picture subdivided into blocks and sub-streams. [Figure 7] Here is shown a schematic diagram of a multi-layer video data stream exemplarily comprising three layers, where options 1 and 2 for arranging NAL units belonging to each time and each layer within the data stream are illustrated in the lower half of Figure 7. [Figure 8] A schematic representation of part of one data stream is presented by illustrating these two options in the exemplary case of two layers. [Figure 9] As a comparative embodiment, a simplified block diagram of a decoder configured to process a multi-layer video data stream according to FIGS. 7 and 8 of Option 1 is shown. [Figure 10] 10 shows a simplified block diagram of an encoder suitable for the decoder of FIG. 9. [Figure 11] An example syntax of part of the VPS syntax extension including a flag indicating interleaved transmission of DUs of changing layers is shown. [Figure 12a] 10 illustrates an exemplary syntax of an SEI message containing timing control information that enables DU de-interleaving of DUs delivered interleaved from a buffer of a decoder according to one embodiment. [Figure 12b] 12a shows an exemplary syntax of an SEI message that is to be interspersed at the beginning of interleaved DUs and that also carries timing control information according to an alternative embodiment. [Figure 12c] 10 illustrates an exemplary syntax of an SEI message that reveals timing control information that enables buffer search for a DU when maintaining interleaving DUs according to one embodiment. [Figure 13] Shown is a schematic representation of the bitstream order of DUs of three layers over time, with layer indices exemplified as registered numbers 0 to 2. [Figure 14] Compared with FIG. 13, a schematic diagram of the bitstream order of the DUs of the three layers in which the DUs are interleaved over time is shown. [Figure 15] 1 shows a diagram illustrating the distribution of multi-layer DUs for multiple CPBs according to one embodiment. [Figure 16] 10 shows a diagram illustrating memory address illustrations for multiple CPBs according to one embodiment. [Figure 17] 8 shows a schematic diagram of a multi-layer video data stream with an encoder modified in relation to FIG. 7 to accommodate an embodiment of the present application; [Figure 18] 8 shows a schematic diagram of a multi-layer video data stream with an encoder modified in relation to FIG. 7 to accommodate an embodiment of the present application; [Figure 19] 8 shows a schematic diagram of a multi-layer video data stream with an encoder modified in relation to FIG. 7 to accommodate an embodiment of the present application; [Figure 20] 8 shows a schematic diagram of a multi-layer video data stream with an encoder modified in relation to FIG. 7 to accommodate an embodiment of the present application; [Figure 21] 10 shows a schematic diagram of a multi-layer video data stream together with a modified decoder associated with the one exemplarily shown in FIG. 9 to correspond to an embodiment of the present application. [Figure 22] 10 shows a schematic diagram of a multi-layer video data stream together with a modified decoder associated with the one exemplarily shown in FIG. 9 to correspond to an embodiment of the present application. [Figure 23]10 shows a schematic diagram of a multi-layer video data stream together with a modified decoder associated with the one exemplarily shown in FIG. 9 to correspond to an embodiment of the present application. [Figure 24] 10 shows a schematic diagram of a multi-layer video data stream together with a modified decoder associated with the one exemplarily shown in FIG. 9 to correspond to an embodiment of the present application. [Figure 25] 1 shows a block diagram illustrating an intermediate network located upstream from a decoder buffer. DETAILED DESCRIPTION OF THE INVENTION
[0015] First, as an overview, an example of an encoder / decoder structure is presented, suitable for the embodiments presented later, i.e., an encoder can be implemented to utilize the concepts outlined later, and the same applies for the decoder.
[0016] FIG. 1 shows the overall structure of an encoder according to one embodiment. The encoder 10 may be implemented in a multi-threaded manner or not, i.e., capable of operating solely in a single thread. That is, the encoder 10 may be implemented, for example, using multiple CPU cores. In other words, the encoder 10 may support parallel processing, but it need not. The generated bitstream could also be generated / decoded by a single-threaded encoder / decoder. The encoding concept of this application, however, allows parallel processing to be applied efficiently to a parallel processing encoder, but without compromising compression efficiency. With regard to parallel processing capabilities, similar statements are valid for the decoder described later with reference to FIG. 2.
[0017] Encoder 10 is a video encoder. Picture 12 of video 14 is shown entering encoder 10 at input 16. Picture 12 represents a particular scene, i.e., picture content. However, encoder 10 also receives at its input 16 another picture 15 related to the same moment as both pictures 12 and 15, but belonging to a different layer. For illustrative purposes only, picture 12 is shown as belonging to layer 0, while picture 15 is shown as belonging to layer 1. FIG. 1 illustrates that layer 1 may include a higher spatial resolution relative to layer 0, i.e., represent the same scene with a greater number of picture samples. However, this is for illustrative purposes only, and picture 15 of layer 1 may alternatively have the same spatial resolution but differ, for example, in view direction relative to layer 0. That is, pictures 12 and 15 may be captured from different viewpoints. It is noted that the terms base and extended used in this document may refer to either a reference and a dependent layer in a layer hierarchy.
[0018] Encoder 10 is a hybrid encoder, i.e., pictures 12 and 15 are predicted by a predictor 18 of encoder 10 and a prediction residual 20 obtained by a residual determiner 22 of encoder 10, and are subjected to a transformation, such as a spectral decomposition, such as a DCT, and quantization in a transform / quantization module 24 of encoder 10. A transformed and quantized prediction residual 26 thus obtained is subjected to entropy coding, such as arithmetic coding with context adaptivity or variable-length coding, in an entropy encoder 28. A reconstructable version of the residual is available to the decoder: a dequantized and retransformed residual signal 30 is recovered by a retransform / requantization module 31 and recombined with a prediction signal 32 of predictor 18 by a combiner 33, thereby resulting in a reconstruction 34 of pictures 12 and 15, respectively. However, encoder 10 operates block-wise. Accordingly, reconstructed signal 34 suffers from discontinuities at block boundaries, and accordingly, filter 36 may be applied to reconstructed signal 34 to produce reference pictures 38 based on which predictor 18 predicts the next coded picture of a different layer, for pictures 12 and 15, respectively. As indicated by the dashed line in FIG. 1 , predictor 18, however, may also utilize reconstructed signal 34 directly, without filter 36 or an intermediate version, as in other prediction modes, such as spatial prediction mode.
[0019] Predictor 18 may select among different prediction modes for predicting a particular block of picture 12. One such block 39 of picture 12 is representative of any block of picture 12 into which picture 12 is partitioned, as exemplarily shown in FIG. 1 . There may be a temporal prediction mode according to which block 39 is predicted based on a pre-coded picture of the same layer, such as picture 12′. A spatial prediction mode exists according to which block 39 is predicted based on pre-coded portions of neighboring blocks 39 of the same picture 12. Block 41 of picture 15 is also exemplarily shown in FIG. 1 because it is representative of any other block into which picture 15 is partitioned. For block 41, predictor 18 may support the prediction modes just discussed, i.e., temporal and spatial prediction modes. In addition, predictor 18 may provide for an inter-layer prediction mode according to which block 41 is predicted based on a corresponding portion of picture 12 of a lower layer. The "corresponding" in "corresponding portion" means spatial correspondence, i.e., a portion in picture 12 that shows the same part of the scene as predicted block 41 in picture 15.
[0020] The prediction of the predictor 18 may of course not be limited to picture samples. Prediction may also be applied to any coding parameter, i.e., prediction mode, motion vectors in temporal prediction, disparity vectors in multiview prediction, etc. Simply, the residual may then be coded in the bitstream 40. That is, coding parameters may be predictively coded / decoded using spatial and / or inter-layer prediction. Even here, disparity compensation may be used.
[0021] A specific syntax is used to compile quantized residual data 26, i.e., to convert both coefficient levels and other residual data as well as coding parameters, including, for example, prediction modes and prediction parameters, for individual blocks 39 and 41 of pictures 12 and 15 as determined by predictor 18. Syntax elements of this syntax are also subject to entropy coding by entropy encoder 28. The data stream thus obtained as output by entropy encoder 28 forms the bitstream 40 output by encoder 10.
[0022] FIG. 2 shows a decoder adapted to the encoder of FIG. 1, i.e., a decoder capable of decoding the bitstream 40. The decoder of FIG. 2 is generally indicated by a reference signal 50 and comprises an entropy decoder, a transform / inverse quantization module 54, a combiner 56, a filter 58, and a predictor 60. The entropy decoder 42 receives the bitstream and performs entropy decoding to receive residual data 62 and coding parameters 64. The retransform / inverse quantization module 54 inverse quantizes and retransforms the residual data 62 and forwards the thus obtained residual signal to the combiner 56. The combiner 56 also receives a prediction signal from the predictor 60. The predictor 60, in turn, forms a prediction signal using the coding parameters 64 based on a reconstructed signal 68 determined by the combiner 56 by combining the prediction signal 66 and the residual signal 65. The prediction mirrors the prediction ultimately selected from the predictor 18. That is, the same prediction modes are available, and these modes are selected for individual blocks of pictures 12 and 15 and steered according to prediction parameters. As already explained above with respect to Figure 1, predictor 60 may use a filtered version of reconstructed signal 68, or alternatively or in addition, the same intermediate versions thereof. The pictures of the different layers to be finally reproduced and output at output 70 of decoder 50 are likewise determined on the unfiltered version of combined signal 68, or the same filtered versions thereof.
[0023] Encoder 10 of FIG. 10 supports the tile concept. According to the tile concept, pictures 12 and 15 are subdivided into tiles 80 and 82, respectively, and within these tiles 80 and 82, predictions of at least blocks 39 and 41, respectively, are restricted to using only data related to the same tile of the same picture as the basis for spatial prediction. This means that the spatial prediction of block 39 is restricted to using pre-coded portions of the same tile, but the temporal prediction mode is not restricted to relying on information from pre-coded pictures such as picture 12′. Similarly, the spatial prediction mode of block 41 is restricted to using pre-coded data only from the same tile, but the temporal and inter-layer prediction modes are not restricted. Predictor 60 of decoder 50 is also configured to handle tile boundaries; in particular, the selection and / or adaptation of predictor and entropy context is performed only within one tile, without crossing any tile boundaries.
[0024] The subdivision of pictures 15 and 12 into six tiles, respectively, was chosen solely for illustrative purposes. The subdivision into tiles can be selected and signaled individually in bitstream 40 for pictures 12, 12′ and 15, 15′, respectively. The number of tiles for pictures 12 and 15 can be 1, 2, 3, 4, 6, etc., respectively. The partitioning of tiles can be limited to regular partitioning into rows and columns of tiles only. For completeness, note that the scheme for separately coding tiles can be limited to intra-prediction or spatial prediction, but can also encompass any prediction of coding parameters across tile boundaries and context selection in entropy coding. That is, the latter can also be limited to relying only on data from the same tile. Thus, the decoder can perform the operations just described in parallel, i.e., in units of tiles.
[0025] The encoder and decoder of Figures 1 and 2 may alternatively or additionally use / support the WPP (wavefront parallel processing) concept. Referring to Figure 3, WPP substream 100 represents a spatial partitioning of pictures 12, 15 into WPP substreams. In contrast to tiles and slices, WPP substreams impose no restrictions on prediction and context selection across the WPP substream 100. The WPP substream 100 extends row-wise across LCUs (Largest Coding Units) 101, i.e., rows of blocks for which a predictive coding mode is most likely to be sent individually in the bitstream. Also, to enable parallel processing, only one compromise is made with respect to entropy coding. In particular, instructions 102 are defined within WPP substream 100, which exemplarily lead from top to bottom, and for each WPP substream 100, except for the first WPP substream, in instruction 102, the probability estimates for the symbol alphabet, i.e., entropy probabilities, are completely reset, but each other, in the LCU row direction, as indicated by arrow 106 and leading to the same side of pictures 12 and 15, respectively, as shown on the left-hand side, using an LCU instruction, or substream decoder instruction, for each WPP substream, as indicated by line 104, where the immediately preceding WPP substream, up to the second LCU, has its entropy taken from the resulting probability with the encoded / decoded entropy equal to, or equal to. Accordingly, by following some coding delay between sequences of WPP substreams of the same pictures 12 and 15, these WPP substreams 100 can be decoded / encoded in parallel, with portions being coded / decoded in parallel in each picture 12, 15, i.e., together forming a kind of wavefront 108 that moves across the picture as if tiled from left to right.
[0026] Briefly note that instructions 102 and 104 also define a raster scan order among the LCUs, leading row by row from the top-left LCU 101 to the bottom-right LCU. A WPP substream corresponds to one LCU row each. Briefly referring back to tiles, the latter may also be restricted to being aligned to an LCU boundary. A substream may be broken into one or more slices without being tied to an LCU boundary as far as the boundary between two slices inside the substream is concerned. Entropy probability, however, is employed in this case when passing from one slice of a substream to the next. In the case of tiles, the entire tile may be collapsed into one slice, or one tile may be broken into one or more slices that are not re-attached to an LCU as far as the boundary between two slices inside the tile is concerned. In the case of tiles, the order among the LCUs is modified to first traverse the tile in the raster scan order before proceeding to the next tile in the tile order.
[0027] As explained so far, picture 12 may be partitioned into tiles or WPP substreams, and likewise picture 15 may be partitioned into tiles or WPP substreams. In theory, the partitioning / concept of the WPP substreams can be selected for one of pictures 12 and 15 when the partitioning / concept of the tiles is selected for the two. Alternatively, a restriction can be imposed on the bitstream according to the concept type, i.e., that the tiles or WPP substreams should be the same within a layer.
[0028] Another example of a spatial segment encompasses a slice. Slices are suitable for dividing the bitstream 40 for transmission purposes. Slices are packed into NAL units, which are the smallest entities for transmission. Each slice is independently codable / decodable; that is, any prediction across slice boundaries is prohibited, as are context selection, etc.
[0029] These are just three examples of spatial segments: slices, tiles, and WPP substreams. In addition, all three parallelization concepts, tiles, WPP substreams, and slices, can be used in combination. That is, picture 12 or picture 15 can be divided into tiles, and each tile can be divided into multiple WPP substreams. Slices can also be used to partition the bitstream into multiple NAL units, for example (but not limited to) at tile or WPP boundaries. If pictures 12 and 15 are partitioned using tiles or WPP substreams and, in addition, slices, and the partitioning of slices deviates from the partitioning of other WPPs / tiles, then a spatial segment would be defined as the smallest independently decodable section of picture 12 and 15. Alternatively, restrictions can be imposed on the bitstream that combinations of concepts can be used within a picture (12 or 15) and / or if boundaries should be aligned between the different used concepts.
[0030] Various prediction modes are supported by the encoder and decoder, as well as by restrictions imposed on the prediction modes and context derivation, to enable parallel processing concepts such as the tile and / or WPP concepts described above. It was also mentioned above that the encoder and decoder may operate on a block-by-block basis. For example, the prediction modes described above are selected on a block-by-block basis, i.e., at a finer granularity than the picture itself. Before proceeding with the described aspects of the present application, the relationship between slices, tiles, WPP sub-streams, and the just-mentioned blocks according to one embodiment will be explained.
[0031] FIG. 4 illustrates a picture, which may be a layer 0 picture, such as layer 12, or a layer 1 picture, such as picture 15. The picture is regularly subdivided into an array of blocks 90. These blocks 90 are sometimes referred to as largest coding blocks (LCBs), largest coding units (LCUs), coding tree blocks (CTBs), etc. This subdivision of the picture into blocks 90 may form a type of basis or coarsest granularity upon which the prediction and residual coding described above is performed. This coarsest granularity, i.e., the size of the blocks 90, may also be signaled and set by the encoder independently for layers 0 and 1. For example, a multi-tree, such as a 4-tree subdivision, may be used and signaled within the data stream to subdivide each block 90 into a prediction block, a residual block, and / or a coding block, respectively. In particular, the coding block may be a leaf block of a recursive multi-tree subdivision of block 90, and some prediction relationship decisions, such as the prediction mode, are signaled at the granularity of the coding block, and the prediction block is coded at the granularity of prediction parameters, such as the motion vector in the case of inter-layer prediction, for example, temporal inter-prediction and the disparity vector in the case of inter-layer prediction, and the residual block may be a leaf block of a multi-tree subdivision of a separate recursive code block, at the granularity of the prediction residual being coded.
[0032] Raster scan encoding / decoding instructions 92 may be defined within block 90. The encoding / decoding instructions 92 limit the availability of neighboring portions for spatial prediction purposes: only portions of the picture according to the encoding / decoding instructions 92 that precede the current portion, such as block 90 or some smaller block thereof, are available for spatial prediction within the current picture, since the currently predicted syntax element is associated with that portion. Within each layer, the encoding / decoding instructions 92 traverse the entire block 90 of the picture, then continue with traversing blocks of the next picture in each layer, in a picture encoding / decoding instruction that does not necessarily follow the temporal regeneration of the picture. Within each block 90, the encoding / decoding instructions 92 are refined to scan within smaller blocks, such as encoding blocks.
[0033] With respect to the blocks 90 and smaller blocks just outlined, each picture is further subdivided into one or more slices in accordance with the just-described encoding / decoding instructions 92. Slices 94a and 94b exemplarily shown in FIG. 4 cover each picture gaplessly accordingly. The boundary or interface 96 between consecutive slices 94a and 94b of one picture may or may not be aligned with the boundary of adjacent blocks 90. More precisely, consecutive slices 94a and 94b within one picture, as illustrated on the right-hand side of FIG. 4, may abut each other at the boundary of a smaller block, such as a coding block, i.e., a leaf block of one subdivision of block 90.
[0034] Slices 94a and 94b of a picture may form the smallest unit in the portion of the data stream into which the picture is coded and may be packetized into packets, i.e., NAL units. Further possible properties of slices, i.e., restrictions on slices with respect to, for example, prediction across slice boundaries and entropy context determination, have been described above. Slices with such restrictions may be referred to as "standard" slices. In addition to standard slices, "dependent slices" may also exist, as outlined in more detail below.
[0035] The encoding / decoding instructions 92 defined within the array of blocks 90 may change if a tile partitioning concept is used for the picture. This is exemplarily shown in FIG. 5, where a picture is shown partitioned into four tiles 82a-82d. As illustrated in FIG. 5, the tiles themselves are defined as regular divisions of the picture in units of blocks 90. That is, each tile 82a-82d is composed of an array of n-by-m blocks 90, with n set individually for each row of the tile and m set individually for each column of the tile. Following the encoding / decoding instructions 92, the blocks 90 in the first tile are first scanned in a raster scan order, before proceeding with the next tile 82b, etc., where tiles 82a-82d are themselves scanned in a raster scan order.
[0036] According to the WPP stream partitioning concept, a picture is subdivided into WPP substreams 98a-98d in units of one or more rows of a block 90 along the encoding / decoding instructions 92. Each WPP substream covers one complete row of a block 90, for example, as illustrated in FIG.
[0037] The tile concept and the WPP substream concept may, however, also be mixed, in which case each WPP substream covers, for example, one row of blocks 90 within each tile.
[0038] Even slice partitioning of a picture can be used in common with tile partitioning and / or WPP substream partitioning. In terms of tiles, one or more slices of a picture can each be subdivided into one complete tile, one or more complete tiles, or simply a subportion of a tile according to the encoding / decoding instructions 92. Slices can also be used to form WPP substreams 98a-98d. To this end, slices that form the smallest units for packetization can comprise standard slices on the one hand and dependent slices on the other hand: while standard slices impose the above-described restrictions on prediction and entropy context derivation, dependent slices do not impose such restrictions. A dependent slice that starts at a picture boundary because the encoding / decoding instructions 92 are far enough away row-wise to adopt the entropy context resulting from the entropy decoding block 90 in the row immediately preceding block 90. Also, a dependent slice starting somewhere else may adopt the entropy coding context resulting from entropy coding / decoding of the immediately preceding slice up to its end. In this manner, each of the WPP substreams 98a-98d may consist of one or more dependent slices.
[0039] That is, the encoding / decoding instructions 92 defined in the blocks 90 here lead linearly from the first side of each picture, exemplarily on the left, to the opposite side, exemplarily on the right, and down to the next row of the blocks 90 in a downward / bottom direction. The available, i.e., previously encoded / decoded portions of the current picture, are accordingly located primarily to the left and above the current encoding / decoding portion, as in the current block 90. Due to the disruption of prediction and entropy context derivation across tile boundaries, tiles of one picture can be processed in parallel. The encoding / decoding of tiles of one picture can even be initiated simultaneously. The limitations arising from the aforementioned in-loop filtering in the case of identical tiles are allowed across tile boundaries. Initiating the encoding / decoding of WPP substreams is performed in a staggered fashion from top to bottom. The intra-picture delay between successive WPP substreams is measured in multiple blocks 90, or two blocks 90.
[0040] However, it may be preferable to even process the encoding / decoding of pictures 12 and 15 in parallel, i.e., at different layer instants. Clearly, the encoding / decoding of picture 15 of the dependent layer must be delayed relative to the encoding / decoding of the base layer to ensure that a "spatially corresponding" portion of the base layer is already available. These considerations are valid even without parallel processing of any of the encoding / decoding within either of pictures 12 and 15 individually. Non-tiled and non-WPP substream processing, encoding / decoding of pictures 12 and 15, respectively, can be used to parallelize even when using one slice to cover the entire pictures 12 and 15. The signaling described next, i.e., the sixth aspect, makes it possible to express such encoding / decoding delays between multiple layers even in such cases, or regardless of whether tiled or WPP processing is used for any of the layer pictures.
[0041] Before discussing the above-mentioned concepts of the present application, please refer again to Figures 1 and 2 and note that the block structures of the encoder and decoder in Figures 1 and 2 are for illustrative purposes only and the structures may also differ.
[0042] There are applications such as videoconferencing and industrial surveillance where end-to-end delay is as low as possible, but multi-layer (scalable) coding would be beneficial. The embodiments described in more detail below allow for lower end-to-end delay in multi-layer video coding. In this regard, it should also be noted that the embodiments described below are not limited to multi-view coding. The layers described below may include different views, but may also represent the same view with varying degrees of spatial resolution, such as SNR accuracy. Possible scalable dimensions along the layers discussed below increase the information content conveyed by varying the number of views, spatial resolution, and SNR accuracy of the previous layer.
[0043] As explained above, NAL units are composed of slices. The tile and / or WPP concepts can be freely selected individually for different layers of a multi-layer video data stream. Accordingly, each NAL unit having slices packetized therein can be spatially attributed to the area of the picture to which each slice refers. Accordingly, to enable low-delay encoding in the case of intra-layer prediction, it would be preferable for the encoder and decoder to be able to interleave NAL units of different layers that coincide at the same time, so as to enable parallel processing of these pictures of different layers and to start encoding, transmitting, and decoding, respectively, slices packeted into these NAL units in a concomitant manner at the same time. However, depending on the application, an encoder may be better off, beyond the ability to enable parallel processing in the layer dimension, with the ability to use different coding instructions among pictures of different layers, such as using different GOP structures for different layers. The syntax of a data stream according to a comparative embodiment is described below with reference to FIG. 7.
[0044] 7 shows a multi-layer video material 201 consisting of a sequence of pictures 204, each of which represents a different layer. Each layer may describe a different characteristic of the scene (video content) described by the multi-layer video material 201. That is, the significance of a layer may be selected, for example, among color components, depth maps, and / or viewpoints. Without loss of generality, we refer to the video material 201 as a multi-view video, with different layers corresponding to different viewpoints.
[0045] For applications requiring low latency, the encoder may decide to signal long-term high-level syntax elements. In that case, the data stream generated by the encoder may look like the one shown in the center of FIG. 7, with a circle around it. In that case, the multi-layer video stream 200 is composed of a sequence of NAL units 202, such as NAL units 202 belonging to one access unit 206 for a picture of one temporal instant and NAL units 202 of different access units for different instants. That is, the access unit 206 collects the NAL units 202 of one instant, i.e., the ones associated with the access unit 206. Within each access unit 206, for each layer, at least some of the NAL units for each layer are grouped into one or more decoding units 208. This means that among the NAL units 202, there are subsequently different types of NAL units, such as VCL NAL units on the one hand and non-VCL NAL units on the other hand, as shown above. More particularly, the NAL units 202 may be of different types, and these types may comprise:
[0046] 1) NAL units that carry syntax elements related to slices, tiles, WPP substreams, etc., i.e., prediction parameters and / or residual data that describe the picture content in terms of picture sample scale / granularity. There can be one or more such types. The VCL NAL unit is of such a type. Such NAL units are removable. 2) Parameter set NAL units may carry information that changes infrequently, such as the long-term coding settings described above. Such NAL units may be interspersed in the data stream, for example, to some extent and repeatedly. 3) Supplemental Enhancement Information (SEI) NAL units can carry arbitrary data.
[0047] As an alternative to the term "NAL unit", "packet" is sometimes used to denote the first type of NAL unit, i.e., a VCL unit, followed by a "payload packet", while "packet" also encompasses non-VCL units, to which packets of types 2 and 3 in the above list belong.
[0048] A decoding unit may consist of the first of the NAL units mentioned above. More precisely, a decoding unit may consist of one or more VCL NAL units in an access unit and associated non-NAL units. A decoding unit therefore describes a particular area of a picture, i.e., an area that is coded into one or more slices contained therein.
[0049] The decoding units 208 for NAL units associated with different layers are interleaved because, for each decoding unit, the intra-layer prediction previously used to encode the decoding unit is based on a portion of a picture of a layer other than the layer to which the decoding unit is associated, which portion is encoded in the decoding unit preceding the decoding unit in the access unit. See, for example, decoding unit 208a in FIG. 7. Exemplarily, imagine an area 210 of a picture of dependent layer 2 and this decoding unit for a particular instant. The co-located area in the base layer picture for the same instant is denoted by 212, and an area of this base layer picture slightly exceeding this area 212 may be required to fully decode decoding unit 208a using intra-layer prediction. This slight excess may be the result of, for example, disparity-compensated prediction. This similarly means that the decoding units 208b preceding decoding unit 208a in access unit 206 should completely cover the area required for intra-layer prediction. Reference is made to the above discussion regarding delay indications that can be used like boundaries for interleaving granularity.
[0050] However, if an application takes advantage of the freedom to select different decoding orders for pictures in different layers, the encoder may prefer the case depicted at the bottom of FIG. 7 with its two circles around it. In this case, the multi-layer video data stream has individual access units for each picture belonging to one or more specific combinations of layer ID values and a single temporal instant. As shown in FIG. 7, at the (i-1)th decoder instruction, i.e., instant t(i-1), each layer may consist of access units AU1, AU2 (and so on), or none (cp instant t(i)), in which all layers are contained in a single access unit AU1. However, interleaving is not possible in this case. Access units are arranged in the data stream 200 following the access units of decoding instruction index i, i.e., the decoding instruction i for each layer, followed by the access units associated with the pictures of those layers corresponding to decoding instruction i+1, etc. Temporal intra-layer prediction signaling in a data stream signal for either the same coding order or different picture coding orders refers to different layers. Also, the signaling can even be located overlapping in more than one position in the data stream, for example, in a slice packetized into NAL units. In other words, case 2 subdivides the access unit scope: a separate access unit is opened for each combination of moment and layer.
[0051] Note that with respect to NAL unit types, the ordering rules defined between them may enable a decoder to determine where the boundary between consecutive access units is located regardless of the NAL units of the removable packet type that have been removed while being transmitted or not transmitted to the decoder. Removable packet type NAL units may comprise, for example, SEI NAL units, or overlapping picture data NAL units, or other specific NAL unit types. That is, the boundary between access units remains stationary, and the ordering rules are still followed within each access unit, but are broken at each boundary between any two access units.
[0052] For completeness, Figure 18 illustrates that case 1 of Figure 7 allows packets belonging to different layers but the same instant (i-1) to be distributed within one access unit, for example. Case 2 of Figure 16 is similarly depicted with a circle around it.
[0053] The fact that the NAL units contained in each access unit are actually interleaved or not with respect to their association with the layers of the data stream can be decided at the discretion of the encoder. To facilitate the handling of the data stream, a syntax element may signal to the decoder the interleaving or deinterleaving of NAL units within an access unit that collects all NAL units of a particular time stamp, since the latter makes the NAL units easier to process. For example, whenever interleaving is signaled to be switched on, the decoder may use one or more coded picture buffers, as briefly illustrated with reference to FIG. 9.
[0054] FIG. 9 illustrates a decoder 700 that may be implemented as outlined above with reference to FIG. 2. Exemplarily, the multi-layer video data stream of FIG. 9, option 1 with a circle around it, is shown as input to the decoder 700. To more easily perform deinterleaving of NAL units belonging to different layers, but at a common instant, per access unit (AU), the decoder 700 uses two buffers 702 and 704 for each access unit AU, with a multiplexer 706 transferring NAL units of the access unit AU, e.g., belonging to the first layer, to buffer 702, and NAL units of the second layer, e.g., to buffer 704. A decoding unit 708 then performs the decoding. For example, in FIG. 9, NAL units belonging to the base / first layer are shown, e.g., not hatched, while NAL units of the dependent / second layer are shown hatched. If the interleaving signaling outlined above is present in the data stream, decoder 700 may respond to this interleaving signaling in the following manner: If interleaving signaling signals NAL unit interleaving to be switched on, i.e., NAL units of different layers are interleaved with each other within one access unit AU, and decoder 700 uses buffers 702 and 704 with multiplexer 706 distributing NAL units over these buffers as just outlined. Otherwise, however, encoder 700 simply uses one of buffers 702 and 704 for all NAL units comprised by any access unit, e.g., buffer 702.
[0055] To more easily understand the embodiment of FIG. 9, reference is made to FIG. 9 in conjunction with FIG. 10, which illustrates an encoder configured to generate a multi-layer video data stream as outlined above. The encoder of FIG. 10 is generally designated with reference numeral 720 and exemplarily encodes inbound pictures, here two layers, designated for ease of understanding as layer 12 forming the base layer and layer 1 forming the dependent layer. These form different perspectives, as outlined previously. The overall encoding instructions along with encoder 720 encoding pictures of layers 12 and 15 traverse these layers substantially together with their temporal (presentational) instructions, where encoder instructions 722 may deviate from the presentational time instructions of pictures 12 and 15 in units of groups of pictures. At each temporal instant, encoding instructions 722 pass through layers 12 and 15 and their dependencies, i.e., pictures from layer 12 to layer 15.
[0056] Encoder 720 encodes the pictures of layers 12 and 15 into data stream 40 in units of the aforementioned NAL units, each of which is associated with a respective portion of the picture in a spatial sense. Therefore, NAL units belonging to a particular picture spatially subdivide or partition each picture, and as already explained, intra-layer prediction renders portions of layer 15 pictures that depend on portions of temporally aligned pictures of layer 12 that are substantially co-located with respective portions of layer 15 pictures by "substantially" surrounding the disparity displacement. In the example of FIG. 10, encoder 720 is selected to exploit interleaving possibilities when forming access units that collect all NAL units belonging to a particular instant in time. Portions of data stream 40 outside the illustrated one in FIG. 10 are inbound to the decoder of FIG. 9. That is, in the example of FIG. 10, decoder 720 uses intra-layer parallelism when encoding layers 12 and 15. As far as instant t(i-1) is concerned, encoder 720 starts encoding the layer 1 picture as soon as NAL unit 1 of the layer 12 picture has been encoded. Each NAL unit that has been encoded is output by encoder 720 and is provided with an arrival timestamp that corresponds to the time that the NAL unit was output by encoder 720. After encoding the first NAL unit of the layer 12 picture at instant t(i-1), encoder 720 starts encoding the content of the layer 12 picture and outputs the second NAL unit of the layer 12 picture, which is provided with an arrival timestamp that follows the arrival timestamp of the first NAL unit of the layer 15 temporally aligned picture. That is, encoder 720 outputs the NAL units of the layer 12 and 15 pictures, which all belong to the same instant, in an interleaved manner, and in this interleaved manner, the NAL units of data stream 40 are actually transmitted. The circumstances in which the encoder 720 has chosen to utilize the interleaving possibilities may be indicated by the encoder 720 in the data stream 40 by respective methods of interleaving signaling 724 .The end-to-end delay between the decoder of FIG. 9 and the encoder of FIG. 10 can be reduced so that encoder 720 is able to output the first NAL unit of dependent layer 15 at an earlier instant t(i-1) compared to the deinterleaving scenario according to which the output of the first NAL unit of layer 15 would be delayed until the completion of encoding and output of all NAL units of the time-aligned base layer picture.
[0057] As already mentioned above, according to an alternative, in the case of deinterleaving, i.e., in the case of signaling 724 indicating a deinterleaving alternative, the definition of the access units may remain the same, i.e., an access unit AU may collect all NAL units belonging to a particular moment in time. In that case, signaling 724 simply indicates whether, within each access unit, NAL units belonging to different layers 12 and 15 are interleaved or not.
[0058] As explained above, depending on signaling 724, the decoder of Figure 9 uses either one buffer or two buffers. If interleaving is switched on, the decoder 700 distributes the NAL units over two buffers 702 and 704, such that, for example, layer 12 NAL units are buffered in buffer 702, while layer 15 NAL units are buffered in buffer 704. Buffers 702 and 704 are emptied for access units. This is true in both cases of signaling 724 indicating interleaving or deinterleaving.
[0059] It is preferable if encoder 720 sets the removal time within each NAL unit so that decoder unit 708 takes advantage of the possibility of decoding layers 12 and 15 from data stream 40 using inter-layer parallelism. End-to-end delay, however, is already reduced even if decoder 700 does not apply inter-layer parallelism.
[0060] As already explained above, NAL units can be of different NAL unit types. Each NAL unit can have a NAL unit type index indicating the type of each NAL unit outside of the possible types and within each access unit. The NAL unit types of each access unit can follow the ordering rules within the NAL unit types, while between two consecutive access units, the ordering rules are broken so that the decoder 700 can identify the access unit boundary by examining this rule. For more detailed information, please refer to the H.264 standard.
[0061] With reference to Figures 9 and 10, a decoding unit (DU) can be identified as consecutive NAL unit operations within one access unit that belong to the same layer. The NAL units indicated by "3" and "4" in Figure 10 in access unit AU(i-1) form, for example, one DU. All other decoding units in access unit AU(i-1) comprise only one NAL unit. Together, access unit AU(i-1) in Figure 19 exemplarily comprises six decoding units DU that are interleaved within access unit AU(i-1), i.e., they are composed of NAL unit operations of one layer, with one layer alternating between layer 1 and layer 0.
[0062] 7-10, mechanisms for enabling and controlling CPBs are provided for operation in a multi-layer video codec that meets the required ultra-low delays possible with current single-layer video codecs such as HEVC. Based on the bitstream instructions illustrated in the just-mentioned figures, the following describes a video decoder operating on an incoming bitstream buffer, i.e., a picture buffer, coded at the decoding unit level, where, in addition, the video decoder operates multiple CPBs at the DU level. In particular, also in a manner applicable to HEVC extensions, an operating mode is described in which additional timing information is provided for operation of the multi-layer codec in a low-delay manner. This timing provides a mechanism for controlling CPBs for interleaved coding of different multiple layers in the stream.
[0063] In the embodiments described below, case 2 of Figures 7 and 10 is not necessary, or in other words, does not need to be implemented: the access unit may maintain its function as a container that collects all payload packets (VCL NAL units) that carry information about pictures belonging to a particular time stamp or instant—regardless of layer. Nevertheless, the embodiments described below achieve compatibility with different types of encoders or encoders that prefer different strategies when decoding inbound multi-layer video data streams.
[0064] That is, the video encoders and decoders described below are still scalable, multiview, or 3D video encoders and decoders. Following the above description, the term layer is used collectively for scalable video coding layers as well as for view- and / or depth maps of multiview coded video streams.
[0065] The DU-based decoding mode, i.e., DU CPB removal in a continuous manner, can still be used by a single-layer (basic specification) ultra-low delay decoder, according to some of the embodiments outlined below. A multi-layer ultra-low delay decoder would use the interleaved DU-based mode of decoding to achieve low-delay operation with multiple layers, as described for Case 1 in Figures 7 and 10 and subsequent figures. On the other hand, a multi-layer decoder that does not decode interleaved DUs may fall back to an AU-based decoding process, or DU-based decoding in a deinterleaved manner, according to various embodiments, which would provide low-delay operation between the interleaved and AU-based approaches. The resulting three operation modes are as follows: AU-based decoding: All DUs of an AU are removed from the CPB at the same time. Decoding based on consecutive DUs: Each DU in a multi-layer AU is attributed to the CPB removal time following the DU removal in the consecutive order of layers. That is, all DUs in layer m are removed from the CPB before the DUs in layer (m+1) are removed from the CPB. Decoding based on interleaved DUs: Each DU of a multi-layer AU is attributed to the CPB removal time following the DU removal in an interleaved order across multiple layers, i.e., the DUs of layer m can be removed from the CPB later than the DUs of layer (m+1) are removed from the CPB.
[0066] The additional timing information for interleaved operation enables a system layer device to determine the arrival time at which a DU must arrive at the CPB when the transmitter sends multi-layer data in an interleaved manner, regardless of the decoder operating mode, which is necessary for proper operation of the decoder to prevent buffer overflow and underflow. How a system layer device (e.g., an MPEG-2 TS receiver) can determine the time at which data must arrive at the decoder CBP is exemplarily shown at the end of the following section, CPB Operation Alone.
[0067] The following table in Figure 11 gives an exemplary embodiment for signaling the presence in the bitstream of DU timing information for operation in interleaved mode.
[0068] Another embodiment would be a suggestion to correspond to an interleaved operating mode in which an indication of DU timing information is provided so that a device that cannot operate in interleaved DU mode can operate in AU mode and override DU timing.
[0069] In addition, another operating mode featuring layer-by-layer, DU-based CPB removal, i.e., DU CPB removal in a deinterleaved manner across layers, allows the same low-latency CPB operation based on DUs as in the interleaved mode for the base layer, but removes DUs from layer (m+1) only after completing removal of DUs in layer m. Therefore, non-base layer DUs may remain in the CPB for a longer period than when removed in the interleaved CPB operating mode. The tables in Figures 12a and 12b provide an exemplary embodiment for signaling additional timing information in an additional SEI based on either the AU level or the DU level. Other possibilities include suggesting that timing provided by other means guides CPB removal from CPBs interleaved across layers.
[0070] A further aspect is the possibility of applying the decoder operation modes described for the following two cases:
[0071] 1. A single CPB is used to accommodate data from all layers. NAL units of different layers may be interspersed within an access unit. This mode of operation is hereafter referred to as single CPB operation. 2. One CPB per layer. The NAL units of each layer are located in consecutive positions. This mode of operation is subsequently referred to as multiple CPB operation.
[0072] -Standalone CPB operation In Figure 13, the arrival of decoding unit (1) at access unit (2) is shown for a bitstream ordered by layers. The numbers in the boxes indicate the layer IDs. As shown, all DUs of layer 0 arrive first, followed by DUs of layer 1, and then layer 2. Three layers are shown in the example, however, more layers may follow.
[0073] In Figure 14, the arrival of decoding unit (1) at access unit (2) is shown for the interleaved bitstream according to Figures 7 and 8, Case 1. The numbers in the boxes indicate layer IDs. As shown, DUs from different layers can be mixed within an access unit.
[0074] A CPB removal time is associated with each decoding unit, which is the start time of the decoding process. This decoding time can be lower than the final arrival time of the decoding unit, exemplarily shown as (3) for the first decoding unit. The final arrival time of the first decoding unit in the second layer, labeled (4), can be lowered by using interleaved bitstream ordering as shown in Figure 14.
[0075] One embodiment is a video encoder that generates a decoder hint in a bitstream that indicates the lowest possible CPB removal (and therefore decoding time) in a bitstream that uses high-level syntax elements for an interleaved bitstream.
[0076] A decoder utilizing the decoder hints described for lower arrival times will remove decoding units from the CPB immediately or soon after their arrival, thus allowing parts of a picture to be fully coded (through all layers) sooner and therefore displayed sooner than for a deinterleaved bitstream.
[0077] A lower-cost implementation of such a decoder can be achieved by constraining the signal timing in the following way for any DUn that precedes a DUm in bitstream order: The CPB removal time for DUn will be lower than or equal to the CPB removal time of DUm. When arriving, packets are stored in consecutive memory addresses (typically a circular buffer) in the CPB; this constraint avoids fragmenting free memory in the CPB. Packets are removed in the same order as they are received. Instead of keeping a list of used and free memory blocks, the encoder can be implemented to keep only the beginning and ending addresses of used memory blocks. This ensures that newly arriving DUs do not need to be split into several memory locations because the used and free memory are consecutive blocks.
[0078] In the following, we describe an embodiment based on the actual current HRD definition as used by the HEVC extension, where timing information for interleaving is provided through an additional DU-level SEI message as mentioned above. The described embodiment allows DUs to be transmitted in order to be interleaved across layers, either consecutively or AU-wise, to be removed DU-wise from the CPB in an interleaved manner.
[0079] In a single CPB solution, the CPB removal times in Annex [1]C should be extended as follows (underlined):
[0080] Multiple tests may be required to verify the conformance of a bitstream, referred to as the bitstream under test. For each test, the following steps are applied in the listed order: For each access unit in BitstreamToDecode, starting with access unit 0, a buffer time SEI message (present in BitstreamToDecode or available through external means not specified in this specification) associated with the access unit and applied to the TargetOp is selected; a picture timing SEI message (present in BitstreamToDecode or available through external means not specified in this specification) associated with the access unit and applied to the TargetOp is selected; and when SubPicHrdFlag is 1 and sub_pic_cpb_params_in_pic_timing_sei_flag is 0, a decoding unit information SEI message (present in BitstreamToDecode or available through external means not specified in this specification) associated with a decoding unit in the access unit and applied to the TargetOp is selected; and sub_pic_interleaved_hrd_params_present_flag is 1, a Decoding Unit Interleaving Information SEI message (present in BitstreamToDecode or available through external means not specified in this specification) associated with the decoding unit in the access unit and applied to the TargetOp is selected. When sub_pic_interleaved_hrd_params_present_flag is 1 in the selected syntax structure, the CPB will operate either at the AU level (when the variable SubPicInterleavedHrdFlag is set to 0) or at the interleaved DU level (when the variable SubPicInterleavedHrdFlag is set to 1). variable SubPicInterleavedHrdPreferredFlag is set to 0 either when specified by external means or when not specified by external means. If the value of the variable SubPicInterleavedHrdFlag was not set in step 9 in this subclause above, it is derived as follows. SubPicInterleavedHrdFlag = SubPicHrdPreferredFlag && SubPicInterleavedHrdPreferredFlag && sub_pic_interleaved_hrd_params_present_flag
[0081] SubPicHrdFlag, and SubPicInterleavedHrdFlag If 0, HRD operates at the access unit level and each decoding unit is an access unit, otherwise HRD operates at the sub-picture level and each decoding unit is a subset of an access unit.
[0082] For each bitstream conformance test, CPB operation is specified in subclause C.2, instantaneous decoder operation is specified in clauses 2-10, DPB operation is specified in subclause C.3, and output cropping is specified in subclause C.3.3 and subclause C.5.2.2.
[0083] HSS and HRD information for the number of enumerated delivery schedules, and their associated bit rates and buffer sizes, are specified in subclauses E.1.2 and E.2.2. The HRD is initialized as specified by the buffer time SEI message specified in subclauses D.2.2 and D.3.2. The removal times of decoding units from the CPB and the output timing of coded pictures from the DPB are specified in the decoding unit information SEI message (specified in subclauses D.2.21 and D.3.21). or in a Decoding Unit Interleaving Information SEI message (specified in subclauses D.2.XX and D.3.XX), It is specified using information in the Picture Timing SEI message (specified in subclauses D.2.3 and D.3.3). All timing information for a particular decoding unit will arrive before the CPB removal time of the decoding unit.
[0084] When SubPicHrdFlag is 1, the following applies: The variable duCpbRemovalDelayInc is derived as follows: - If SubPicInterleavedHrdFlag is 1, duCpbRemovalDelayInc is set equal to the value of du_spt_cpb_interleaved_removal_delay_increment in the Decoding Unit Interleaving Information SEI message, and subordinate clause C1 selected as specified in associated with decoding unit m . - Otherwise , sub_pic_cpb_params_in_pic_timing_sei_flag is If it is 0 and sub_pic_interleaved_hrd_params_present_flag is 0, duCpbRemovalDelayInc is set equal to the value of du_spt_cpb_removal_delay_increment in the decoding unit information SEI message, selected as specified in subclause C1, and associated with decoding unit m. - Otherwise, if sub_pic_cpb_params_in_pic_timing_sei_flag is 0 and sub_pic_interleaved_hrd_params_present_flag is 1, then duCpbRemovalDelayInc is set equal to the value of du_spt_cpb_removal_delay_increment in the Decoding Unit Information SEI message, and duCpbRemovalDelayIncInterleaved is set equal to the value of du_spt_cpb_interleaved_removal_delay_increment in the Decoding Unit Interleaving Information SEI message, selected as specified in subclause C1 and associated with decoding unit m. - Otherwise, du_common_cpb_removal_delay_flag is If it is 0 and sub_pic_interleaved_hrd_params_present_flag is 0, duCpbRemovalDelayInc is set equal to the value of du_cpb_removal_delay_increment_minus1[i]+1 for decoding unit m in the picture timing SEI message, selected as specified in subclause C.1, and associated with access unit n, where the first num_nalus_in_du_minus1[0]+1, where i has the value 0, consecutive NAL units in the access unit include decoding unit m, the next num_nalus_in_du_minus1[1]+1 NAL units in the same access unit where i has the value 1, the next num_nalus_in_du_minus1[2]+1 NAL units in the same access unit where i has the value 2, etc. - Otherwise, if du_common_cpb_removal_delay_flag is 0 and sub_pic_interleaved_hrd_params_present_flag is 1, then duCpbRemovalDelayInc is set equal to the value of du_cpb_removal_delay_increment_minus1[i]+1 in the Picture Timing SEI message for decoding unit m, and duCpbRemovalDelayIncInterleaved is set equal to the value of du_spt_cpb_interleaved_removal_delay_increment in the Decoding Unit Interleaving Information SEI message, selected as specified in subclause C.1 and associated with access unit n. where the first num_nalus_in_du_minus1[0]+1 consecutive NAL units in an access unit with i value 0 include decoding unit m, the next num_nalus_in_du_minus1[1]+1 NAL units in the same access unit with i value 1, the next num_nalus_in_du_minus1[2]+1 NAL units in the same access unit with i value 2, etc. - Otherwise, duCpbRemovalDelayInc is set equal to the value of du_common_cpb_removal_delay_increment_minus1+1 in the picture timing SEI message, selected as specified in subclause C.1 and associated with access unit n. The nominal removal time of a decoding unit m from the CPB is specified as follows, where AuNominalRemovalTime[n] is the nominal removal time of access unit n: If decoding unit m is the last decoding unit in access unit n, then the nominal removal time DuNominalRemovalTime[m] of decoding unit m is set equal to AuNominalRemovalTime[n]. Otherwise (decoding unit m is not the last decoding unit in access unit n), the nominal removal time DuNominalRemovalTime[m] of decoding unit m is derived as follows: if( sub_pic_cpb_params_in_pic_timing_sei_flag && !SubPicInterleavedHrdFlag) DuNominalRemovalTime[m] = DuNominalRemovalTime[m+1] - ClockSubTick*duCpbRemovalDelayInc (C13) else DuNominalRemovalTime[m] = AuNominalRemovalTime(n) - ClockSubTick*duCpbRemovalDelayInc
[0085] The DU operation mode used, SubPicInterleavedHrdFlag, determines either interleaved or non-interleaved operation mode, and DUNominalRemovalTime[m] is the removal time of the DU for the selected operation mode. In addition, the earliest arrival time of DUs will be different from what is currently defined if sub_pic_interleaved_hrd_params_present_flag is 1, regardless of the operation mode. The earliest arrival time is therefore derived as follows: if(!SubPicInterleavedHrdFlag&& sub_pic_interleaved_hrd_params_present_flag) DuNominalRemovalTimeNonInterleaved[ m ] = AuNominalRemovalTime(n) - ClockSubTick*duCpbRemovalDelayIncInterleaved if( !subPicParamsFlag ) tmpNominalRemovalTime = AuNominalRemovalTime[m] (C6) else if(!sub_pic_interleaved_hrd_params_present_flag || SubPicInterleavedHrdFlag) tmpNominalRemovalTime = DuNominalRemovalTime[m] else tmpNominalRemovalTime = DuNominalRemovalTimeNonInterleaved[m] "
[0086] With respect to the above embodiment, it is worth noting that the operation of the CPB accounts for the arrival time of data packets at the CPB as well as the removal time of the data packets, which is explicitly signaled. Such arrival times affect the behavior of intermediate devices that constitute buffers along the data packet transmission chain, for example, elementary stream buffers in an MPEG-2 transport stream receiver, which act as the CPB of the decoder. The HRD model based on the above embodiment derives the first arrival time based on the variable tmpNominalRemovalTime, thereby considering either the removal time for DUs in the case of interleaved DU operation (as if data were removed from the CPB in an interleaved manner) or the equivalent removal time "DuNominalRemovalTimeNonInterleaved" for continuous DU operation mode, for the exact first arrival time of a data packet at the CPB (see C-6) (as if data were removed from the CPB in an interleaved manner).
[0087] A further embodiment is layer-wise reordering of DUs for AU-based decoding operations. When a single CPB operation is used and data is received in an interleaved manner, it may be desirable for the decoder to operate based on AUs. In such a case, data read from the CPB corresponding to several layers will be interleaved and immediately transmitted to the decoder. When AU-based decoding operations are performed, AUs are reordered / rearranged in such a way that all DUs from layer m precede DUs from layer m+1 before being transmitted for decoding, because a reference layer is always decoded before the enhancement layer that references it.
[0088] Multi-CPB operation Instead, the decoder is described with one coded picture buffer for each DU in a layer.
[0089] Figure 15 shows the allocation of DUs to different CPBs. For each layer (numbered in the box), when its CPB is run, the DUs are stored in different memory locations for each CPB. Exemplarily, the arrival times of interleaved bitstreams are shown. The allocation works in the same way as for non-interleaved bitstreams, based on layer identifiers.
[0090] Figure 16 shows memory usage in different CPBs. DUs of the same layer are stored in consecutive memory locations.
[0091] Multi-layer decoders can take advantage of this memory allocation because DUs belonging to the same layer can be accessed at consecutive memory addresses. DUs arrive in the decoding order of each layer. Removal of DUs from different layers cannot create any "holes" in the used CPB memory area. Used memory blocks always cover consecutive blocks in each CPB. The multiple CPB concept also has the advantage of layer-wise split bitstreams at the transport layer. If different layers are transmitted using different channels, multiplexing of DUs into a single bitstream can be avoided. Therefore, multi-layer video decoders do not need to implement this additional step, and implementation costs can be reduced.
[0092] When multi-CPB operation is used, in addition to the timing described for the single CPB case also applies:
[0093] A further aspect is the reordering of DUs from multiple CPBs when they share the same CPB removal time (DuNominalRemovalTime[m]). In both interleaved and non-interleaved modes of operation for DU removal, it may occur that DUs from different layers, and therefore different CPBs, share the same CPB removal time. In such cases, the DUs are ordered by increasing LayerId numbers before being sent to the decoder.
[0094] The above-described embodiment also describes a mechanism for synchronizing with multiple CPBs below. In the present text [1], the reference time or anchor time is described as the first arrival time of the first decoding unit entering a (unique) CPB. In the case of multiple CPBs, there is one master CPB and multiple slave CPBs, leading to dependencies between the multiple CPBs. A mechanism for the master CPB to synchronize with the slave CPBs is also described. This mechanism is advantageous because CPBs receiving DUs remove their DUs at the unique time, i.e., using the same time reference. More specifically, the first DU that initializes the HRD synchronizes with the other CPBs, and the anchor time is set equal to the first arrival time of the DU for the extended CPB. In a specific embodiment, the master CPB is the CPB for base layer DUs, while if a random access point for the enhancement layer is enabled and the HRD is initialized, the master CPB can correspond to the CPB that receives enhancement layer data.
[0095] Therefore, in accordance with the considerations outlined above following Figures 7-10, the comparative embodiments of these figures are modified in the manner outlined below with respect to the following figures. The encoder of Figure 17 operates similarly to the one discussed above with respect to Figure 10. Signaling 724, however, is optional. Accordingly, the scope of the above description also applies to the following embodiment, and similar statements will apply to the decoder embodiment described subsequently.
[0096] In particular, the encoder 720 of FIG. 17 exemplarily encodes video content including video from layers 12 and 15 into a multi-layer video data stream 40, with the video content encoded therein in units of subportions of pictures of the video content using intra-layer prediction for each of layers 12 and 15. In the example, the subportions are denoted 1-4 for layer 0 and 1-3 for layer 1. Each subportion is encoded into one or more payload packets of a sequence of packets of the video data stream 40, each associated with one of multiple layers, and the sequence of packets is divided into a sequence of access units (AUs) so that an access unit collects payload packets associated with a common instant. Two AUs are exemplarily shown, one at instant i-1 and the other at instant i. The access units (AUs) are subdivided into decoding units (DUs) so that each access unit is subdivided into two or more decoding units, with each decoding unit solely comprising payload packets associated with one of multiple layers. Decoding units with packets associated with different layers are interleaved with each other. In simple terms, the encoder 720 controls the interleaving of decoding units within an access unit to reduce—or keep as low as possible—end-to-end delay by traversing and encoding common instants in the traversal order, first layers, then sub-parts. So far, the encoder's modes of operation have already been provided above with respect to FIG. 10.
[0097] However, the encoder of Figure 17 provides each access unit AU with two times of time control information: a first time control information 800 signals the respective decoder buffer search times for the access unit as a whole, and a second time control information 802 signals the decoder buffer search times for each of the decoding units DU of the access unit AU, which correspond to their sequential order in the multi-layer video data stream.
[0098] Figure 17 Illustrated 8, the encoder 720 encodes the second timing information 802 into the respective The demodulation timing control packet is associated with No. Unit DU precedes And each Timing Control Packets The preceding No. Unit DU of Second Decoder Search Buffer Time of Show vinegar , some Timing Control Packets spread it to. FIG. 12c shows an example of such a timing control packet. As can be seen here, Timing Control Packets teeth Associated Ta Forms the beginning of a decoding unit can be ,and Each Associated with DU Ta index, i.e., decoding_unit_idx, and Each DU About decoder search buffer time, i.e., Indicates the search time or DPB removal time in a given time unit (increment) du_spt_cpb_removal_delay_interleaved_increment Therefore, , the second timing control information 802 is the current Layers of moments but encoding It has been It can be output by the encoder 720 in between. Nota For teeth , the encoder 720 calculates the spatial complexity in the pictures 12 and 15 of the various layers during encoding. is changing React.
[0099] The encoder 720 may evaluate the first timing control information 800 at the decoding buffer search time for each access unit AU, i.e., prior to encoding the layer at the current moment and location, at the beginning of each AU, or - if possible according to the criteria - at the end of the AU.
[0100] In addition to or instead of providing the timing control information 800, the encoder 720 provides, to each access unit, third timing control information signaling a third decoder buffer search time for each decoding unit of the access unit, as shown in Figure 18. As a result, according to the third decoder buffer search times for the decoding units DU of each access unit, the decoding units DU in each access unit AU are ordered according to the layer order defined among the multiple layers, so that a non-decoded unit comprising packets associated with a first layer follows any decoding unit in each access unit comprising packets associated with a second layer that follows the first layer according to the layer order. That is, according to the buffer search times of the third timing control information 804, the DUs shown in Figure 18 are reordered at the decoder side so that the DUs of portions 1, 2, 3, and 4 of picture 12 precede the DUs of portions 1, 2, and 3 of picture 15. The encoder 720 may estimate the decoder buffer search time according to the first timing control information 802 at the beginning of each AU, the third timing control information 80 for DUs prior to encoding the layer at the current moment and place. This possibility is exemplarily depicted in FIG. 8 and in FIG. 12a. FIG. 12a shows that ldu_spt_cpb_removal_delay_interleaved_increment_minus1 is transmitted for each DU and for each layer. Although the number of decoding units per layer could be equally limited for all layers, i.e., one num_layer_decoding_units_minus1 would be used as illustrated in FIG. 12a, an alternative to FIG. 12a would be that the number of decoding units per layer could be set individually for each layer. In the latter case, the syntax element num_layer_decoding_units_minus1 may be read for each layer, in which case the reading would be moved from the location shown in Figure 12a to, for example, between the two next for loops in Figure 12a.As a result, num_layer_decoding_units_minus1 will be read for each layer in the next for loop using the counter variable j. If permitted to follow the criteria, the encoder 720 may instead set the timing control information at the end of the AU. Even instead, the encoder 720 may set the third timing control information at the beginning of each DU, just like the second timing control information. This is shown in Figure 12b. Figure 12b shows an example for a timing control packet that is set at the beginning of each DU (in their interleaved state). As can be seen from the fact that the timing control packet carrying the timing control information 804 for a particular DU can be set at the beginning of the associated decoding unit and indicates one index, namely, layer_decoding_unit_idx, associated with each DU, it is layer-specific, i.e., all DUs belonging to the same layer are attributed to the DU index of the same layer. Furthermore, ldu_spt_cpb_removal_delay_interleaved_increment, which indicates the decoder search buffer time for each DU, i.e., the search time or DPB removal time in a given temporal unit (increment), is signaled in such a packet. According to these timings, the DUs are re-partitioned to follow the layer order: first, the DUs of layer 0 are removed from the DPB, then the DUs of layer 1, etc. Accordingly, timing control information 808 can be output by the encoder 720 during the encoding of the current layer.
[0101] As explained, information 802 and 804 may exist simultaneously in the data stream, as illustrated in FIG.
[0102] Finally, as illustrated in Figure 20, a decoding unit interleaving flag 808 can be inserted into the data stream by the encoder 720 for either timing control information 802 or timing control information 808 to be transmitted in addition to timing control information 800 that operates like 804. That is, if the encoder 720 decides to interleave DUs of different layers as depicted in Figure 20 (and Figures 17-19), then the decoding unit interleaving flag 808 is set to indicate that information 808 is equal to information 802, and the above description of Figure 17 applies with respect to the remainder of the encoder functionality of Figure 20. However, if the encoder does not interleave packets of different layers within an access unit as depicted in FIG. 13, then the decoding unit interleaving flag 808 is set by the encoder 720 to indicate that information 808 is equal to information 804. With the difference from the depiction of FIG. 18 regarding the generation of timing control information 804, the encoder therefore does not need to evaluate timing control information 804 in addition to information 802, but can determine the buffer retrieval time of timing control information 804 on the fly by reacting to variations in layer-specific coding complications within the layer sequence on the fly while encoding the access unit in a manner similar to the procedure for generating timing control information 802.
[0103] Figures 21, 22, 23, and 24 show the data streams of Figures 17, 18, 19, and 20, respectively, as they enter decoder 700. If decoder 700 is configured as the one described with respect to Figure 9 above, then decoder 700 can decode the data stream in the same manner as described above with respect to Figure 9 using timing control information 802. That is, both the encoder and decoder contribute to minimal delay. In the case of Figure 24, the decoder receives the stream of Figure 20, which is obviously only possible in the case of DU interleaving, which is used and indicated by flag 806.
[0104] For whatever reason, the decoder, however, decodes the multi-layer video data stream by emptying the decoder buffer for buffering the multi-layer data stream on an access unit basis using the first timing control information 800 in the cases of Figures 21, 23, and 24 and regardless of the second timing control information 802. For example, the decoder may not be able to perform parallel processing. The decoder may not have, for example, more than one buffer. Because the decoder buffer operates at the complete AUs rather than at the DU level, delays are increased on both the encoder and decoder sides compared to using an interleaved layer order of DUs according to the timing control information 802.
[0105] As already discussed above, the decoder of Figures 21-24 does not need to have two buffers. One buffer such as 702 is sufficient, especially if timing control information 802 is not used, rather than any alternative in the form of timing control information 800 and 804. On the other hand, the decoder buffer may consist of a partial buffer for each layer. This is useful when timing control information is used, because the decoder can buffer decoding units with packets associated with each layer in partial buffers for each layer. 705 illustrates the possibility of having more than two buffers. The decoder may empty decoding units from different partial buffers to decoder entities of different coders. Instead, the decoder uses fewer partial buffers compared to the number of layers, i.e., each partial buffer for a subset of layers by transferring the Ds of a particular layer to the partial buffer associated with the set of layers to which each DU belongs. One partial buffer such as 702 may be synchronized with other partial buffers such as 704 and 705.
[0106] 22 and 23, the decoder decodes the multi-layer video data stream by emptying the decoder buffer, controlled via timing control information 804, i.e., by deinterleaving according to the layer order, thereby removing the decoding units of the access unit. By this measure, the decoder—guided by timing control information 804—effectively recombines the decoding units associated with the same layer and belonging to the access unit, and reorders them according to a specific rule, such as DUs of layer n before DUs of layer n+1.
[0107] As illustrated in FIG. 24, a decoding unit that interleaves flags 806 in a data stream can operate as if timing control information 808 were either timing control information 802 or 804. In this case, a decoder receiving the data stream can be configured to respond to the decoding unit that interleaves flags 806 to empty the decoder's buffer for buffering the multi-layer data stream on an access unit basis using the first timing control information 800 and regardless of information 806 if the information 806 is the second timing control information according to 802, and to empty the decoder's buffer for buffering the multi-layer data stream on a decoding unit basis using information 806 if the information is timing control information according to 804. That is, in this case, the end-to-end delay between DUs that would otherwise be achieved by using 802 and DU operations ordered using timing control information 808 is not interleaved, and the maximum delay achievable by timing control information 800 will result.
[0108] Whenever timing control information 800 is used as an alternative, i.e., the encoder chooses to empty its buffer on an access unit basis, the decoder can remove the decoded units of the access unit from buffer 702—or even fill buffer 702 with DUs—in a deinterleaving manner. This results in AUs with DUs being ordered according to layer order. That is, the decoder can recombine decoded units associated with the same layer and belonging to the access unit and reorder them according to a specific rule, such as DUs of layer n before DUs of layer n+1, before removing all AUs from the buffer for decoding. This deinterleaving is not necessary in the case of a decoding unit interleaving flag 806 of FIG. 24, which indicates that deinterleaving has already been used, and timing control information 808 operates like timing control information 804.
[0109] Although not specifically discussed above, the second timing control information 802 may be defined as an offset for the first timing control information 800 .
[0110] 21-24 operates as an intermediate network device configured to forward a multi-layer video data stream for a coded picture buffer of a decoder. According to an embodiment, the intermediate network device is configured to receive information identifying a decoder capable of handling second timing control information 802, and, if the decoder can handle the second timing control information 802, derive the earliest arrival time for scheduling transmission from the timing control information 802 and 800 according to a first calculation rule, i.e., DuNominalRemovalTime; and, if the decoder cannot handle the second timing control information 802, derive the earliest arrival time for scheduling transmission from the timing control information 802 and 800 according to a second calculation rule, i.e., DuNominalRemovalTimeNonInterleaved. To explain the just outlined problem in more detail, an intermediate network device, also denoted by reference numeral 706, is shown disposed between the inbound multi-layer video data stream and the output leading to a decoding buffer, generally denoted by reference numeral 702, but reference is made to FIG. 25 outlined above. The decoding buffer may be composed of several partial buffers. As shown in FIG. 25, the intermediate network device 706 internally comprises a buffer 900 for buffering inbound DUs and therefore forwarding them to the decoding buffer 702. The above-described embodiment has been concerned with the decoding buffer removal time of inbound DUs, i.e., the time when these DUs need to be transferred from the decoding buffer 702 to the decoding unit 708 of the decoder. However, the storage capacity of the decoding buffer 702 is guaranteed as much as possible in addition to the removal time, i.e., the amount of time for DUs buffered in the buffer 702 to be removed, for which the earliest arrival time will also be managed.This is the purpose of the "earliest arrival time" mentioned earlier, and according to the embodiment outlined here, the intermediate network device 706 is configured to calculate these earliest arrival times based on the acquired timing control information according to different calculation rules, with the calculation rule being selected according to information about the decoder's ability to operate on the inbound DUs in their interleaved format, i.e., depending on whether the decoder can operate in the same interleaved manner or in a different manner. In principle, the intermediate network device 706 can determine the earliest arrival time based on the DU removal time of the timing control information 802 for a decoder that can decode the inbound data stream using the DU interleaving concept, by providing a constant time offset between the earliest arrival time and the removal time for each DU. The intermediate network device 706 similarly provides a fixed time offset between AU removal times as indicated by the timing control information 800 to derive the earliest arrival time of the access unit in case of a decoder that prefers a treatment with respect to the inbound data stream access unit, i.e., selects the first timing control information 800. Instead of using a fixed time offset, the intermediate network device 706 may also consider the size of the individual DUs and AUs.
[0111] The problem of Figure 25 will also be used as an opportunity to demonstrate possible improvements to the embodiments described above. In particular, the embodiments discussed above treated timing control information 800, 802, and 804 as signaling "decoder buffer removal times" for the decoding unit and the access unit, directly signaling "removal times," i.e., the times when each DU and AU needs to be transferred from buffer 702 to decoding unit 708, respectively. However, as is clear from the discussion of Figure 25, the arrival times and removal times are interrelated through the size of decoder buffer 702, as their guaranteed minimum size, and, on the other hand, the size of individual DUs in the case of timing control information 802 and 804, and the size of an access unit in the case of timing control information 800, respectively. Accordingly, all of the above-described embodiments may be interpreted such that the "decoder buffer removal time" signaled by "timing control information" 800, 802, and 804, respectively, includes explicit signaling by way of both the earliest arrival time or buffer removal time. All of the above discussions directly translate from the explanations discussed above using explicit signaling of the buffer removal time as the decoder buffer search time in the alternative embodiment in which the earliest arrival time is used as the decoder buffer removal time: interleaved transmission DUs may be repartitioned according to timing control information 804. The only difference: the repartitioning or deinterleaving may be performed upstream, i.e., before the buffer 702, i.e., between the buffer 702 and the decoding unit 708, rather than downstream thereof.In the case of an intermediate network device 706 that calculates earliest arrival times from inbound timing control information, the intermediate network device 706 will use these earliest arrival times to instruct a network entity located upstream with respect to the intermediate network device 706, such as the encoder itself or the same intermediate network entity, to follow these earliest arrival times in supplying the buffer 900, and in the alternative case to derive a buffer removal time from the inbound timing control information. Thus, with explicit signaling of the earliest arrival time, the intermediate network device 706 will activate access units from the buffer 900 according to the removal time of the DUs, or in the alternative case, the derived removal time.
[0112] Summarizing the very outline of the embodiment alternatives outlined above, this can be done by using the timing control information directly or indirectly to empty the decoder buffer: if the timing control information is implemented as direct signaling of decoder buffer removal times, then emptying the buffer can be done to be directly scheduled according to these decoder buffer removal times, and if the timing control information is implemented using decoder buffer arrival times, then a recalculation can be done to deduce the decoder buffer removal times from these decoder buffer arrival times according to the removal of DUs or AUs.
[0113] As noted above in common with the various embodiments and drawings illustrating "interleaved packet" transmission, it is presented that "interleaving" does not necessarily involve combining packets belonging to DUs of different layers on a common channel. Rather, transmission can occur fully parallel on separate channels (logically or physically separate channels): packets of different layers, thus forming different DUs, are output by the encoder in parallel, with output time interleaved as discussed above, and in addition to the DUs, the time control information described above is transmitted to the decoder. Among this timing control information, timing control information 800 indicates when DUs forming a complete AU need to be transferred from the decoder buffer to the decoder, timing control information 802 indicates for each individual DU when each DU needs to be transferred from the decoder buffer to the decoder, these removal times corresponding to the output time order of the DUs at the encoder, and timing control information 804 indicates for each individual DU when each DU needs to be transferred from the decoder buffer to the decoder, these search times deviating from the output time order of the DUs at the encoder and leading to repartition: instead of being transferred from the decoder buffer to the decoder, in the output interleaving order, DUs of layer i are transferred before DUs of layer i+1 for all layers. As described above, DUs can be distributed to separate buffer portions according to the layer relationships.
[0114] The multi-layer video data streams of Figures 19 to 23 may be in accordance with AVC or HEVC or any extension thereof, but this does not exclude other possibilities.
[0115] Although some aspects have been described in the context of an apparatus, it will be apparent that these aspects also represent a description of a corresponding method, where a block or apparatus corresponds to a method step or a function of a method step. Similarly, when an aspect is described in the context of a method step, it also represents a description of a corresponding block or item or feature of a corresponding apparatus. Some or all of the method steps may be performed by (or using) a hardware apparatus, such as, for example, a microprocessor, a programmable computer, or an electronic circuit. In some embodiments, one or more of the most significant method steps may be performed by such an apparatus.
[0116] The original encoded video signal may be stored on a digital storage medium or transmitted over a transmission medium such as a wireless transmission medium or a wired transmission medium such as the Internet.
[0117] Depending on specific implementation requirements, embodiments of the present invention may be implemented in hardware or software. Implementations may be performed using, for example, digital storage media such as floppy disks, DVDs, Blu-rays, CDs, ROMs, PROMs, EPROMs, EEPROMs, or FLASH memory having electronically readable control signals stored thereon, which cooperate (or are capable of cooperating) with a programmable computer system on which the respective methods are executed. Thus, the digital storage media may be computer-readable.
[0118] Some embodiments of the present invention comprise a data carrier having electronically readable control signals, which can cooperate with a programmable computer system to perform one of the methods described herein.
[0119] In general, embodiments of the present invention may be implemented as a computer program product with program code that, when run on a computer, causes the computer program product to perform one of the methods. The program code may, for example, be stored on a device-readable carrier.
[0120] Other embodiments comprise the computer program for performing one of the methods described herein, stored on a machine readable carrier.
[0121] In other words, an embodiment of the inventive method is, therefore, a computer program having a program code for performing one of the methods described herein, when the computer program runs on a computer.
[0122] A further embodiment of the inventive method is therefore a data carrier (or digital storage medium, or computer readable medium) comprising stored thereon a computer program for performing one of the methods described herein. The data carrier, digital storage medium or stored medium is typically tangible and / or non-transitory.
[0123] A further embodiment of the inventive method is, therefore, a data stream or sequence of signals representing a computer program for performing one of the methods described herein, which data stream or sequence of signals can for example be adapted to be transmitted via a data communications connection, for example via the Internet.
[0124] A further embodiment comprises a processing means, for example a computer, or a programmable logic device configured to or adapted to perform one of the methods described herein.
[0125] A further embodiment comprises a computer having installed thereon a computer program for performing one of the methods described herein.
[0126] Further embodiments according to the invention comprise an apparatus or system configured to transmit (e.g., electronically or optically) a computer program for performing one of the methods described herein to a receiver. The receiver may be, for example, a computer, a mobile device, a memory device, etc. The apparatus or system may comprise, for example, a file server for transmitting the computer program to the receiver.
[0127] In some embodiments, a programmable logic device (e.g., a field programmable gate array) may be used to perform some or all of the functionality of the methods described herein. In some embodiments, a field programmable gate array may cooperate with a microprocessor to perform one of the methods described herein. In general, the methods are preferably performed by any hardware device.
[0128] The above-described embodiments are merely illustrative of the principles of the present invention. It is understood that modifications and variations of the arrangements and details described herein will be apparent to others skilled in the art. It is therefore intended to be limited only by the scope of the appended claims, and not by the specific details provided by the method of description and interpretation of the embodiments herein.
Claims
1. 1. A decoder comprising: a processor configured to decode a multi-layer video data stream including video content encoded using inter-layer prediction in units of portions of pictures of the video content for a plurality of layers, wherein each portion is encoded into a payload packet of a sequence of packets of the multi-layer video data stream, the sequence of packets being partitioned into a sequence of access units; each access unit includes payload packets associated with a common instant in time and is subdivided into two or more decoding units, each decoding unit including at least one payload packet associated with one of the plurality of layers; Each access unit further comprises: first timing control information signaling a first decoder buffer search time for retrieving the respective access unit; second timing control information signaling, for each decoding unit of the respective access unit, a second decoder buffer search time for searching the decoding unit within the respective access unit according to a sequential order associated with the decoding unit in the multi-layer video data stream; Equipped with The processor: emptying a decoder buffer for buffering the multi-layer video data stream on an access unit basis according to a first condition based on the first timing control information and regardless of the second timing control information signaling the second decoder buffer search time for each decoding unit; configured to empty a decoder buffer for buffering the multi-layer video data stream in units of decoding units in the sequential order based on the second decoder buffer search time in accordance with a second condition different from the first condition; the decoding units associated with different layers are interleaved in the multi-layer video data stream; the processor is configured to remove decoding units of the access units in accordance with the layer order by deinterleaving when emptying the decoder buffer on an access unit basis; Decoder.
2. 1. A decoder comprising: a processor configured to decode a multi-layer video data stream including video content encoded using inter-layer prediction in units of portions of pictures of the video content for a plurality of layers, wherein each portion is encoded into a payload packet of a sequence of packets of the multi-layer video data stream, the sequence of packets being partitioned into a sequence of access units; each access unit includes payload packets associated with a common instant in time and is subdivided into two or more decoding units, each decoding unit including at least one payload packet associated with one of the plurality of layers; Each access unit further comprises: first timing control information signaling a first decoder buffer search time for retrieving the respective access unit; second timing control information signaling, for each decoding unit of the respective access unit, a second decoder buffer search time for searching the decoding unit within the respective access unit according to a sequential order associated with the decoding unit in the multi-layer video data stream; Equipped with The processor: emptying a decoder buffer for buffering the multi-layer video data stream on an access unit basis according to a first condition based on the first timing control information and regardless of the second timing control information signaling the second decoder buffer search time for each decoding unit; configured to empty a decoder buffer for buffering the multi-layer video data stream in units of decoding units in the sequential order based on the second decoder buffer search time in accordance with a second condition different from the first condition; the decoder buffer consists of one partial buffer for each layer, and the processor is configured to buffer, for each layer, the decoding unit including packets associated with the respective layer in the partial buffer for the respective layer. Decoder.
3. 1. A decoder comprising: a processor configured to decode a multi-layer video data stream including video content encoded using inter-layer prediction in units of portions of pictures of the video content for a plurality of layers, wherein each portion is encoded into a payload packet of a sequence of packets of the multi-layer video data stream, the sequence of packets being partitioned into a sequence of access units; each access unit includes payload packets associated with a common instant in time and is subdivided into two or more decoding units, each decoding unit including at least one payload packet associated with one of the plurality of layers; Each access unit further comprises: first timing control information signaling a first decoder buffer search time for retrieving the respective access unit; second timing control information signaling, for each decoding unit of the respective access unit, a second decoder buffer search time for searching the decoding unit within the respective access unit according to a sequential order associated with the decoding unit in the multi-layer video data stream; Equipped with The processor: emptying a decoder buffer for buffering the multi-layer video data stream on an access unit basis according to a first condition based on the first timing control information and regardless of the second timing control information signaling the second decoder buffer search time for each decoding unit; configured to empty a decoder buffer for buffering the multi-layer video data stream in units of decoding units in the sequential order based on the second decoder buffer search time in accordance with a second condition different from the first condition; the decoder buffer comprises a plurality of partial buffers, each partial buffer being associated with a subset of the layers, and the processor is configured to, for each layer, buffer the decoding unit including packets associated with the respective layer in the partial buffer for the subset of layers to which the respective layer belongs; Decoder.
4. The decoder of claim 3 , wherein the processor is configured to synchronize one partial buffer with another partial buffer.
5. 2. The decoder of claim 1, wherein the second timing control information is spread over several timing control packets each preceding a decoding unit with which the timing control packet is associated, and each timing control packet indicates the second decoder buffer search time for the preceding decoding unit.
6. 1. An encoder comprising: a processor configured to encode video content into a multi-layer video data stream using inter-layer prediction, the video content being encoded as a plurality of layers in units of portions of pictures of the video content, the video content being encoded as a plurality of layers, the sequence of packets being divided into a sequence of access units, each access unit including payload packets associated with a common instant in time and being subdivided into two or more decoding units, each decoding unit including at least one payload packet associated with one of the plurality of layers; Each access unit further comprises: first timing control information signaling a first decoder buffer search time for retrieving the respective access unit; second timing control information signaling, for each decoding unit of the respective access unit, a second decoder buffer search time for searching the decoding unit within the respective access unit according to a sequential order associated with the decoding unit in the multi-layer video data stream; Equipped with The processor: emptying a decoder buffer for buffering the multi-layer video data stream on an access unit basis according to a first condition based on the first timing control information and regardless of the second timing control information signaling the second decoder buffer search time for each decoding unit; configured to empty a decoder buffer for buffering the multi-layer video data stream in units of decoding units in the sequential order based on the second decoder buffer search time in accordance with a second condition different from the first condition; the decoding units associated with different layers are interleaved in the multi-layer video data stream; the processor is configured to remove decoding units of the access units in accordance with the layer order by deinterleaving when emptying the decoder buffer on an access unit basis; encoder.
7. 7. The encoder of claim 6, wherein the decoder buffer consists of one partial buffer for each layer, and the processor is configured to buffer, for each layer, the decoding unit including packets associated with the respective layer in the partial buffer for the respective layer.
8. 7. The encoder of claim 6, wherein the decoder buffer comprises a plurality of partial buffers, each partial buffer of the plurality of partial buffers being associated with a subset of the layers, and the processor is configured to buffer, for each layer, the decoding unit including packets associated with the respective layer in the respective partial buffer for the subset of layers to which the respective layer belongs.
9. The encoder of claim 8 , wherein the processor is configured to synchronize one partial buffer with another partial buffer.
10. 7. The encoder of claim 6, configured to spread the second timing control information, indicating the second decoder buffer search time for the decoding unit preceded by each timing control packet, to several timing control packets each preceding a decoding unit with which the timing control packet is associated.