An encoding concept that enables efficient multi-view / layer coding.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- DOLBY VIDEO COMPRESSION LLC
- Filing Date
- 2025-12-11
- Publication Date
- 2026-05-11
AI Technical Summary
Existing multi-view/layer coding technologies, such as HEVC SVC/MVC, do not efficiently utilize redundancy between views and can lead to increased decoding/encoding latency due to inter-view prediction across spatial segment boundaries and strict sequential processing of NAL units, which hinders parallel processing and increases end-to-end delay.
Implementing modifications to inter-view prediction by restricting it to a single spatial segment, allowing parallel decoding/encoding of views, and optimizing NAL unit arrangement to reduce latency and enable efficient parallel processing without compromising encoding efficiency.
Reduces decoding/encoding latency and maintains encoding efficiency by minimizing inter-view prediction across spatial segment boundaries and optimizing NAL unit processing, facilitating parallel processing across multiple views/layers.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical Field]
[0001] This application relates to an encoding concept that enables efficient multiview / layer encoding, such as multiview image / video encoding. [Background technology]
[0002] In prior art, the concept of scalable coding is known. In video coding, for example, H.264 allows for the addition of supplemental enhancement layer data to the base layer encoded video data stream to improve the playback quality of base layer quality video in different aspects such as spatial resolution, signal-to-noise ratio (SNR), and / or, importantly, the number of views. The recently completed HEVC standard is also extended by the SVC / MVC profile (SVC = Scalable Video Coding, MVC = Multi-View Coding). HEVC differs from its predecessor H.264 in many aspects, such as its suitability for parallel decoding / coding and low-latency transmission. As far as parallel coding / coding is concerned, HEVC supports WPP (Wavefront Parallel Processing) coding / coding as well as the tile parallel processing concept. According to the WPP concept, individual images are segmented into substreams in a row-wise manner. The coding order within each substream is oriented from left to right. Substreams have a predetermined decoding order between them, leading from the top substream to the bottom substream. Entropy coding of substreams is performed using probabilistic adaptation. Probabilistic initialization is performed individually for each substream, or based on a pre-adapted state of probabilities used in the entropy coding of the immediately preceding substream up to a specific position from the left edge of each preceding substream, such as the end of a second CTB (Coding Tree Block). Spatial prediction does not need to be restricted; that is, spatial prediction can traverse the boundaries of immediately following substreams. Thus, this type of substream can be coded / decoded in parallel depending on the current coding / decoding arrangement, forming a wavefront that follows in a biased manner, leading from left to right and from bottom left to top right. According to the tile concept, the image is segmented into tiles, and spatial prediction across tile boundaries is prohibited in order to provide objects that can be processed in parallel for coding / decoding of these tiles. Simply in-loop filtering across tile boundaries is permitted. To support low-latency processing, the slice concept is extended.A slice may be allowed to switch to a new entropy probability initialization that either employs an entropy probability preserved during processing of the previous substream, i.e., the substream preceding the substream to which the current slice belongs, or an entropy probability that is continuously updated up to the end of the nearest preceding slice. This means that WPP and tile concepts become more suitable for low-latency processing.
[0003] However, it is preferable to have readily available concepts that further improve upon the multi-view / layer coding concept. [Disclosure of the Invention]
[0004] Therefore, an object of the present invention is to provide a concept that further improves the multi-view / layer coding concept.
[0005] This objective is achieved by the subject matter of the independent claims in the application.
[0006] A first aspect of this application relates to multi-view coding. In particular, the idea underlying the first aspect is as follows: Inter-view prediction, on the one hand, is useful in utilizing redundancy between multiple views in which a particular scene is captured, thereby increasing coding efficiency. On the other hand, inter-view prediction prevents multiple views from being decoded / coded completely independently of each other, i.e., from being decoded / coded in parallel to benefit from, for example, a multi-core processor. More precisely, inter-view prediction makes a portion of a second view dependent on the corresponding reference portion of a first view, and this interrelationship between the portions of the first and second views requires that the decoding / coding offset / delay between particular views match when the first and second views are decoded / coded in parallel. The idea underlying the first aspect is that, with respect to how inter-view prediction is performed at the spatial segment boundary of the spatial segment in which the first / reference view is divided, this inter-view coding offset can be substantially reduced by merely reducing coding efficiency in a minor way when coding and / or decoding is changed. Modifications can be made such that the inter-view prediction from the first view to the second view does not combine any information for different spatial segments of the first view, but rather predicts the second view and its syntax elements from information arising from only one spatial segment of the first view. According to embodiments, the modifications are made more strictly so that the inter-view prediction does not further cross spatial segment boundaries, i.e., one spatial segment has a location or portion that occupies the same location. When considering the consequences of combined information arising from two or more spatial segments of the first view in the inter-view prediction, the benefits of modifying the inter-view prediction at segment boundaries become clear. In that case, the encoding / decoding of any portion of the second view, including this type of combination of inter-layer predictions, must be postponed until the encoding / decoding of all spatial segments of the first view that are combined by the inter-layer predictions.Changes in inter-view predictions at the spatial segment boundaries of the spatial segments of the first view, however, solve this problem, and each part of the second view can be encoded / decoded immediately as soon as one spatial segment of the first view is decoded / encoded. Encoding efficiency is reduced only slightly, however, since inter-layer predictions are still substantially acceptable, and the limitation applies only to the spatial segment boundaries of the spatial segments of the first view. According to the embodiment, the encoder addresses changes in inter-layer predictions at the spatial segment boundaries of the spatial segments of the first view in order to avoid the coupling of two or more spatial segments of the first view as just outlined, signals this avoidance / circumstance to the decoder, and subsequently uses signaling as a corresponding guarantee to reduce the delay of inter-view decoding in response to the signaling, for example. According to other embodiments, the decoder can also modify the method of inter-layer prediction triggered by signaling in the data stream, thereby reducing the amount of side information required to control inter-layer prediction as far as these spatial segment boundaries are concerned. Thus, the constraints on setting inter-layer prediction parameters at the spatial segment boundaries of the first view can be utilized in the formation of the data stream.
[0007] A second aspect of the present invention relates to multi-layer video encoding and circumstances under which NAL units, typically representing images from multiple layers, are aggregated into access units such that NAL units related to one time instant form one access unit regardless of the layer, or, regardless of the possibility of selection, the NAL units for each different pair of time instants and layers are processed separately and ordered without interleaving, so that there is one access unit for each different pair of time instants and layers. That is, NAL units belonging to one particular time instant and layer are sent before proceeding to NAL units for other pairs of time instants and layers. Interleaving is not permitted. However, since the encoder prevents the sending of NAL units belonging to dependent layers between NAL units belonging to the base layer, this hinders further reduction of end-to-end delay, and this opportunity arises, however, from inter-layer parallel processing. A second aspect of the present invention abandons the strict sequential non-interleaved arrangement of NAL units in the transmitted bitstream and, to achieve this objective, collects all NAL units in one time instant: all NAL units in one time instant are collected in one access unit, reusing the first possibility of defining an access unit such that the access unit is still arranged in an uninterleaved manner within the transmitted bitstream. However, interleaving of NAL units in one access unit is permitted, and NAL units of one layer are distributed by NAL units of other layers. A run of NAL units belonging to one layer within one access unit forms a decoding unit. Interleaving is permitted to the extent that, for each NAL unit in one access unit, the necessary information for inter-layer prediction is contained in any of the preceding NAL units within that access unit.The encoder can signal whether interleaving has been applied within the bitstream, and the decoder can then, based on the signaling, use, for example, multiple buffers to resort the interleaved NAL units of different layers for each access unit, or simply use one buffer if there is no interleaving. However, this does not result in a penalty to encoding efficiency due to the reduction in end-to-end latency.
[0008] A third aspect of the present application relates to the signaling of layer indices per bitstream packet, such as per NAL unit. According to the third aspect of the present application, the inventors recognized that applications can be divided into two types. Typical applications require a moderate number of layers and are therefore not troubled by layer ID fields in each packet configured to cover a moderate number of layers overall. More complex applications requiring an excessive number of layers occur only rarely. Therefore, according to the third aspect of the present application, signaling of a layer identification extension mechanism in a multilayer video signal is used to signal that a layer identification syntax element in each packet, together with layer identification extension in a multilayer data stream, either completely or simply partially determines the layer of each packet, or is completely replaced / discarded by layer identification extension. This means consumes bitrate only in rare applications where layer identification extension is necessary, but enables efficient signaling of layer associations in most cases.
[0009] A fourth aspect of the present application relates to signaling of inter-layer prediction dependencies between different levels of information in which video material is encoded into a multi-layer video data stream. According to the fourth aspect, the first syntax structure is the number of dependency dimensions and the maximum N rank levels per dependency dimension i. iThis defines a bijective mapping, a mapping of each level to at least one of the subsets of available points in the dependency space, and a second syntax structure per dependency dimension i. The latter defines the dependencies in the layer. Each syntax structure defines N of the dependency dimension i to which each second syntax structure belongs. i This describes dependencies within the rank level. Thus, the effect of defining dependencies increases linearly with the number of dependency dimensions, but the constraints on interdependencies between individual layers imposed by this signaling are relatively low.
[0010] Naturally, all of the above embodiments can be combined in pairs, triplets, or all of them. [Brief explanation of the drawing]
[0011] Preferred embodiments of the present invention will be described below with reference to the following drawings. [Figure 1] Figure 1 shows a video encoder that serves as a useful example of implementing one of the multilayer encoders described in more detail in the following figures. [Figure 2] Figure 2 shows a schematic block diagram illustrating a video decoder compatible with the video encoder in Figure 1. [Figure 3] Figure 3 shows a schematic diagram of the image after it has been subdivided into substreams for WPP processing. [Figure 4] Figure 4 shows a schematic diagram illustrating the image of one of the layers that has been subdivided into blocks by instructing further subdivision of the image into spatial segments. [Figure 5] Figure 5 shows a schematic diagram of the image of either layer subdivided into blocks or tiles. [Figure 6] Figure 6 shows a schematic diagram of the image subdivided into blocks and subsystems. [Figure 7]Figure 7 shows a schematic diagram of a base view image and a dependent view image, both registered with each other, with a dependent view image placed in front of the base view image to illustrate the limitation of the domain of assumed disparity vector values for a base view block near a spatial segment boundary, in comparison with a base view block positioned further away from the spatial segment boundary. [Figure 8] Figure 8 shows a schematic block diagram of an encoder supporting inter-view prediction limits and spatial segment boundaries according to an embodiment. [Figure 9] Figure 9 shows a schematic block diagram of a decoder that fits the encoder in Figure 8. [Figure 10] Figure 10 shows a schematic diagram illustrating possible modifications to the inter-view prediction and / or inter-view predicted dependent view block process using disparity vector values to illustrate how the domain constraints of the assumed disparity vector are determined. [Figure 11] Figure 11 shows a schematic diagram illustrating inter-view prediction, which predicts the parameters of a dependent view, in order to illustrate the application of inter-view prediction. [Figure 12] Figure 12 shows a schematic diagram of an image subdivided into code blocks and tiles, respectively, by tiles consisting of code blocks and an integer multiple of the decoding order, as defined within the code blocks that follow the subdivision of the image into tiles. [Figure 13a] Figure 13a shows the modified VPS syntax portion as an example of incorporating the embodiments of Figures 8-11 into HEVC. [Figure 13b] Figure 13b is the corresponding part of Figure 13a, but this part indicates the portion that belongs to the SPS syntax. [Figure 13c] Figure 13c shows another exemplary portion of the modified VPS syntax. [Figure 13d] Figure 13d shows an example of the modified VPS syntax that implements signaling of changes in cross-view predictions. [Figure 13e]FIG. 13e shows a further example for a modified VPS syntax that signals changes in inter-view prediction at spatial segment boundaries. [Figure 13f] FIG. 13f shows a further example for a part from the modified VPS syntax. [Figure 13g] FIG. 13g shows a part of the modified VPS syntax as a further possibility for signaling changes / limitations in inter-view prediction. [Figure 14] FIG. 14 shows an example of a part of a modified VPS syntax for a further embodiment that signals changes / limitations in inter-view prediction. [Figure 15] FIG. 15 shows a schematic diagram of an overlay of a dependent view image and a base view image shown in the upper part of the other for illustrating possible modifications to a base layer filter process at a spatial segment boundary, where the modifications can be triggered in parallel with changes / limitations in inter-view prediction according to an embodiment. [Figure 16] FIG. 16 shows a schematic diagram of a multi-layer video data stream, where Option 1 and Option 2 for arranging NAL units belonging to each time instance and each layer within the data stream are illustrated in the lower half of the figure, and here three layers are exemplarily provided. [Figure 17] FIG. 17 shows a schematic diagram of a part of a data stream that illustrates these two options in an exemplary case of two layers. [Figure 18] FIG. 18 shows a schematic block diagram of a decoder configured to process the multi-layer video data stream according to Option 1 of FIGS. 16 and 17. [Figure 19] FIG. 19 shows a schematic block diagram of an encoder that fits the decoder of FIG. 18. [Figure 20] FIG. 20 shows a schematic diagram of an image sub-streamed for WPP processing by additionally indicating a resulting wavefront when decoding / encoding images in parallel using WPP according to an embodiment. [Figure 21]Figure 21 shows a multilayer video data stream for three views, each from three decoding units that are not interleaved. [Figure 22] Figure 22 shows a schematic diagram of the multilayer video data stream described in Figure 21, with an interleaved view configuration. [Figure 23] Figure 23 shows a schematic diagram of the portion of the multilayer video data stream in the internal sequence of the NAL unit to illustrate the constraints that are likely to be observed during layer interleaving within the access unit. [Figure 24] Figure 24 shows an example of a portion of the modified VPS syntax illustrating the possibility of signaling interleaving in the decoding unit. [Figure 25] Figure 25 shows an example portion of the NAL unit header, which includes a fixed-length layer identification syntax element. [Figure 26] Figure 26 shows a portion of the VPS syntax that indicates the possibility of implementing layer identification extension mechanism signaling. [Figure 27] Figure 27 shows a portion of the VPS syntax illustrating other possibilities for implementing layer identification extension mechanism signaling. [Figure 28] Figure 28 shows a portion of the VPS syntax to illustrate a further embodiment of implementing layer identification extension mechanism signaling. [Figure 29] Figure 29 shows a portion of the slice segment header to illustrate the possibility of implementing layer identification extension in the data stream. [Figure 30] Figure 30 shows a portion of the slice segment header to illustrate further possibilities for implementing layer identification extension. [Figure 31] Figure 31 shows the portion of the VPS syntax illustrating the implementation of layer identification extension. [Figure 32] Figure 32 shows a portion of the datastream syntax illustrating other possibilities for achieving layer identification extension. [Figure 33]Figure 33 schematically illustrates a camera setup to illustrate the possibility of combining a layer identification syntax element with a layer identification extension according to an embodiment. [Figure 34a] Figure 34a shows the portion of the VPS extension syntax that signals the framework for the layer extension mechanism within the data stream. [Figure 34b] Figure 34b shows the portion of the VPS extension syntax that signals the framework for the layer extension mechanism within the data stream. [Figure 35] Figure 35 shows a schematic diagram of a decoder configured to process a multilayer video data stream, each of which provides signaling for the layers of the data stream, the arrangement of these layers in the dependency space, and the dependencies between the layers. [Figure 36] Figure 36 shows a schematic diagram illustrating the dependency space, which shows the direct dependency structure of layers in a two-dimensional space by each dimension in the space, using a specific prediction structure. [Figure 37] Figure 37 shows a schematic diagram of an array of direct dependency flags that identify dependencies between different layers. [Figure 38] Figure 38 shows a schematic diagram of two arrays of direct position dependency flags that identify dependencies between different positions and different dimensions. [Figure 39] Figure 39 shows a portion of the datastream syntax illustrating how to signal the portion of the first syntax structure that defines the dependency space. [Figure 40] Figure 40 shows a portion of the data stream illustrating the potential signaling of a portion of the first syntax structure regarding the mapping between data stream layers and available points in the dependency space. [Figure 41] Figure 41 shows a portion of the data stream illustrating the possibility of defining a second syntactic structure that describes dependencies at the dependency dimension level. [Figure 42] Figure 42 shows other possibilities for defining the second syntax structure. [Modes for carrying out the invention]
[0012] Firstly, as an overview, an example of an encoder / decoder structure that fits one of the concepts presented subsequently is given.
[0013] Figure 1 shows the general structure of an encoder according to an embodiment. The encoder 10 can be implemented to operate in a multithreaded manner or not, i.e., simply in a single thread. That is, the encoder 10 can be implemented, for example, using multiple CPU cores. In other words, the encoder 10 can support parallel processing, but is not necessarily required. The generated bitstream can also be generated / decoded by a single-threaded encoder / decoder. The coding concept of the present invention enables a parallel-processing encoder that efficiently applies parallel processing without compromising compression efficiency. Regarding parallel processing capabilities, a similar description is valid for the decoder described later with respect to Figure 2.
[0014] Encoder 10 is a video encoder, but it can also typically be an image encoder. Image 12 of video 14 is presented as input to encoder 10 at input 16. Image 12 represents a specific scene, i.e., image content. However, encoder 10 also receives other images 15 related to the same time instant, both images 12 and 15 belonging to different layers, at its input 16. For the sake of explanation only, image 12 is shown as belonging to layer 0, while image 15 is shown as belonging to layer 1. Figure 1 shows that layer 1 can contain a higher spatial resolution than layer 0, i.e., it can represent the same scene with a larger number of image samples, but this is for the sake of explanation only, and alternatively, image 15 of layer 1 could have the same spatial resolution but differ in the direction of view compared to layer 0, i.e., images 12 and 15 could be captured from different viewpoints. It should be noted that the terms base and augmented layers used in this document can refer to any set of reference layers and subordinate layers in a layer hierarchy.
[0015] Encoder 10 is a hybrid encoder, where images 12 and 15 are predicted by a predictor 18, and the predicted residuals 20 obtained by the residual decisionor 22 are subject to transformations such as spectral decomposition like DCT, and quantization in the transform / quantization module 24. The thus obtained transformed and quantized predicted residuals 26 are subject to entropy coding in the entropy encoder 28, such as arithmetic coding with context adaptation or variable-length coding. A reconstructible version of the residuals becomes available to the decoder, i.e., the inversely quantized and retransformed residual signal 30 is reconstructed by the retransform / inversely quantization module 31 and recombined with the predicted signal 32 of the predictor 18 by the coupler 33, thereby resulting in the reconstructions 34 of images 12 and 15, respectively. However, encoder 10 operates on a block basis. Therefore, the reconstructed signal 34 is affected by discontinuities at block boundaries, and thus a filter 36 is applied to the reconstructed signal 34 to produce reference images 38 for images 12 and 15, respectively, and based on this, the predictor 18 subsequently predicts the encoded images of the different layers. As shown by the dashed line in Figure 1, the predictor 18 can, however, directly utilize the reconstructed signal 34 without the filter 36 or an intermediate version, as in other prediction modes such as the spatial prediction mode.
[0016] The predictor 18 can select from different prediction modes to predict a particular block of image 12. Such a block 39 of image 12 is illustrated in Figure 1. Block 39, which represents any block of image 12 into which image 12 is divided, can be a temporal prediction mode in which it is predicted based on a previously encoded image of the same layer, such as image 12'. There can also be a spatial prediction mode in which block 39 is predicted based on a previously encoded portion of the same image 12, an adjacent block 39. Block 41 of image 15 is also illustrated in Figure 1 in which it represents any other block into which image 15 is divided. For block 41, the predictor 18 can support the prediction modes just described, namely the temporal and spatial prediction modes. In addition, the predictor 18 can provide an interlayer prediction mode in which block 41 is predicted based on a corresponding portion of image 12 of a lower layer. "Corresponding portion" means a spatial correspondence, i.e., a portion in image 12 represents a portion of the same scene as block 41 predicted in image 15.
[0017] The predictions of the predictor 18 can, of course, not be limited to image samples. The predictions can also be applied to coding parameters, namely the prediction mode, the motion vector for temporal predictions, the disparity vector for multi-view predictions, and so on. Simply put, the residuals can then be coded in the bitstream 40. That is, the coding parameters can be predictively coded / decoded using spatial and / or inter-layer predictions. Furthermore, disparity compensation can be used here.
[0018] A specific syntax is used to edit the quantized residual data 26, as well as the coding parameters, including the prediction mode and prediction parameters for individual blocks 39 and 41 of images 12 and 15 determined by the predictor 18, the transformation coefficient levels and other residual data, and the syntax elements of this syntax are subject to entropy coding by the entropy encoder 28. The data stream 40 thus obtained as the output of the entropy encoder 28 forms the bitstream 40 output by the encoder 10.
[0019] Figure 2 shows a decoder that fits the encoder of Figure 1 and can decode the bitstream 40. The decoder of Figure 2 is typically indicated by reference numeral 50 and comprises an entropy decoder, a retransform / inverse quantization module 54, a coupler 56, a filter 58, and a predictor 60. The entropy decoder 42 receives the bitstream and performs entropy decoding to reconstruct the residual data 62 and coding parameters 64. The retransform / inverse quantization module 54 inverse quantizes and retransforms the residual data 62 and transfers the residual signal thus obtained to the coupler 56. The coupler 56 also receives the prediction signal 66 from the predictor 60 and then forms a prediction signal based on the reconstructed signal 68 determined by the coupler 56 combining the prediction signal 66 and the residual signal 65 using the coding parameters 64. The prediction mirrors the prediction ultimately selected by the predictor 18, i.e., the same prediction modes are available, and these modes are selected for individual blocks of images 12 and 15 and guided (steering) according to the prediction parameters. As already described above with respect to Figure 1, the predictor 60 may, or in addition, use a filtered version or some intermediate version of the reconstructed signal 68. The images of the different layers to be ultimately reconstructed and the output images of the decoder 50's output 70 can similarly be determined for an unfiltered version or some filtered version of the combined signal 68.
[0020] According to the tile concept, images 12 and 15 are subdivided into tiles 80 and 82, respectively, and the predictions for blocks 39 and 41 within these tiles 80 and 82 are restricted to using data from the same tiles of the same images 12 and 15, respectively, as the basis for spatial prediction. This means that the spatial prediction for block 39 is restricted to using previously encoded portions of the same tile, but the temporal prediction mode is not restricted to relying on information from previously encoded images such as image 12'. Similarly, the spatial prediction mode for block 41 is restricted to using only previously encoded data from the same tile, but the temporal and inter-layer prediction modes are not restricted. The subdivision of images 15 and 12 into six tiles is selected simply for explanatory purposes. The subdivision into tiles can be individually selected and signaled within bitstream 40 for images 12', 12 and 15, and 15', respectively. The number of tiles per image 12 and 15 can be 1, 2, 3, 4, 6, etc., and tile division can be limited to regular divisions of tiles into rows and columns. For completeness, it should be noted that the method of encoding tiles separately cannot be limited to intra-prediction or spatial prediction, but can include any prediction of encoding parameters across tile boundaries and context selection in entropy coding. That is, the latter can be limited to dependent only on data within the same tile. In this way, the decoder can perform the parallel, i.e., tile-based operations just mentioned.
[0021] The encoders and decoders in Figures 1 and 2 may, or in addition, utilize the WPP concept. See Figure 3. The WPP substream 100 represents the spatial partitioning of images 12 and 15 into WPP substreams. In contrast to tiles and slices, the WPP substream does not impose any restrictions on prediction and context selection across the WPP substream 100. The WPP substream 100 extends row-wise across rows of LCUs (Largest Encoding Units) 101, i.e., the largest possible blocks that the predictive coding mode can individually transmit in the bitstream, and only one compromise is made with respect to entropy coding to enable parallel processing. In particular, order 102 is defined within the WPP substreams 100, followed exemplarily from top to bottom, and for each WPP substream 100, except for the first WPP substream in order 102, the probability estimates for the symbol alphabet, i.e., the entropy probabilities, are not completely reset, but are set to be adopted from or equal to the resulting probability after entropy encoding / decoding the preceding WPP stream up to its second LCU, by the LCU order, or for each substream, the decoding order of the substreams, such as the left-handed side indicated by arrow 106, which starts on the same side of images 12 and 15 respectively and leads to the other side in the row direction of the LCU, as indicated by line 104. Therefore, by preserving some encoding delays between sequences of the same WPP substreams 12 and 15, these WPP substreams 100 can be decoded / encoded in parallel, i.e., simultaneously, such that the portions of each image 12, 15 that are encoded / decoded move across the image in a tilting manner from left to right, forming a kind of wavefront 108.
[0022] It is noteworthy that orders 102 and 104 also define a raster scan that leads row by row from top to bottom within the LCU, from the top-left LCU 101 to the bottom-right LCU. Each WPP substream can correspond to one LCU row. Returning to tiles for a brief reference, the latter can also be restricted to be aligned to LCU boundaries. A substream can be fragmented into one or more slices without being constrained to LCU boundaries, as far as the boundaries between two slices within the substream. Entropy probabilities are employed, however, in the case of transitions from one slice to the next within a substream. In the case of tiles, all tiles can be combined into one slice, or a single tile can be fragmented into one or more slices by again not being constrained to LCU boundaries, as far as the boundaries between two slices within the tile. In the case of tiles, the order within the LCU is first modified in the raster scan order to traverse the tile in the tile order before moving to the next tile in the tile order.
[0023] As described above, image 12 can be divided into tiles or WPP substreams, and similarly, image 15 can also be divided into tiles or WPP substreams. Theoretically, the WPP substream division / concept can be selected for one of images 12 and 15, while the tile division / concept can be selected for the other of the two. Alternatively, restrictions can be imposed on the bitstream by requiring that the concept type, i.e., tile or WPP substream, be the same within the layer. Another embodiment of spatial segmentation involves slices. Slices are used to segment bitstream 40 for transmission purposes. Slices are packed into NAL units, which are the smallest entities for transmission. Each slice is independently encodeable / decodeable; that is, any predictions crossing slice boundaries are prohibited, just as with context selection, etc. These, together, are three embodiments of spatial segmentation: slices, tiles, and WPP substreams. In addition, all three parallelization concepts—tile, WPP substream, and slice—can be used in combination, and image 12 or image 15 can be divided into tiles, each tile being divided into multiple WPP substreams. Slices can also be used to divide a bitstream into multiple NAL units, for example (but not limited to) tile or WPP boundaries. When images 12 and 15 are divided using tiles or WPP substreams and also using slices, and the slice divisions are displaced from other WPP / tile divisions, the spatial segments are defined as the smallest, independently decodeable sections of images 12 and 15. Alternatively, combinations of concepts can be used within an image (12 or 15), and / or restrictions can be imposed on the bitstream if boundaries must be aligned between concepts used differently.
[0024] To enable parallel processing concepts such as tiling and / or WPP concepts, various prediction modes supported by encoders and decoders, as well as limitations imposed on these prediction modes, and contextual derivations for entropy coding / decoding have been described above. It has also been mentioned that encoders and decoders can operate in a block-based manner. For example, the prediction modes described above can be selected in a block-based manner, i.e., with a granularity finer than the image itself. Before proceeding to describe embodiments of the present application, the relationships between slices, tiles, WPP substreams and the blocks just mentioned in the embodiments will be explained.
[0025] Figure 4 shows an image that can be either a Layer 0 image like Image 12 or a Layer 1 image like Image 15. The image is regularly subdivided into an array of Block 90. Sometimes, these blocks 90 are referred to as the Largest Coding Block (LCB), Largest Coding Unit (LCU), Coding Tree Block (CTB), etc. Subdividing the image into blocks 90 can form a kind of base or coarsest granularity on which the prediction and residual coding described above is performed. This coarsest granularity, i.e., the size of the blocks 90, can be signaled and set individually for layer 0 and layer 1 by the encoder. For example, a multi-tree such as a quad-tree subdivision can be used, and each block 90 can be signaled within the data stream to be subdivided into a prediction block, a residual block, and / or coding block, respectively. In particular, the coded block can be a leaf block of a recursive multitree subdivision of block 90, and several prediction-related decisions can be signaled at the granularity of the coded block, such as the prediction mode, and prediction blocks in which prediction parameters such as motion vectors in the case of temporal interpretation and disparity vectors in the case of interlayer prediction are coded at that granularity, and residual blocks in which prediction residuals are coded at that granularity can be separate leaf blocks of a recursive multitree subdivision of the coded block.
[0026] The raster scan coding / decoding order 92 can be defined within a block 90. The coding / decoding order 92 restricts the availability of adjacent parts for the purpose of spatial prediction: simply, according to the coding / decoding order 92, parts of the image preceding the current part, such as block 90 or some smaller blocks, to which the syntax element currently being predicted relates, are available for spatial prediction within the current image. Within each layer, the coding / decoding order 92 traverses all blocks 90 of the image in order to continue traversing the next block of the image in each layer, in an image coding / decoding order that does not necessarily follow the temporal playback order of the image. Within individual blocks 90, the coding / decoding order 92 is refined into a scan within fewer blocks, such as coding blocks.
[0027] In the relationship between the block 90 and fewer blocks outlined, each image is further subdivided into one or more slices along the coding / decoding order 92 mentioned. Thus, the slices 94a and 94b shown exemplarily in Figure 4 cover their respective images without gaps. The boundary or interface 96 between consecutive slices 94a and 94b of an image may or may not align with the boundary of an adjacent block 90. More precisely, and illustrated on the right-hand side of Figure 4, consecutive slices 94a and 94b within an image may touch each other at the boundary of a coding block, i.e., a smaller block such as a leaf block of one subdivision of block 90.
[0028] Image slices 94a and 94b can form the smallest unit into which the portion of the data stream from which the image is encoded can be packetized, i.e., into packets, i.e., NAL units. Further possible properties of slices, namely, restrictions on slices regarding, for example, the determination of prediction and entropy context across slice boundaries, have been described above. Slices with this kind of restriction can be called "normal" slices. In addition to normal slices, "dependent slices" can also exist, as will be outlined in more detail below.
[0029] When the tile division concept is applied to an image, the encoding / decoding order 92 defined within the sequence of blocks 90 can be changed. This is illustrated in Figure 5, which illustrates that the image has been divided into four tiles 82a-82d. As shown in Figure 5, each tile is defined as a regular subdivision of the image in units of blocks 90. That is, each tile 82a-82d consists of an n×m sequence of blocks 90, where n is set individually for each row of the tile and m is set individually for each column of the tile. Following the encoding / decoding order 92, the blocks 90 in the first tile are first scanned in raster scan order before proceeding to the next tile 82b and so on, and tiles 82a-82d are then scanned in raster scan order themselves.
[0030] According to the WPP stream partitioning concept, the image is subdivided into WPP substreams 98a-98d, each consisting of one or more rows of block 90, in accordance with the encoding / decoding order 92. Each WPP substream can cover one complete row of block 90, for example, as illustrated in Figure 6.
[0031] The tile concept and the WPP substream concept can, however, be combined. In that case, each WPP substream would cover, for example, one row of 90 blocks within each tile.
[0032] Even image slicing can be used in conjunction with tiling and / or WPP substreaming. With respect to tiles, each of the one or more slices into which the image is subdivided can consist precisely of one complete tile, one or more complete tiles, or simply a sub-part of one tile, in accordance with the encoding / decoding order 92. Slices can also be used to form WPP substreams 98a-98d. For this purpose, the slices forming the minimum unit for packetization may consist of a normal slice on the one hand and a dependent slice on the other: the normal slice imposes the aforementioned restrictions on prediction and entropy context derivation, while the dependent slice does not impose these kinds of restrictions. A dependent slice that begins at an image boundary where the coding / decoding order 92 points substantially away from the row direction adopts the entropy context resulting from the entropy-decoded block 90 in the row immediately preceding block 90, while a dependent slice that begins somewhere else may adopt the entropy coding context resulting from the entropy coding / decoding of the preceding slice to its end. By this means, each WPP substream 98a-98d can consist of one or more dependent slices.
[0033] That is, the encoding / decoding order 92 defined within block 90 proceeds linearly from the first side of each image, here exemplary the left side, to the opposite side, exemplary the right side, and then steps to the next row in the downward / bottom direction of block 90. The available, i.e., already encoded / decoded portion of the current image is therefore essentially to the left and above the currently encoded / decoded portion, such as the current block 90. Due to the destruction of prediction and entropy context derivation across tile boundaries, tiles of one image can be processed in parallel. The encoding / decoding of tiles of one image can even be started simultaneously. The limitation arises from the in-loop filtering described above in the cases where it is permissible to cross tile boundaries. Starting the encoding / decoding of WPP substreams in order is performed in a staggered manner from top to bottom. The intra-image delay between consecutive WPP substreams is measured in block 90, which is two blocks 90.
[0034] However, parallelizing the encoding / decoding of images 12 and 15, i.e., even the time instants of different layers, is advantageous. Clearly, the encoding / decoding of image 15 in the dependent layer must be delayed compared to the encoding / decoding of the base layer to ensure that there is a "spatially corresponding" portion of the base layer already available. These considerations are valid even in cases where no parallelization of individual encoding / decoding is used within images 12 and 15. Even in cases where one slice each is used to cover all of images 12 and 15, the encoding / decoding of images 12 and 15 can be parallelized by not using tiling and WPP substream processing. The signaling described below, i.e., aspect 6, may represent this type of decoding / encoding delay between layers, even in this type of case where tiling or WPP processing is used for any of the images in the layers, or not.
[0035] Before discussing the presented concept of this application, it should be noted that the block structures of the encoder and decoder shown in Figures 1 and 2 are for illustrative purposes only, and the structures may differ.
[0036] With regard to the above description relating to the minimum coding delay between consecutive layers of coding, it should be noted that the decoder can determine the minimum decoding delay based on short syntax elements. However, in cases where long syntax elements are used to signal this inter-layer time delay in advance for a predetermined period, the decoder can plan for the future with the provided guarantees, and workload allocation within the parallel decoding of bitstream 40 can be performed more easily.
[0037] The first aspect relates to limiting inter-layer predictions within a view, particularly parallax-corrected inter-view predictions, in order to benefit from lower overall encoding / decoding latency or parallelization capabilities. Details are readily available from the following figures. For a brief explanation, please refer to Figure 7.
[0038] The encoder can restrict the available domain 301 of the disparity vector so that, for example, the current block 302 of the dependent view is inter-layer predicted at the boundary 300 of the base layer segment. 303 indicates the restriction. For comparison, Figure 7 shows another block 302' of the dependent view, where the available domain of its disparity vector is not restricted. The encoder can signal this behavior, i.e., restriction 303, in the data stream so that the decoder can take advantage of its low-latency sensing. That is, the decoder can operate just normally, ensuring that the "unavailable segment" portion is not needed as far as inter-layer prediction is with respect to the encoder, i.e., the decoder can keep the inter-layer delay lower. Alternatively, both the encoder and decoder can change their operating modes as far as inter-layer prediction is with respect to boundary 300, for example, to take advantage of a lower manifold of available states of inter-layer prediction parameters at boundary 300.
[0039] Figure 8 shows a multi-view encoder 600 configured to encode multiple views 12 and 15 into a data stream 40 using inter-view prediction. In the case of Figure 8, the number of views is exemplary selected to 2 by inter-view prediction leading from the first view 12 to the second view 15, as illustrated by arrow 602. Extensions to two or more views are readily conceivable. The same applies to the embodiments described below. The multi-view encoder 600 is configured to modify the inter-view prediction at the spatial segment boundary 300 of the spatial segment 301 into which the first view is divided.
[0040] As far as possible implementation details relating to the encoder 600 are concerned, for example, the proposed description above with respect to Figure 1. That is, the encoder 600 can be an image or video encoder and can operate in a block-based manner. In particular, the encoder 600 can be a hybrid encoding type configured to subject the first view 12 and the second view 15 to predictive encoding, insert predictive parameters into the data stream 40, convert and encode the predictive residuals into the data stream 40 using spectral decomposition, and switch between different prediction types, at least with respect to the second view 15, including spatial and inter-view predictions 602. As previously stated, the units in which the encoder 600 switches between different prediction types / modes can be called encoding blocks, and the size of these encoding blocks can vary so that they can represent, for example, leaf blocks of a hierarchical multi-tree subdivision of the image of the second view 15 or tree root blocks of the image of the second view 15 that can be periodically pre-divided. Interview prediction can result in predicting samples within each encoded block using a disparity vector 604 that indicates the displacement applied to a portion 606 of the first view 12 image, which is spatially located in the same location as the interview predicted block 302 of the image of the second view 15, in order for the samples in block 302 to access the portion 608 predicted by copying the restored version of portion 608 to block 302. Interview prediction 602, however, is not limited to that type of sample value of the second view 15. Rather, interview prediction, as supported by encoder 600, can be used to predictively encode prediction parameters itself: imagine encoder 600 supports spatial and / or temporal prediction in addition to the interview prediction modes just outlined. Spatially predicting a particular encoded block, just as temporal prediction does, ends in the prediction parameters into which that encoded block is inserted into the data stream 40.Instead of independently encoding all of these predictive parameters for the encoded blocks of the image of the second view 15 into the data stream 40, however, the encoder 600 can use predictive coding by predicting the predictive parameters used to predictively encode the encoded blocks of the second view 15 based on predictive parameters or other information available from the portion of the data stream 40 encoded by the encoder 600 for the first view 12. That is, predictive parameters for a particular encoded block 302 of the second view 15, such as motion vectors, can be predicted, for example, based on the motion vectors of the temporally predicted encoded blocks of the corresponding first view 12. The "correspondence" can take into account the parallax between views 12 and 15. For example, the first and second views 12 and 15 may each have their associated depth images, and the encoder 600 may be configured to encode texture samples from views 12 and 15, along with the associated depth values of the depth images, into a data stream 40. The encoder 600 may then use the depth estimation of encoding block 302 to determine the "corresponding encoding block" in the first view 12 whose scene content better fits the scene content of the current encoding block 302 in the second view 15. Of course, this type of depth estimation can also be determined by the encoder 600 based on the disparity vector used for the interview predicted encoding block near view 15, regardless of any encoded or unencoded depth images.
[0041] As already mentioned, the encoder 600 in Figure 8 is configured to modify inter-view predictions at spatial segment boundaries 300. That is, the encoder 600 modifies the inter-view prediction method at these spatial segment boundaries 300. The reasons and purposes for this are outlined further below. In particular, the encoder 600 modifies the inter-view prediction method in such a way that each entity of the predicted second view 15, such as the texture sample content of the inter-view predicted coded block 300 or specific prediction parameters of this type of coded block, is precisely dependent on one spatial segment 301 of the first view 12 by the inter-view prediction method 602. The benefit can be easily understood by focusing on the result of the modification of the inter-view prediction for a particular coded block whose sample value or prediction parameter is inter-view predicted. Without modification or limitation of the inter-view prediction 602, encoding this coded block must be postponed until the encoding of two or more spatial segments 301 of the first view 12 participating in the inter-view prediction 602 is completed. Therefore, encoder 600 must adhere to this interview coding delay / offset in any case, and encoder 600 cannot further reduce the coding delay by coding views 12 and 15 in a time-overlapping manner. When the interview prediction 602 is modified / corrected at the spatial segment boundary 301 in the manner outlined, in that case things are different, as the very problematic coding block 302 where several entities are interview predicted can be subject to coding as soon as one (just one) spatial segment 301 of the first view 12 is fully coded. Thereafter, the assumed coding delay is reduced.
[0042] Accordingly, Figure 9 shows a multiview decoder 620 that fits the multiview encoder of Figure 8. The multiview decoder of Figure 9 is configured to reconstruct multiple views 12 and 15 from the data stream 40 using interview predictions 602 from the first view 12 to the second view 15. As described above, the decoder 620 can redo the interview predictions 602 in the same way as expected to be done by the multiview encoder 600 of Figure 8 by reading and applying prediction parameters contained in the data stream 40, such as prediction modes indicated for each coding block of the second view 15, some of which are interview predicted coding blocks. As already described above, the interview predictions 602 may also relate to the prediction of the prediction parameters themselves, and the data stream 40 may point to a list of prediction residuals or predictors for this type of interview predicted prediction parameter, one of which may have an index that is interview predicted according to 602.
[0043] As already described with respect to Figure 8, the encoder can modify the method of inter-view prediction at boundary 300 to avoid inter-view prediction 602 which combines information from the two segments 301. The encoder 600 can achieve this in a manner transparent to the decoder 620. That is, the encoder 600 can simply impose an autonomy on its selection from possible coding parameter settings so that the combination of information from the two distinguishable segments 301 in inter-view prediction 602 is essentially avoided by the decoder 620, which simply applies set coding parameters transmitted within the data stream 40.
[0044] In other words, unless the decoder 620 is not interested in or able to apply parallel processing to the decoding of the data stream 40 by decoding views 12 and 15 in parallel, the decoder 620 can simply ignore the signaling of the encoder 600 inserted into the data stream 40, which signals the aforementioned changes in the inter-view prediction. More precisely, according to one embodiment of the present invention, the encoder in Figure 8 signals within the data stream 40 whether there is a change in the inter-view prediction at the segment boundary 300, i.e., whether there is a change or no change at the boundary 300. When signaled to apply, the decoder 620 may take a modification in the interview prediction 602 at the spatial segment boundary 300 of the spatial segment 301 as a guarantee that the interview prediction 602 is restricted at the spatial segment boundary 300 of the spatial segment 301 so as to ensure that the interview prediction 602 does not include any dependency of any part 302 of the second view 15 on any spatial segment other than the spatial segment in which the part 306 of the first view 12, which is positioned in the same location for each part 302 of the second view 15, is positioned. That is, when signaled to apply a modification in the interview prediction 602 at the boundary 300, the decoder 620 takes this as a guarantee: the interview prediction 602 does not introduce any dependency on any “adjacent spatial segment” to any block 302 of the dependent view 15 for which the interview prediction 602 is used to predict any of its samples or any of its prediction parameters. This means the following: for each part / block 302, there is a part 606 in the first view 12 that is located in the same place as each block 302 in the second view 15. "Same place" is intended to refer to a block in view 12 whose circumference precisely locally exhibits the same index as the circumference of block 302.Alternatively, “same location” is not measured with sample precision, but determining “same location” blocks results in the division of the layer 12 image into blocks and the selection of those blocks, i.e., the one incorporating the location will incorporate the same location location relative to the upper left corner of block 302 or select another representative location of block 302, measured by the granularity of the blocks into which the layer 12 image is divided. “Same location portion / block” is indicated by 606. Recall that due to the different view orientations of views 12 and 15, the same location portion 606 cannot have the same scene content as portion 302. Nevertheless, in the case of signaling changes in inter-view predictions, decoder 620 assumes that any portion / block 302 of the second view 15 subject to inter-view prediction 602 is subject to inter-view prediction 602 to the spatial segment 301 in which the simply same location portion / block 606 is located. That is, when considering 12 of 15 images of the first and second views, each registered to the other, the inter-view prediction 602 does not cross the segment boundaries 300 of the first view 12, and the segments 301 in which each block / part 302 of the second view 15 is located remain within those segments 301 that view 15. For example, the multi-view encoder 600 appropriately restricts the signaling / selected disparity vectors 604 of the inter-view predicted parts / blocks 302 of the second view 15, and / or appropriately encodes / selects the index in the predictor list so as not to index predictors containing the inter-view prediction 602 from the information of "adjacent spatial segments 301".
[0045] Before continuing to describe various possible details representing different embodiments of the encoder and decoder in Figures 8 and 9 that may or may not be able to be combined with each other, the following should be noted. From the descriptions of Figures 8 and 9, it becomes clear that there are different ways in which the encoder 600 can implement “modification / restriction” of its inter-view prediction 602. In a looser restriction, the encoder 600 simply restricts the inter-view prediction 602 in such a way that it does not combine information from two or more spatial segments. The description in Figure 9 features a stricter restriction example, which further restricts the inter-view prediction 602 from traversing spatial segments 302: that is, any part / block 302 of the second view 15 to which the inter-view prediction 602 is subject obtains its inter-view prediction exclusively from information from its spatial segment 301 of the first view 12 in which its “same-location block / part 606” is located. The encoder operates accordingly. The latter type of restriction represents a variation described in Figure 8, which is even stricter than those described earlier. In both variations, decoder 602 can take advantage of limitations. For example, decoder 620 can take advantage of limitations of inter-view prediction 602 by reducing / decreasing the inter-view decoding offset / delay in decoding the second view 15 in relation to the first view 12, when signaled to apply. Or, in addition, decoder 602 can take into account assurance signaling when deciding to perform a trial of decoding views 12 and 15 in parallel: when assurance is signaled to apply, the decoder opportunistically performs inter-view parallel processing, and otherwise refrains from the trial. For example, in the embodiment shown in Figure 9, where the first view 12 is periodically divided into four spatial segments 301, each representing one-quarter of the 12 images of the first view, decoder 620 can begin decoding the second view 15 as soon as the first spatial segment 301 of the first view 12 is fully decoded.Otherwise, assuming the disparity vector 604 is only in the horizontal natural world, the decoder 620 must wait at least for the complete decoding of both upper spatial segments 301 of the first view 12. Tighter changes / limitations on interview predictions along the segment boundary 300 make the use of guarantees easier.
[0046] The signaling of the aforementioned guarantee can have a scope / effectiveness that encompasses, for example, simply a single image or even a sequence of images. Therefore, as will be discussed later, it can be signaled in a video parameter set, a sequence parameter set, or even an image parameter set.
[0047] Up to this point, embodiments have been provided with respect to Figures 8 and 9 in which changes in the inter-view prediction 602 do not change, except for the guaranteed signaling, the data stream 40, and the method of encoding / decoding it by the encoder and decoder in Figures 8 and 9. Rather, the method of decoding / encoding the data stream remains the same regardless of whether or not self-limiting in the inter-view prediction 602 is applied. According to alternative embodiments, however, the encoder and decoder even modify the method of encoding / decoding the data stream 40 to take advantage of the guaranteed case, i.e., the limiting of the inter-view prediction 602 at the spatial segment boundary 300. For example, the domain of assumed disparity vectors that can be signaled in the data stream 40 may be limited to the inter-view predicted block / part 302 of the second view 15 near the same location on the spatial segment boundary 300 of the first view 12. See, for example, Figure 7 again. As already stated above, Figure 7 shows two exemplary blocks 302' and 302 of the second view 15, one of which, namely block 302, is close to the same location on the spatial segment boundary 300 of the first view 12. The same location on the spatial segment boundary 300 of the first view 12 is shown as 622 when it is changed to the second view 15. As shown in Figure 7, block 306, which is the same location as block 302, is close to the spatial segment boundary 300, and the vertically separated spatial segment 301a has a block 606 that is the same location, and the vertically adjacent spatial segment 301b results in an interview prediction block 302 that is copied at least partially from a sample of this adjacent spatial segment 301b, to such an extent that the disparity vector is too large and shifts the block / part 606 that is the same location to the right, i.e., toward the adjacent spatial segment 301b, in which case the interview prediction 602 crosses the spatial segment boundary 300. Therefore, in the "guarantee case," the encoder 600 cannot select this type of disparity vector for block 302, and thus the domain of the assumed disparity vectors that can be encoded for block 302 can be limited.For example, when using Huffman coding, the Huffman code used to encode the disparity vector for the inter-view predicted block 302 can be modified to take advantage of the limited domain of disparity vectors that block 302 can have. In the case of arithmetic coding, for example, other binarizations combined with a binary arithmetic scheme can be used to encode the disparity vector, or other probability distributions among the assumed disparity vectors can be used. According to this embodiment, the slight reduction in coding efficiency resulting from the limitations of inter-view prediction at the spatial segment boundary 300 can be partially compensated by reducing the amount of side information transmitted in the data stream 40 with respect to the transmission of disparity vectors to spatial segments 302 near the same location on the spatial segment boundary 300.
[0048] Thus, according to the embodiments described above, both the multiview encoder and the multiview decoder modify their methods for decoding / encoding the disparity vector from the data stream depending on whether a guarantee case applies. For example, both modify the Huffman coding used to decode / encode the disparity vector, or modify the binarization and / or probability distribution used to arithmetically decode / encode the disparity vector.
[0049] For a concrete example, Figure 10 is referenced to more clearly describe how the encoder and decoder in Figures 8 and 9 restrict the domain of assumed disparity vectors that can be signaled in the data stream 40. Figure 10 again illustrates the normal behavior of the encoder and decoder for the interview predicted block 302: the disparity vector 308 from the domain of assumed disparity vectors is determined for the current block 302, which is therefore the predicted block that is disparity corrected and predicted. The first view 12 is then sampled at reference portion 304, which is displaced by the determined disparity vector 308 from the portion 306 of the first view 12 that is currently positioned at the same location relative to block 302. The restriction of the domain of assumed disparity vectors that can be signaled in the data stream is made as follows: the restriction is made such that the reference portion 304 is entirely within the spatial segment 301a in which the portion 306 that is positioned at the same location is spatially located. The disparity vector 308 illustrated in Figure 10 does not fulfill this restriction, for example. It is therefore outside the domain of the disparity vector assumed for block 302, and according to one embodiment, as far as block 302 is concerned, it is not signalable in the data stream 40. According to an alternative embodiment, however, the disparity vector 308 is signalable in the data stream, but the encoder 600 avoids applying this disparity vector 308 in guaranteed cases and chooses to apply other prediction modes to block 302, for example, such as the spatial prediction mode.
[0050] Figure 10 also illustrates that the kernel half-width 10 of the interpolation filter can be considered to perform domain restriction of the disparity vector. More precisely, in a copy of the sample content of block 302 predicted by disparity correction from the image of the first view 12, each sample of block 302 can be obtained from the first view 12 by applying interpolation using an interpolation filter having a specific interpolation filter kernel size, in the case of the sub-percentage disparity vector. For example, the sample value illustrated with "x" in Figure 10 can be obtained by combining the samples in the filter kernel 311 where the sample position "x" is located at its center, and thus the assumed disparity vector for block 302 can be restricted so that, in that case, there are no samples in the reference portion 304, and the filter kernel 311 overlays onto the adjacent spatial segment 301b but remains within the current spatial segment 301a. The signalable domain can therefore be restricted or not restricted. According to an alternative embodiment, a sample of the filter kernel 311 placed within an adjacent spatial segment 301b can be simply loaded with respect to the sub-perspective disparity vector, except that it follows some exceptional rules to avoid additional constraints on the assumed domain of the disparity vector, while the decoder allows for replacement loading in cases where it is simply signaled to apply a guarantee.
[0051] The latter embodiment reveals that whether or not the decoder 620 can, and in addition, or otherwise, changes in the entropy decoding of the data stream, alters how interview predictions are performed at the spatial segment boundary 300 in response to the data stream, such as those inserted into the data stream by the signaling and encoder 600. For example, as mentioned above, both the encoder and decoder can differently fill the interpolation filter kernel in the portion extending beyond the spatial segment boundary 300, depending on whether or not the guarantee case applies. The same can be applied to the reference portion 306 itself. The reference portion 306 can be at least partially extended into the adjacent spatial segment 301b by filling each portion with information that is independent of any information outside of the current spatial segment 301a. In effect, the encoder and decoder can treat the spatial segment 300 as an image boundary by filling portions of the reference portion 304 and / or the interpolation filter kernel 311 by extrapolation from the current spatial segment 301a in the guarantee case.
[0052] Furthermore, as mentioned above, interview prediction 602 is not limited to predicting the sample-level content of the interview-predicted block 302. Rather, interview prediction can also be applied to predicting prediction parameters, such as predicting motion parameters related to the temporally predicted prediction of block 302 in view 15, or spatial prediction parameters related to the spatially predicted prediction of block 302. To illustrate possible modifications and limitations imposed on this type of interview prediction 602 at boundary 300, Figure 11 is referenced. Figure 11 shows a block 302 of a dependent view 15 whose parameters are predicted at least between aliases using interview prediction. For example, a list of several predictors for the parameters of block 302 can be determined by interview prediction 602. For this purpose, the encoder and decoder work, for example, as follows: The reference portion of the first view 12 is now selected for block 302. The selection or derivation of the reference portion / block 314 is performed from blocks such as the encoded block, predicted block, etc., where the 12 images of the first layer are divided. For that derivation, the representative position 318 in the first view 12 can be determined to be in the same location relative to the representative position 628 of block 302, or the representative position 630 of an adjacent block 320 adjacent to block 302. For example, the adjacent block 320 may be the block relative to the top of block 302. The determination of block 320 may include a selected block 320 from a block that is divided such that the image 15 of the second view layer has a sample immediately adjacent to the top of the sample in the upper left corner of block 302. The representative positions 628 and 630 may be the sample in the upper left corner or the sample in the center of the block, etc. The reference position 318 in the first view 12 is then the position that is in the same location relative to 628 or 630. Figure 11 illustrates the same location relative to position 628. Next, the encoder / decoder estimates the disparity vector 316. This can be done, for example, based on the estimated distance image of the current scene, or using disparity vectors that have already been decoded and are in the spatial-temporal neighborhood of block 302 or block 320.The disparity vector 316 thus determined is applied to the representative position 318 such that the head of vector 316 points to the position 632. In the division of the image of the first view 12 into blocks, the reference portion 314 is selected such that it is the portion comprising the position 632. As just mentioned, the division in which the selection of portion / block 314 is made can be a division of the view 12 into encoded blocks, predicted blocks, residual blocks and / or transformed blocks.
[0053] According to one embodiment, the multiview encoder simply checks whether the reference portion 314 is in an adjacent spatial segment 301b, i.e., in a spatial segment that does not have a block located in the same place as the same location of the reference point 628. When the encoder signals the decoder the guarantee outlined above, the encoder 600 restrains any application to the parameters of the current block 302. That is, the list of predictors for the parameters of block 302 may include interview predictors that lead to crossing the boundary 300, but the encoder 600 avoids selecting those predictors and chooses indices for block 302 that do not point to unnecessary predictors. When both the multiview encoder and decoder check, in the guarantee case, whether the reference portion 314 is in an adjacent spatial segment 301b, both the encoder and decoder may replace an interview predictor that "crosses the boundary" with other predictors, or simply remove it from a list of predictors that may also include, for example, spatially and / or temporally predicted parameters and / or one or more non-performing predictors. The status, i.e., whether or not the reference portion 314 is part of spatial segment 301a, and the conditional substitution or exclusion checks are performed only in guaranteed cases. In non-guaranteed cases, any check of whether the reference portion 314 is within spatial segment 301a can be discontinued, and the application of the predictor block 302 parameters derived from the attributes of the reference portion 314 to the prediction can be done regardless of whether or not the reference portion 314 is within spatial segment 301a or 301b. In cases where no predictor derived from the attributes of block 314 is added to the list of predictors for the current block 302, or in cases of adding an alternative predictor, each modification of the normal interview prediction is performed by the encoder and decoder 620 according to the reference block 314, whether or not it is within spatial segment 301a. By this means, any predictor index to the list of predictors for block 302 thus determined points to the same list of predictors in the decoder.The signalable domain of the index for block 302 can be restricted or not restricted depending on whether a guarantee case applies. In the case where a compensation case applies, but the encoder simply performs a check, the multiview encoder forms a list of predictors for block 302 regardless of whether the reference portion 314 is within spatial segment 301a (and even regardless of whether a guarantee case applies), however, in the prediction case, the list of predictors is derived from attributes of block 314 that are outside spatial segment 301a, and in the guarantee case, the index is restricted from selecting a predictor from there. In that case, the decoder 620 can form a list of predictors for block 302 in the same manner, i.e., in the guarantee and non-guarantee cases, in the same manner that the encoder 600 has already dealt with, that is, that interview prediction does not require any information from the adjacent spatial segment 301b.
[0054] Regarding the parameters of block 302 and the attributes of reference section 314, it should be noted that they can be motion vectors, disparity vectors, residual signals such as transformation coefficients, and / or depth values.
[0055] The inter-view predictive change concept described in relation to Figures 8-11 can be introduced into the currently envisioned extensions of the HEVC standard, in the manner described below. To that extent, the description proposed below is also interpreted as the basis for possible implementation details of the above proposed description with respect to Figures 8-11.
[0056] As an intermediate note, it should be noted that the spatial segment 301 discussed above, which forms a unit at a boundary where inter-view predictions are altered / restricted, does not necessarily form this type of spatial segment in which intra-layer parallel processing is mitigated or enabled within that unit. In other words, the spatial segments discussed above in Figures 8-11 can be tiles into which the base layer 12 is divided, but other examples are equally possible, such as the embodiment in which the spatial segment 301 forms the coded tree root block CTB of the base layer 12. In embodiments described later, the spatial segment 301 is combined with the definition of a tile, i.e., the spatial segment is a tile or a group of tiles.
[0057] Due to the limitations described for ultra-low latency and parallelization in HEVC, inter-layer prediction is constrained in a way that ensures the base layer image, particularly the tile division, is maintained.
[0058] HEVC allows the encoded base layer image's CTB to be divided into rectangular regions called tiles, which can be processed independently except for in-loop filtering, via a grid of vertical and horizontal boundaries. In-loop filtering can be turned off at the tile boundaries to make them completely independent.
[0059] When configured to reduce tile boundary artifacts, in-loop filters can traverse tile boundaries, but the dependencies of analysis and prediction are broken, especially at tile boundaries such as image boundaries. Therefore, the processing of individual tiles does not depend entirely on other tiles in the image, or on the dependencies of a wide range of filtering configurations. A constraint is imposed in that all CTBs in a tile belong to the same slice, or all CTBs in a slice belong to the same tile. As can be seen in Figure 1, the tiles force the CTB scan order to note the order of the tiles, i.e., all CTBs belonging to the first tile, e.g., the top-left tile, are passed through before the CTBs belonging to the second tile, e.g., the top-right tile. The tile structure is defined through the number and size of CTBs in the rows and columns of each tile that make up the grid in the image. This structure can either change per frame base or remain constant across the encoded video sequence.
[0060] Figure 12 shows an exemplary division of the CTB into nine tiles within the image. Thick black lines represent tile boundaries, and the numbering indicates the scan order of the CTB and also the tile order.
[0061] The HEVC extension's augmented layer tiles can be decoded as soon as all tiles covering their corresponding image area in the base layer bitstream have been decoded.
[0062] The following section describes modifications to the constraints, signaling, and coding / decoding processes that enable lower inter-layer coding offsets / delays, using the concepts shown in Figures 7-11.
[0063] The modified decoding process related to tile boundaries in HEVC can be seen as follows:
[0064] a) Motion or disparity vectors must not cross tiles in the base layer. If restraint is possible, the following applies: When cross-layer predictions (e.g., predictions of sample values, motion vectors, residual data, or other data) use the base view (layer 12) as a reference image, the disparity or motion vector is constrained to belong to the same tile as the base layer CTU on which the referenced screen area is located. In certain embodiments, the motion or disparity vector 308 is clipped in the decoding process so that the referenced image area is located within the same tile and the referenced sub-pel position is predicted only from information within the same tile. More specifically in the current HEVC sample interpolation process, this constrains the motion vector to point to a sub-pel position clipped 3-4 pixels away from the tile boundary 300, or in the cross-view motion vector, cross-view residual prediction process, this constrains the disparity vector to point to a position within the same tile. Alternative embodiments adjust the sub-pel interpolation filter to handle tile boundaries similar to the image boundary, allowing the motion vector to point to a sub-pel position located closer to the tile boundary than the kernel size 310 of the sub-pel interpolation filter. An alternative embodiment means a bitstream constraint that does not allow the use of motion or disparity vectors that are clipped in the aforementioned embodiment.
[0065] b) Blocks adjacent to a block placed in the base layer are not used if they are on different tiles. If restrictions are available, the following applies: When the base layer is used for predictions from adjacent blocks (e.g., TMVP or derivation of the disparity of adjacent blocks) and tiles are used, the following applies: If CTU B belongs to the same tile as the placed base layer CTU A, then only predictor candidates arising from a different CTU B than the placed CTU A in the base layer are used. For example, in the current HEVC derivation process, CTU B is placed to the right of the placed CTU A. In certain embodiments of the present invention, the predictor candidates are replaced by different predictors. For example, the placed PU can be used for the predictor instead. In other embodiments of the present invention, the use of the relevant predictor mode in the encoded bitstream is not permitted.
[0066] As a variation of the HEVC modification possibilities outlined in Figures 8 and 11, with respect to the replacement of the predictor in Figure 11, the same can be selected so that it becomes the attribute of each of its blocks in the first layer 12, and it will be noted that it has itself located at the same location as the reference position 628 of the current block 302.
[0067] c) Signaling In certain embodiments, as shown in Figures 13a and 13b, for example, the following high-level syntax can be used in a VPS or SPS to enable the constraints / restrictions described above using N flags.
[0068] Here, PREDTYPE, RESTYPE, and SCAL in inter_layer_PREDTYPE_RESTYPE_SCAL_flag_1 through inter_layer_PREDTYPE_RESTYPE_SCAL_flag_N may be replaced with different values as described below: PREDTYPE indicates the prediction type to which restrictions / constraints apply, and may be one of the following or other prediction types not listed: - For example, the prediction of the temporal motion vector of a block placed in the base view from adjacent blocks using temporal_motion_vector_prediction. - For example, disparity_vector_prediction for predicting the disparity vector of a block placed in the base view from adjacent blocks. - For example, depth_map_derivation for predicting depth values from the base view. - For example, inter_view_motion_predition for predicting motion vectors from the base view. - For example, inter_view_residual_prediction for predicting residual data from the base view - For example, inter_view_sample_prediction for predicting sample values from the base view
[0069] Alternatively, the restrictions / constraints may not be explicitly signaled for the prediction types to which they apply, or they may be applied to all prediction types, or the restrictions / constraints may be signaled only for sets of prediction types that use only one flag per set.
[0070] RESTYPE indicates the type of restriction. It may be one of the following: - For example, constraints (indicating bitstream constraints; indicating that flags can be included in the VUI) - For example, a restriction (instructing clipping (a) or selection of a different predictor (b))
[0071] SCAL indicates whether the constraints / restrictions apply only to layers of the same type: - For example, same_scal (indicates that only the limit applies when the base layer has the same scalability type as the augmentation layer) - For example, diff_sca (indicates that the limit applies regardless of the scalability type of the base layer and augmentation layer)
[0072] As an alternative embodiment related to Figure 14, for example, in a VPS or SPS, the use of all described restrictions can be signaled as an ultra-low delay mode in high-level syntax, such as the ultra_low_delay_decoding_mode_flag.
[0073] A value of 1 for `ultra_low_delay_decoding_mode_flag` indicates the use of a decoding process modified at tile boundaries.
[0074] The constraints implied by this flag may also include constraints on tile boundary alignment and constraints on upsampling filters across tile boundaries.
[0075] In other words, referring to Figure 1, assurance signaling can be additionally used to signal an assurance that the image 15 of the second layer will be subdivided such that the boundaries 84 between the spatial segments 82 of the image of the second layer overlay all the boundaries 86 of the spatial segments 80 of the first layer (perhaps after upsampling, if spatial scalability is considered) over a predetermined period, such as a period extending across a sequence of images. The decoder still periodically determines the actual subdivision of the images 12 and 15 of the first and second layers into spatial segments 80 and 82 based on short-term syntax elements of the multi-layer video data stream 40 at time intervals smaller than the predetermined period, for example, on an image pitch interval, such as on an individual image basis, but knowledge of alignment already helps in planning the workload assignment for parallel processing. The solid line 84 in Figure 1 represents, for example, an embodiment where the tile boundary 84 is perfectly spatially aligned with the tile boundary 86 of layer 0. The guarantee mentioned above, however, also allows for a finer-grained tiling of Layer 1 than that of Layer 0, such that the tiling of Layer 1 includes additional tile boundaries that do not spatially overlap any of the tile boundaries 86 of Layer 0. In any case, knowledge of tile registration between Layer 1 and Layer 0 helps the decoder in allocating the workload or processing power available in the spatial segments being processed in parallel and simultaneously. Without a long syntax element structure, the decoder would have to perform workload allocation at smaller time intervals, i.e., per image, thereby consuming computing power to perform the workload allocation. Another aspect is "opportunistic decoding": a decoder with multiple CPU cores can leverage knowledge of layer parallelism to decide whether or not to attempt to decode more complex layers, i.e., layers with a higher number of layers or, in other words, layers with a greater number of views.Bitstreams exceeding the capabilities of a single core can be decoded by utilizing all cores of the decoder. This information is particularly useful when profile and level indicators do not include this type of indication regarding minimum parallelism.
[0076] As previously mentioned, the signaling of the guarantee (exemplary, ultra_low_delay_decoding_mode_flag) can also be used to steer the upsampling filter 36 in the case of multilayer video with a base layer image 12 having a different spatial resolution from the dependent view image 15. If the upsampling filtering is performed across the spatial segment boundary 86 in layer 0, the delay encountered when decoding / encoding the spatial segment 82 of layer 1 in parallel with the encoding / decoding of the spatial segment 80 of layer 0 is increased because the upsampling filtering combines information from adjacent spatial segments of layer 0, thus making them dependent on each other and serving as a predictive reference 38 used in the interlayer prediction of block 41 of layer 1. See Figure 15. Both images 12 and 15 are shown in an overlay method, where they are resized and registered to the required size relative to each other according to their spatial correspondence, i.e., by overlaying parts of each other that represent the same part of the scene. Images 12 and 15 are shown exemplarily divided into tile-like spatial segments 6 and 12, respectively. The filter kernel is illustrated to move across the upper-left tile of image 12 to obtain an upsampled version that serves as a basis for inter-layer predictions of any block within the tile of image 15 that spatially overlays the upper-left tile. In some intermediate instances, such as 202, kernel 200 overlaps with an adjacent tile of image 12. The central sample value of kernel 200 at position 202 of the upsampled version is therefore dependent on both the sample of the upper-left tile of image 12 and the sample of the tile of image 12 to its right. If the upsampled version of image 12 serves as a basis for inter-layer predictions, the inter-layer latency increases in the parallel processing of the layer segments. Limitations can help increase the amount of parallelization across different layers and therefore help reduce the overall coding latency. Of course, syntax elements can also be long-running syntax elements that are valid for sequences of images.The constraint can be achieved in one of the following ways: for example, by loading the overlap portion of kernel 200 at overlap position 202 using the central trend of the sample values within the non-dashed portion of kernel 200, and then extrapolating the non-dashed portion to the dashed portion, etc., using a linear or other function.
[0077] An alternative embodiment is given as an example in a VPS as follows: The aforementioned restrictions / constraints are controlled by ultra_low_delay_decoding_mode_flag, but alternatively (when the flag is unavailable) each restriction / constraint can be enabled individually. Figures 13c and 13d are referenced for this embodiment. This embodiment can also be included in other non-VCLNAL units (e.g., SPS or PPS). In Figures 13c and 13d,
[0078] If ultra_low_delay_decoding_mode_flag is equal to 1, it is assumed that du_interleaving_enabled_flag, interlayer_tile_mv_clipping_flag, depth_disparity_tile_mv_clipping_flag, inter_layer_tile_tmvp_restriction_flag, and independent_tile_upsampling_idc are equal to 1, and that they do not exist in VPS, SPS, or PPS.
[0079] When parallelization techniques such as tiling are used in layered encoded video sequences, controlling the limitations of encoding tools, such as inter-view prediction, in HEVC extensions is beneficial from a latency perspective, in order to avoid crossing tile boundaries in a unified manner.
[0080] In the embodiment, the value of independent_tiles_flag determines the presence of syntax elements that control individual restrictions / constraints, such as inter_layer_PREDTYPE_RESTYPE_SCAL_flag_x or independent_tile_upsampling_idc. independent_tiles_flag can be included in the VPS as illustrated in Figure 13e. Here
[0081] An independent_tiles_flag equal to 1 is assumed to be inter_layer_PREDTYPE_RESTYPE_SCAL_flag_1 ~ inter_layer_PREDTYPE_RESTYPE_SCAL_flag_N, and independent_tile_upsampling_idc is assumed to be equal to 1, identifying that they do not exist in VPS, SPS, or PPS.
[0082] An alternative embodiment, shown in Figure 13f, is that in a VPS, the constraints described above are controlled by the independent_tiles_flag, but alternatively (when the flag is enabled), each constraint can be enabled individually. This embodiment can also be included in other non-VCLNAL units (e.g., SPS or PPS), as shown in Figure 13g.
[0083] With respect to Figures 8 to 15, the aforementioned embodiments can be summarized, and the signaling of guarantees in the data stream can be used by decoder 620 to optimize the inter-layer decoding offset between the decoding of different layers / views 12 and 15, or the guarantees can be used by decoder 620 to suppress or allow inter-layer parallel processing attempts as described above by referring to “opportunistic decoding”.
[0084] The next aspect of the present application discussed relates to the problem of enabling lower end-to-end latency in multilayer video coding. It is worth noting that the aspects described below can be combined with the aspects described above, and vice versa, and embodiments relating to the aspects described herein can also be implemented without the details mentioned above. In this regard, it should also be noted that the embodiments described hereafter are not limited to multiview coding. The multiple layers referred to below in the second aspect of the present application may contain different views, but may also represent the same view in terms of varying spatial resolution, SNR accuracy, etc. The possible scalability dimensions by which the multiple layers discussed below increase the information content transmitted by the previous layers are diverse, including, for example, the number of views, spatial resolution, and SNR accuracy, and further possibilities will become apparent from discussing the third and fourth aspects of the present application, which may also be combined with the aspects currently described according to their embodiments.
[0085] The second aspect of the present invention described herein relates to the problem of actually achieving low coding delay, i.e., embedding the idea of low latency into the framework of NAL units. As described above, NAL units consist of slices. The tile and / or WPP concept is individually and freely selected for different layers of a multilayer video data stream. Thus, each NAL unit having slices packetized into it can be spatially attributed to the area of the image referenced by each slice. Thus, to enable low-latency coding in the case of interlayer prediction, it is advantageous to allow interleaving NAL units of different layers that accompany at the same time instant, and to allow encoders and decoders to initiate coding and transmission, and to allow the slices packetized into these NAL units to be decoded in a manner that accompanies at the same time constant, enabling parallel processing of these images of different layers. However, depending on the application, an encoder may prefer the ability to use different coding orders in images of different layers, such as the use of different GOP structures for different layers, beyond the ability to enable parallel processing at the layer dimension. Therefore, according to the second aspect, the structure of the data stream can be described again below with respect to Figure 16.
[0086] Figure 16 shows a multilayer video material 201 consisting of a sequence of images 204 for each of the different layers. Each layer can describe a different property of this scene described by the multilayer video material 201. That is, the meaning of the layers can be selected from the following: for example, color components, depth image, transparency, and / or viewpoint. Without loss of generality, we assume that the different layers correspond to different views by the video material 201, which is a multiview video.
[0087] In the case of applications requiring low latency, the encoder may decide to signal long-term high-level syntax elements (by setting du_interleaving_enabled_flag, introduced below, to equal 1). In this case, the data stream generated by the encoder can be seen as indicated in the center of Figure 16 by the circle surrounding it. In this case, the multilayer video stream 200 consists of a sequence of NAL units 202 such that each NAL unit 202 belonging to one access unit 206 relates to an image at one time instant, and NAL units 202 of different access units relate to different time instants. Within each access unit 206, for each layer, at least several NAL units related to each layer are grouped into one or more decoding units 208. This means that within the NAL units 202, there are different types of NAL units, such as VCLNAL units on the one hand and non-VCLNAL units on the other, as indicated above. More specifically, NAL units 202 can have different types, and these types can include:
[0088] 1) A NAL unit that carries residual data describing the image content in terms of the scale / granularity of the image samples and syntax elements related to prediction parameters, such as slices, tiles, WPP substreams, etc. One or more of this type may exist. The VCLNAL unit is one of this type. This type of NAL unit is not detachable.
[0089] 2) Parameter set NAL units can carry information that changes infrequently, such as long-term encoding settings, some examples of which are described above. This type of NAL unit can be scattered to some extent and repeatedly within a data stream, for example;
[0090] 3) The Auxiliary Enhancement Information (SEI) NAL unit can carry optional data.
[0091] A decoding unit can consist of the first of the NAL units described above. More precisely, a decoding unit can consist of "one or more VCLNAL units and associated non-VCLNAL units in the access unit." A decoding unit, therefore, describes a particular area of an image, i.e., an area to be encoded into one or more slices contained therein.
[0092] The decoding units 208 of NAL units relating to different layers are interleaved so that, for each decoding unit, the inter-layer prediction used to encode each decoding unit is interleaved based on the portion of the image in layers other than the layer relating to each decoding unit, and that portion is encoded in the decoding unit preceding each decoding unit within each access unit. For example, see decoding unit 208a in Figure 16. Imagine this decoding unit relating, exemplarily, to a dependent layer 2 and area 210 of each image for a particular time instant. An area located at the same location in the base layer image for the same time instant is shown by 212, and the area of this base layer image slightly exceeding this area 212 may be required to fully decode decoding unit 208a by utilizing the inter-layer prediction. The slight excess may be, for example, the result of a disparity-corrected prediction. This means, in turn, that decoding unit 208b preceding decoding unit 208a within access unit 206 must fully cover the area required for the inter-layer prediction. The above description is relevant to the use of delay indicators as boundaries for interleaving granularity.
[0093] However, if the application wants to take advantage of more freedom to choose different decoding orders for images in different layers, in this case, represented by 2 by the circle surrounding it at the bottom of Figure 16, the encoder can choose to set du_interleaving_enabled_flag to equal 0. In this case, the multilayer video data stream has individual access units for each image that belongs to a particular pair of layer IDs and one or more values of a single time instant. As shown in Figure 16, in the (i-1)th decoding order, i.e., time instant t(i-1), each layer can consist of access units AU1, AU2 (etc.), or all layers can be configured to be contained within a single access unit AU1. However, in this case, interleaving is not permitted. The access units are arranged in the data stream 200 according to the decoding order index i, i.e., an access unit for decoding order index i for each layer, followed by access units for the images of those layers corresponding to decoding order i+1. Signaling for temporal inter-image predictions in a data stream signals whether to apply the same coding order to different layers or to apply different image coding orders, and the signaling can be redundantly placed in one or more locations within the data stream, for example, within slices that are packetized into NAL units.
[0094] With respect to NAL unit types, it is noteworthy that the ordering rules defined within them allow the decoder to determine where the boundaries between consecutive access units are located, regardless of whether NAL units of removable packet types are removed during transmission. NAL units of removable packet types may include, for example, SEINAL units, or NAL units of redundant image data, or other specific NAL unit types. That is, the boundaries between access units remain stationary, and the ordering rules are observed within each access unit, but broken at each boundary between two access units.
[0095] For completeness, Figure 17 illustrates the case where du_interleaving_flag = 1, which allows packets to belong to different layers, but where, for example, the same time instant t(i-1) is distributed within a single access unit. The case where du_interleaving_flag = 0 is represented by a circle around it in Figure 16, with a value of 2.
[0096] However, with respect to Figures 16 and 17, it should be noted that the interleaved signaling described above can necessarily result in a multilayer video data stream using access unit definitions in accordance with the case indicated by the circle surrounding it in Figures 16 and 17.
[0097] According to the embodiment, whether or not the NAL units contained within each access unit are actually interleaved can be determined at the encoder's discretion with respect to their relationship to the layers of the data stream. To facilitate handling of the data stream, a syntax element such as du_interleaving_flag can signal to the decoder whether NAL units are interleaved or not within an access unit that aggregates all NAL units for a particular timestamp, so that the decoder can process the NAL units more easily. For example, whenever interleaving is signaled to be switched on, the decoder can use multiple encoded image buffers, as briefly illustrated with respect to Figure 18.
[0098] Figure 18 shows a decoder 700 that can be implemented as outlined above with respect to Figure 2, and can even correspond to the proposed description with respect to Figure 9. Exemplaryly, the multilayer video data stream of Figure 17 is indicated as Option 1 as the entering decoder 700 by the circle surrounding it. To more easily perform deinterleaving of NAL units belonging to different layers in a common time instant per access unit AU, the decoder 700 uses two buffers 702, 704 for each access unit AU, each having a multiplexer 706 that forwards, for example, the NAL units of that access unit AU belonging to the first layer to buffer 702, and for example, the NAL units belonging to the second layer to buffer 704. The decoding unit 708 then performs decoding. For example, in Figure 18, for example, NAL units belonging to the base layer / first layer are hatched and not shown, while NAL units of the dependent layer / second layer are shown using hatching. If the interleaving signaling outlined above is present in the data stream, the decoder 700 can respond to this interleaving signaling in the following ways: if the interleaving signaling signals that the interleaving of NAL units is switched on, i.e., NAL units of different layers are interleaved with each other within one access unit AU, the decoder 700 uses buffers 702 and 704, which have a multiplexer 706 that distributes the NAL units on these buffers, as outlined above. Otherwise, however, the decoder 700 simply uses one of buffers 702 and 704, for example buffer 702, for all NAL units provided in the access unit.
[0099] To better understand the embodiment of Figure 18, Figure 18 is referenced together with Figure 19, which shows an encoder configured to generate a multilayer video data stream as described above. The encoder in Figure 9 is generally indicated by reference numeral 720 and, for ease of understanding, exemplary, encodes input images of two layers, one of which is layer 12 forming the base layer and the other is layer 15 forming the dependent layer. They can form different views, as outlined earlier. A typical encoding order in which encoder 720 encodes images of layers 12 and 15 scans the images of these layers substantially in chronological order (presentation time), while encoding order 722 can deviate from the presentation time order of images 12 and 15 in units of groups of images. At each chronological time instant, encoding order 722 passes images of layers 12 and 15 in chronological order, i.e., from layer 12 to layer 15.
[0100] The encoder 720 encodes the images of layers 12 and 15 into the data stream 40, using the aforementioned NAL units, each spatially associated with a portion of the respective image. Thus, the NAL units belonging to a particular image spatially subdivide or divide the respective image, and as already described, the inter-layer prediction places the portion of the layer 15 image in substantially the same location as the respective portion of the layer 15 image, and subordinates it to the time-matched portion of the layer 12 image, which substantially includes disparity displacement. In the embodiment of Figure 19, the encoder 720 chooses to utilize the possibility of interleaving in the formation of an access unit that aggregates all NAL units belonging to a particular time instant. In Figure 19, the portion from the illustrated data stream 40 corresponds to the input to the decoder in Figure 18. That is, in the embodiment of Figure 19, the encoder 720 uses inter-layer parallel processing in encoding layers 12 and 15. As far as time instant t(i-1) is concerned, the encoder 720 begins encoding the layer 15 image as soon as the NAL unit 1 of the layer 12 image is encoded. Each NAL unit, once its encoding is complete, is output by encoder 720, and the encoder 720 provides an arrival timestamp corresponding to the time each NAL unit was output. After encoding the first NAL unit of the layer 12 image at time instant t(i-1), encoder 720 continues encoding the content of the layer 12 image, outputting a second NAL unit of the layer 12 image, and providing an arrival timestamp that follows the arrival timestamp of the first NAL unit of the time-matched layer 15 image. In other words, encoder 720 outputs the NAL units of the layer 12 and 15 images, all belonging to the same time instant, in an interleaved manner, and the NAL units of the data stream 40 are actually transmitted in this interleaved manner. The circumstances under which encoder 720 chooses to utilize the possibility of interleaving are indicated by the encoder 720, by the signaling 724 for each interleave within the data stream 40.The encoder 720 can output the first NAL unit of the dependent layer 15 at time instant t(i-1) earlier compared to a non-interleaved scenario where the output of the first NAL unit of layer 15 is delayed until the encoding and output of all NAL units of the time-matched base layer image is complete, thus reducing the end-to-end delay between the decoder (Figure 18) and the encoder (Figure 19).
[0101] As already mentioned above, according to the alternative embodiment, in the non-interleaved case, i.e., in the case of signaling 724 indicating the non-interleaved alternative, the definition of the access unit remains the same, that is, the access unit AU can collect all NAL units belonging to a particular time instant. In that case, signaling 724 simply indicates within each access unit whether or not NAL units belonging to different layers 12 and 15 are interleaved.
[0102] As described above, according to signaling 724, the decoding in Figure 18 uses either one or two buffers. In the case where interleaving is switched on, the decoder 700 distributes the NAL units on the two buffers 702 and 704, for example, so that the NAL units of layer 12 are buffered in buffer 702, while the NAL units of layer 15 are buffered in buffer 704. Buffers 702 and 704 are emptied on an access unit basis. This is correct for both signaling 724, which indicate interleaving or non-interleaving.
[0103] If encoder 720 sets the removal time within each NAL unit, it is preferable that decoding unit 708 utilizes the possibility of decoding layers 12 and 15 from data stream 40 using inter-layer parallel processing. Even if decoder 700 does not apply inter-layer parallel processing, end-to-end delay is, however, already reduced.
[0104] As already mentioned above, NAL units can have different NAL unit types. Each NAL unit may have a NAL unit type index that indicates the type of each NAL unit from a set of possible types. Within each access unit, the NAL unit types of each access unit can adhere to the ordering rules within the NAL unit types, but the ordering rules are broken between simply two consecutive access units. Therefore, the decoder 700 can identify access unit boundaries by examining this rule. For more detailed information, refer to the H.264 standard.
[0105] With respect to Figures 18 and 19, a decoding unit DU can be identified as a continuous run of NAL units within a single access unit belonging to the same layer. In access unit AU(i-1) in Figure 19, the NAL units designated "3" and "4" form, for example, one DU. All other decoding units in access unit AU(i-1) consist of simply one NAL unit. Together, access unit AU(i-1) in Figure 19 comprises, exemplary, six decoding unit DUs arranged as alternatives within access unit AU(i-1), i.e., they consist of a run of NAL units in one layer, alternating between layer 1 and layer 0.
[0106] Similar to the first aspect, the second aspect previously described will be outlined below with respect to how it can be incorporated into the HEVC extension.
[0107] Prior to this, however, for completeness, a further aspect of the current HEVC is described, which enables parallel processing between images, i.e., WPP processing.
[0108] Figure 20 describes how WPP is currently implemented in HEVC. That is, this description forms the basis for the optional implementation of WPP processing in any of the embodiments described above or below.
[0109] In the base layer, wavefront parallel processing enables parallel processing of coding tree block (CTB) rows. Predictive dependencies are not destroyed when traversing CTB rows. Regarding entropy coding, as can be seen in Figure 20, WPP modifies the CABAC dependency for the upper-left CTB in each upper CTB row. Once the entropy decoding of the corresponding upper-right CTB is completed, entropy coding of the CTB in the following row can begin.
[0110] In the enhancement layer, as soon as the CTB containing the corresponding image area is fully decoded and available, decoding of the CTB can begin.
[0111] In HEVC and its extensions, the following definition of a decoding unit is given:
[0112] Decoding unit: An access unit when SubPicHrdFlag is equal to 0, or a subset of an access unit consisting of one or more VCLNAL units and associated non-VCLNAL units within the access unit, etc.
[0113] In HEVC, preferably, if external means and lower-level image HRD parameters are available, the HRD (Hypothetical Reference Decoder) can optionally activate CPB and DPB at the decoding unit level (or sub-picture level).
[0114] The HEVC specification [1] describes the concept of a so-called decoding unit, as defined below:
[0115] 3.1 Decoding Unit: An access unit when SubPicHrdFlag is equal to 0, or a subset of an access unit consisting of one or more VCLNAL units and associated non-VCLNAL units within the access unit, etc.
[0116] In layered encoded video sequences, such as those provided in HEVC extensions for 3D[3], multiview[2], or spatial scalability[4] (where additional representations of video data, e.g., high fidelity, spatial resolution, or different camera views, depend on lower layers and are encoded through predictive interlayer / interview coding tools), interleaving related or similarly located decoding units of related or similarly located layers of related layers (in an image areawise manner) or related in the bitstream is beneficial for minimizing end-to-end latency on encoders and decoders.
[0117] In order to enable interleaving of decoding units in the encoded video bitstream, specific constraints on the encoded video bitstream must be signaled and enforced.
[0118] The methods by which the above interleaving concept can be implemented in HEVC are described in detail and demonstrated in the following subsections.
[0119] The current state of the HEVC extension, as far as can be taken from the draft document of the MV-HEVC specification [2], includes a definition for the access unit used, to which the access unit constrains one encoded image (by a specific value nuh_layer_id). One encoded image is defined below and is essentially the same as the view component in MVC. It was an unresolved question whether the access unit should be defined so that it contains all view components by the same POC value.
[0120] The base HEVC specification [1] defines it as follows:
[0121] 3.1 Access Units: A set of NAL units is related to one another according to the decoding order rules, is contiguous in the decoding order, and contains exactly one encoded image.
[0122] Note 1 - In addition to containing VCL NAL units of the encoded image, an access unit may also contain non-VCL NAL units. Decoding of an access unit always results in a decoded image.
[0123] The definition of an Access Unit (AU), which allows only one encoded image in each Access Unit, appears to be interpreted in a way that requires each dependent view to be interpreted as a separate encoded image and contained in a separate Access Unit. This is represented by "2" in Figure 17.
[0124] In the previous standard, an "encoded image" included all layers of the view representation of an image at a specific timestamp.
[0125] Access units cannot be interleaved. This means that if each view is contained within a different access unit, all images of the base view must be received in the DPB before the first decoding unit (DU) of the dependent images can decode them.
[0126] Interleaving the decoding units is advantageous for ultra-low latency operation using dependent layers / views.
[0127] The embodiment in Figure 21 includes three views, each from three decoding units. They are received in the order from left to right: If each view is included in its own access unit, the minimum delay for decoding the first decoding unit of view 3 is fully included in the received views 1 and 2.
[0128] If views can be interleaved and transmitted, the minimum delay can be reduced, as shown in Figure 22 and as already explained with respect to Figures 18 and 19.
[0129] Interleaving NAL units from different layers in HEVC's scalable extension can be achieved as follows:
[0130] A bitstream interleaving mechanism for layer or view representations, and a decoder that can use this bitstream layout to achieve decoding of dependent views with very low latency using parallelization techniques. Interleaving of DU is controlled via a flag (e.g., du_interleaving_enabled_flag).
[0131] In the scalable extension of HEVC, interleaving of NAL units from different layers of the same AU is necessary to enable low-latency decoding and parallelization. Therefore, a definition can be introduced that follows this:
[0132] Access Unit: A set of NAL units, related to each other according to specific identification rules, is sequential in the decoding order and contains exactly one encoded image.
[0133] Encoded layer image component: The encoded representation of the layer image component, including all the encoded tree units of the layer image component.
[0134] Encoded image: An encoded representation of an image that includes all the encoding tree units of an image containing one or more encoded layer image components.
[0135] Image: An image is a set of one or more layered image components.
[0136] Layer image component: An array of luma samples in a monochrome format and two corresponding arrays of chroma samples in 4:2:0, 4:2:2, and 4:4:4 color formats, where the encoded representation consists of NAL units from a specific layer among all NAL units in the access unit.
[0137] NAL units are interleaved according to their dependencies in such a way that each NAL unit can be decoded using data received by the previous NAL unit in the decoding order, i.e., no data is needed from subsequent NAL units in the decoding order for each NAL unit to be decoded (du_interleaving_enabled_flag == 1).
[0138] When DU interleaving is applied (du_interleaving_enabled_flag == 1), and the luminance and chroma components are separated into different color planes, each NAL unit associated with a color plane can be interleaved. Each of these NAL units (related to the unique value of colour_plane_id) must adhere to the VCLNAL unit order, as described below. Since the color planes are expected to have no coding dependencies between them in the access units, they follow the normal order.
[0139] The ordering constraints of NAL units can be expressed using the syntax element min_spatial_segment_delay, which measures and guarantees the worst-case delay / offset between spatial segments in units of CTB. Syntax elements describe spatial region dependencies between spatial segments of the CTB or base and augmentation layers (e.g., tiles, slices, or CTB rows for WPP). Syntax elements are not required for interleaving NAL units or for sequential decoding of NAL units in coding order. Parallel multilayer decoders can use syntax elements to prepare for parallel decoding of layers.
[0140] The following constraints affect the possibility of encoders enabling interleaving of parallelization and decoding units across layers / views, as primarily described with respect to the first aspect:
[0141] 1) Prediction of samples and syntax elements: Interpolation filters for luminance and chroma resampling set constraints on the lower layers to the necessary data in order to generate the upsampled data required for the upper layers. For example, since spatial segments of an image can be upsampled independently, decoding dependencies can be reduced by constraining these filters; signaling of specific constraints for tiling is described above with respect to the first form.
[0142] Motion vector prediction, and more specifically temporal motion vector prediction, for "reference index-based scalable extension" (HLS approach) sets constraints on the required data at lower layers and generates the motion field of the required resampled image. Related inventions and signaling are described above with respect to the first embodiment.
[0143] 2) Motion vector: For SHVC, motion compensation is not used depending on the underlying layer; that is, if the underlying layer is used as a reference image (HLS approach), the resulting motion vector must be a zero vector. However, for MV-HEVC0 or 3D-HEVC0, the disparity vector can be constrained, but it is not necessarily a zero vector. That is, motion compensation can be used for interview prediction. Therefore, constraints on the motion vector can be applied to ensure that only the data received in the previous NAL unit is needed for decoding. Related inventions and signaling are described above with respect to the first embodiment.
[0144] 3) Image segmentation by tile boundaries: When parallel processing and low latency are efficiently desired through interleaving NAL units from different layers, image segmentation in the augmentation layer must be subordinate to image segmentation in the reference layer.
[0145] As far as the order of VCL NAL units and their relationship to encoded images are concerned, the following can be identified:
[0146] Each VCL NAL unit is part of the encoded image.
[0147] The order of VCLNAL units within the encoded layer image component of an encoded image, i.e., VCLNAL units of an encoded image with the same layer_id_nuh value, is constrained as follows:
[0148] - The first VCLNAL unit of the encoded layer image component sets first_slice_segment_in_pic_flag to equal 1.
[0149] - Set sliceSegAddrA and sliceSegAddrB to the slice_segment_address values of two encoded slice segments NAL units A and B within the same encoded layer image component. Encoded slice segment NAL unit A precedes encoded slice segment NAL unit B when any of the following conditions are true:
[0150] - TileId[ CtbAddrRsToTs[ sliceSegAddrA ] ] is less than TileId[ CtbAddrRsToTs[ sliceSegAddrB ] ].
[0151] - TileId[ CtbAddrRsToTs[ sliceSegAddrA ] ] is equal to TileId[ CtbAddrRsToTs[ sliceSegAddrB ] ], and CtbAddrRsToTs[ sliceSegAddrA ] is less than CtbAddrRsToTs[ sliceSegAddrB ].
[0152] When an encoded image consists of multiple layered image components, the order of the VCLNAL units of all image components is constrained as follows:
[0153] - In the encoded layer image component layerPicA used as a reference to another layer image component layerPicB, VCLNALA is defined as the first VCLNAL unit A. In this case, VCLNAL unit A precedes any VCLNAL unit B belonging to layerPicB.
[0154] - Otherwise (not the first VCLNAL unit), if du_interleaving_enabled_flag is equal to 0, VCLNALA is set to one of the VCLNAL units of the encoded layer image component layerPicA used as a reference to another encoded layer image component layerPicB. In this case, VCLNAL unit A precedes any VCLNAL unit B belonging to layerPicB.
[0155] - Otherwise (not the first VCLNAL unit, and du_interleaving_enabled_flag is equal to 1), if ctb_based_delay_enabled_flag is equal to 1 (i.e., CTB-based delay is signaled regardless of whether tiling or WPP is used in the video sequence), then layerPicA is the encoded layer image component used as a reference to the other encoded layer image component layerPicB. Also, NALUsetA is the sequence of continuous slice segment NAL units belonging to layerPicA that directly follows the sequence of continuous slice segment NAL units belonging to layerPicA and layerPicB, and NALUsetB1 and NALUsetB2 are the sequences of continuous slice segment NAL units belonging to layerPicB that directly follow NALUsetA. Let sliceSegAddrA be the slice_segment_address of the first segment NAL unit of NALUsetA, and sliceSegAddrB be the slice_segment_address of the first encoded slice segment NAL unit of NALUsetB2. Then the following condition is true:
[0156] - If NALUsetA exists, then NALUsetB2 also exists.
[0157] - CtbAddrRsToTs[PicWidthInCtbsYA * CtbRowBA(sliceSegAddrB-1) + CtbColBA(sliceSegAddrB-1) + min_spatial_segment_delay] is less than or equal to CtbAddrRsToTs[sliceSegAddrA-1]. See also Figure 23.
[0158] Otherwise (not the first VCLNAL unit, du_interleaving_enabled_flag is equal to 1, and ctb_based_delay_enabled_flag is equal to 0), if tiles_enabled_flag is equal to 0 and entropy_coding_sync_enabled_flag is equal to 0 (i.e., neither tiles nor WPP are used in the video sequence), then layerPicA is the encoded layer image component used as a reference for the other encoded layer image component layerPicB. Also, VCL NAL unit B is one of the VCL NAL units of the encoded layer image component layerPicB, and VCL NAL unit A is the VCLNAL unit immediately preceding layerPicA by a slice_segment_address value equal to sliceSegAddrA, where there is a VCL NAL unit (min_spatial_segment_delay - 1) from layerPicA between VCL NAL unit A and VCL NAL unit B. Furthermore, VCL NAL unit C is defined as the next VCL NAL unit of the encoded layer image component layerPicB that follows VCL NAL unit B by a slice_segment_address value equal to sliceSegAddrC. PicWidthInCtbsYA is defined as the image width in units of CTB oflayerPicA. Then the following condition is true:
[0159] - A VCLNAL unit with min_spatial_segment_delay always exists in layerPicA, which precedes VCL NAL unit B.
[0160] - PicWidthInCtbsYA * CtbRowBA(sliceSegAddrC-1) + CtbColBA(sliceSegAddrC-1) is less than or equal to sliceSegAddrA-1.
[0161] - Otherwise (not the first VCLNAL unit, du_interleaving_enabled_flag is equal to 1, and ctb_based_delay_enabled_flag is equal to 0), if tiles_enabled_flag is equal to 0 and entropy_coding_sync_enabled_flag is equal to 1 (i.e., WPP is used in the video sequence), then sliceSegAddrA is set to the slice_segment_address of any of the encoded layer image components of layerPicA that directly precede slice segment VCL NAL unit B, by a slice_segment_address equal to sliceSegAddrB belonging to the encoded layer image component layerPicB that uses layerPicA as a reference. Also, PicWidthInCtbsYA is the image width of layerPicA in units of CTB. Then the following condition is true:
[0162] - ( CtbRowBA(sliceSegAddrB) - Floor( (sliceSegAddrA) / PicWidthInCtbsYA) + 1) is equal to or greater than min_spatial_segment_delay.
[0163] Otherwise (not the first VCLNAL unit, du_interleaving_enabled_flag is equal to 1, and ctb_based_delay_enabled_flag is equal to 0), if tiles_enabled_flag is equal to 1 and entropy_coding_sync_enabled_flag is equal to 0 (i.e., tiles are used in the video sequence), then sliceSegAddrA is the slice_segment_address of any slice segment VCLNAL unit A of the encoded layer image component layerPicA, and slice segment VCLNAL unit B is the first next VCLNAL unit belonging to the encoded layer image component layerPicB, which uses layerPicA as a reference by a slice_segment_address equal to sliceSegAddrB. Also, PicWidthInCtbsYA is the image width of layerPicA in units of CTB. Then the following condition is true:
[0164] - TileId[ CtbAddrRsToTs[ PicWidthInCtbsYA * CtbRowBA(sliceSegAddrB-1) + CtbColBA(sliceSegAddrB-1) ] ] - TileId[ CtbAddrRsToTs[ sliceSegAddrA-1] ] is equal to or greater than min_spatial_segment_delay.
[0165] The signaling 724 can be placed within the VPS as shown in Figure 24, where:
[0166] When du_interleaving_enabled_flag is equal to 1, du_interleaving_enabled_flag specifies that frames with a single associated encoded image (i.e., a single associated AU) consisting of all encoded layer image components can be interleaved for frames and VCLNAL units corresponding to different layers. When u_interleaving_enabled_flag is equal to 0, frames that may have multiple associated encoded images (i.e., one or more associated AUs) and VCLNAL units of different encoded layer image components are not interleaved.
[0167] To complete the above discussion, the virtual reference decoder associated with decoder 700 may, in the alignment according to the embodiment in Figure 18, operate with one or two buffers, buffers 702 and 704, according to the setting of signaling 724, i.e., it may switch between these options according to signaling 724.
[0168] Other aspects of the present application are described below, which can again be combined with Form 1, Form 2, or both. A third aspect of the present application relates to an extension of scalability signaling for many applications, such as views.
[0169] For the sake of understanding the proposed description below, an overview of existing scalability signaling concepts is provided.
[0170] Most 3D video applications or deployments of a certain level of technology feature stereoscopic content with or without distance images for each of the two camera views or more views (>2) of multiview content.
[0171] The High Efficiency Video Coding (HEVC) standard [1] and its extensions to 3D and multiview video [2][3] feature scalability signaling for network abstraction layers (NALs) that can represent up to 64 different layers, each with a 6-bit layer identifier (uh_layer_id) in the header of each NAL unit given in the syntax table in Figure 25.
[0172] Each value of a layer identifier can be converted to a set of scalable identifier variables (e.g., DependencyID, ViewID, etc.) according to the scalability dimension in use, for example through video parameter set extensions, which allows up to 64 dedicated views to be directed on a NAL layer or 32 dedicated views when the layer identifier is also used to indicate a depth image.
[0173] However, there are applications that require a substantially larger number of views to be encoded, transmitted, decoded, and displayed in video bitstreams, such as in multi-camera arrays with a large number of cameras, as provided in references [5][6][7], or in holographic displays that require a large number of viewpoints. The following sections describe two inventions that address the aforementioned shortcomings of the HEVC high-level syntax for extensions.
[0174] Simply extending the size nuh_layer_id field in the NAL unit header does not seem like a useful solution to the problem. The header is expected to be fixed-length, which is required in very simple (low-cost) devices that perform operations on the bitstream, such as routines and extractions, for easy access. This means that for all cases, even if very few views are used, the extra bits (or bytes) must be appended.
[0175] Furthermore, after the final version of the standard is completed, it is no longer possible to modify the NAL unit header.
[0176] The following description outlines an extension mechanism for an HEVC decoder or intermediate device that extends scalability signaling capabilities to meet the above-mentioned requirements. Activation and extension data can be signaled in HEVC high-level syntax.
[0177] The following describes, in particular, the signaling that indicates the availability of the layer identifier extension mechanism (as described in the next section) in the video bitstream.
[0178] A possible implementation of a third concept in the HEVC framework is first described by describing generalizations of the following embodiments, in addition to those relating to the first and second embodiments. The concept allows for multiple view components with the same existing layer identifier (nuh_layer_id) within the same access unit. Additional identifier extensions are used to distinguish between these view components. These extensions are not encoded in the NAL unit header. Thus, while it is not as easily accessible as in the NAL unit header, it still enables new use cases with more views. In particular, for view clustering (see explanation below), the old extraction mechanism can still be used without any modification to extract groups of views that belong together.
[0179] To extend the existing range of layer identifier values, the present invention describes the following mechanism:
[0180] a. A predetermined value of an existing layer identifier is used as a special value (a so-called "escape code") to indicate that the actual value is determined using an alternative derivation process (in a particular embodiment): the value of the syntax element nuh_layer_id in the NAL unit header (e.g., the highest value of the layer identifier) is used. b. Flags, indices, or bit length indications in high-level syntax structures (for example, in slice header syntax or in extensions of video / sequence / image parameter sets, as given in the following embodiments of the present invention) enable combinations of values for existing layer identifier values in other syntax structures.
[0181] The activation of the expansion mechanism can be carried out as follows:
[0182] For a), explicit activation signaling is not required; that is, reserved escape codes can always be used to signal the use of extension (a1). However, this reduces the number of possible layers / views that do not use extension by 1 (the value of the escape code). Thus, the following switching parameters can be used for both variations (a2).
[0183] The extension mechanism can be enabled or disabled within the bitstream using one or more syntax elements that persist throughout the entire bitstream of a video sequence, the video sequence, or a portion of a video sequence.
[0184] A particular embodiment of the present invention that enables the extension by using a variable LayerId representing an existing layer identifier is: Modification I) Modification I is shown in Figure 26. Here,
[0185] The `layer_id_ext_flag` parameter allows the use of additional LayerId values.
[0186] Modification II) Modification II is illustrated in Figure 27. Here,
[0187] A `layer_id_mode_idc` equal to 1 indicates that the range of LayerId values is extended by using escape codes. A `layer_id_mode_idc` equal to 2 indicates that the range of LayerId values is extended by an offset value. A `layer_id_mode_idc` equal to 0 indicates that no extension mechanism is used for LayerId.
[0188] Note: Different values can be assigned to the mode.
[0189] Modification III) Modification III is illustrated in Figure 28. Here,
[0190] `layer_id_ext_len` indicates the number of bits used to extend the LayerId range.
[0191] The syntax elements described above serve as indicators for the use of the layer identifier extension mechanism to indicate the layer identifier of the corresponding NAL unit or slice data.
[0192] In the following description, the variable LayerIdExtEnabled is used as a logical indicator to show that the extension mechanism was available. The variable is used in the description for simpler references. Examples of variable names and embodiments of the invention may directly use different names or corresponding syntax elements. The variable LayerIdExtEnabled is derived according to the above case as follows:
[0193] For a1), if only predetermined values for the layer identifier syntax element are used to enable the layer identification extension mechanism, the following applies:
[0194] if ( nuh_layer_id == predetermined value ) LayerIdExtEnabled = true else LayerIdExtEnabled = false
[0195] In the case of Modification I, for cases a2) and b), i.e., when a flag (e.g., layer_id_ext_enable_flag) is used to enable the layer identifier extension mechanism, the following applies:
[0196] LayerIdExtEnabled = layer_id_ext_enable_flag
[0197] In the case of Variation II, for cases a2) and b), i.e., when an index (e.g., layer_id_mode_idc) is used to enable the layer identifier extension mechanism, the following applies:
[0198] if ( layer_id_mode_idc == predetermined value ) LayerIdExtEnabled = true else LayerIdExtEnabled = false
[0199] In the case of Variation III, for cases a2) and b), i.e., when a bit length instruction (e.g., layer_id_ext_len) is used to enable the layer identifier extension mechanism, the following applies:
[0200] if (layer_id_ext_len > 0) LayerIdExtEnabled = true else LayerIdExtEnabled = false
[0201] For case a2), if a predetermined value is used in combination with the syntax element to be enabled, the following applies:
[0202] LayerIdExtEnabled &= ( nuh_layer_id == predetermined value )
[0203] Layer identifier extensions can be signaled as follows:
[0204] If an extension mechanism is available (for example, through signaling as described in the previous section), a predetermined or signaled number of bits (layer_id_ext_len) is used to determine the actual LayerId value. For VCLNAL units, additional bits may be included in the slice header syntax (for example, by using existing extensions) or in the SEI message used to extend the signaling range of the layer identifier in the NAL unit header, either by its position in the video bitstream or by an index associated with the corresponding slice data.
[0205] For non-VCL NAL units (VPS, SPS, PPS, SEI messages), additional identifiers may be added for specific extensions or by associated SEI messages. In further description, certain syntax elements are referred to as `layer_id_ext`, regardless of their position in the bitstream syntax. The name is used as an example. The following syntax table and semantics provide examples of possible embodiments.
[0206] Signaling for layer identifier extensions in slice headers is illustrated in Figure 29.
[0207] Alternative signaling for layer identifier extension in slice header extension is shown in Figure 30.
[0208] An example of signaling for the video parameter set (VPS) is shown in Figure 31.
[0209] Similar extensions exist for SPS, PPS, and SEI messages. Additional syntax elements can be added to these extensions in a similar manner.
[0210] The signaling of layer identifiers in related SEI messages (e.g., layer ID extended SEI messages) is illustrated in Figure 32.
[0211] The range of an SEI message can be determined based on its position in the bitstream. In a particular embodiment of the present invention, all NAL units are received between the time a layer ID extended SEI message is associated with the value of layer_id_ext and the start of a new access unit or a new layer ID extended SEI message.
[0212] Dependent on that position, additional syntax elements can be encoded by fixed (denoted here as u(v)) or variable-length (ue(v)) codes.
[0213] The layer identifier for a particular NAL unit and / or slice data is derived from mathematically combined information provided by the layer identifier (nuh_layer_id) in the NAL unit header and the layer identifier extension mechanism (layer_id_ext) which is activated according to the activation of the layer identifier extension mechanism (LayerIdExtEnabled).
[0214] In a particular embodiment, the layer identifier referred to here as LayerId is derived using the existing layer identifier (nuh_layer_id) as the most significant bit and the extended information as the least significant bit, as follows:
[0215] if (LayerIdExtEnabled == true) LayerId = (nuh_layer_id << layer_id_ext_len) + layer_id_ext else LayerId = nuh_layer_id
[0216] This signaling scheme allows for signaling more distinct LayerId values by a smaller range of layer_id_ext values in case b) where nuh_layer_id can represent different values. It also allows for clustering of specific views; that is, views placed close together can use the same nuh_layer_id value to indicate that they belong together. See Figure 33.
[0217] Figure 33 illustrates the configuration of a view cluster, where all NAL units associated by a cluster (i.e., a group of camera views that are physically close) have the same value for nuh_layer_id and unique values for layer_id_ext. Alternatively, in other embodiments of the present invention, the syntax element layer_id_ext can be used to constitute the cluster, and nuh_layer_id can be used to identify the views within the cluster.
[0218] Another embodiment of the present invention derives a layer identifier, referred to here as LayerId, by using the existing layer identifier (nuh_layer_id) as the least significant bit and extended information as the most significant bit, as follows: if (LayerIdExtEnabled == true) LayerId = (layer_id_ext << 6) + nuh_layer_id else LayerId = nuh_layer_id
[0219] This signaling scheme enables signaling through clustering of specific views; that is, views of cameras physically separated from each other can use the same value of nuh_layer_id to instruct them to utilize the same predictive dependency with respect to the camera views by having the same value of nuh_layer_id in different clusters (i.e., the value of layer_id_ext in this embodiment).
[0220] Other embodiments employ an additional scheme to extend the LayerId range (maxNuhLayerId, which refers to the maximum allowable value of the existing layer identifier range (nuh_layer_id)):
[0221] if (LayerIdExtEnabled == true) LayerId = maxNuhLayerId + layer_id_ext else LayerId = nuh_layer_id
[0222] This signaling scheme is particularly useful in case a) where a predetermined value for nuh_layer_id is used to enable expansion. For example, the value of maxNuhLayerId can be used as a predetermined escape code to enable gapless expansion of the LayerId value range.
[0223] In the context of the draft of the test model for the HEVC 3D video coding extension, described as an early draft version of [3], possible embodiments are described in the following paragraphs.
[0224] In the initial version of Section G.3.5 of [3], the view components are defined as follows:
[0225] View Components: The encoded representation of a view in a single access unit A may include depth view components and texture view components.
[0226] The mapping of depth and texture view components is defined in the VPS extension syntax based on the existing layer identifier (nuh_layer_id). This invention adds flexibility to map the range of additional layer identifier values. An exemplary syntax is shown in Figure 34. Changes to the existing syntax are highlighted using shading.
[0227] If layer identifier extensions are used, VpsMaxLayerId is set to be equal to vps_max_layer_id; otherwise, it is set to be equal to vps_max_ext_layer_id.
[0228] If layer identifier extensions are used, VpsMaxNumLayers is set to the maximum number of layers that can be encoded using the extensions (either by a predefined number of bits or based on layer_id_ext_len); otherwise, VpsMaxNumLayers is set to vps_max_layers_minus1 + 1.
[0229] vps_max_ext_layer_id is the maximum used LayerId value.
[0230] `layer_id_in_nalu[i]` identifies the value of the LayerId associated with the VCLNAL unit of the i-th layer. For any i in the range of 0 to VpsMaxNumLayers - 1, comprehensively, if i does not exist, the value of layer_id_in_nalu[i] is assumed to be equal to i.
[0231] When i is greater than 0, layer_id_in_nalu[i] is greater than layer_id_in_nalu[i - 1]. When `splitting_flag` is equal to 1, if the total number of bits in a segment is less than 6, the MSB of `layer_id_in_nuh` must be 0.
[0232] For any value of i in the range of 0 to vps_max_layers_minus1, the variable LayerIdInVps[ layer_id_in_nalu[ i ] ] is inclusively set to equal to i.
[0233] `dimension_id[i][j]` identifies the identifier of the j-th current scalability dimension type for the i-th layer. If it does not exist, the value of `dimension_id[i][j]` is presumed to be equal to 0. The number of bits used for the representation of `dimension_id[i][j]` is `dimension_id_len_minus1[j] + 1` bits. When `splitting_flag` is equal to 1, it is a necessary condition for bitstream matching such that `dimension_id[i][j]` is equal to ((layer_id_in_nalu[i] & ((1 << dimBitOffset[j + 1]) - 1)) >> dimBitOffset[j]).
[0234] The variable ScalabilityId[i][smIdx], which identifies the identifier of the i-th scalability dimension type for the i-th layer, the variable ViewId[layer_id_in_nuh[i]], which identifies the view identifier for the i-th layer, and DependencyId[layer_id_in_nalu[i]], which identifies the spatial / SNR scalability identifier for the i-th layer, are derived as follows:
[0235] for (i = 0; i < VpsMaxNumLayers; i++) { for( smIdx= 0, j =0; smIdx< 16; smIdx ++ ) if( ( i ! = 0 ) && scalability_mask[ smIdx ] ) ScalabilityId[ i ][ smIdx ] = dimension_id[ i ][ j++ ] else ScalabilityId[i][smIdx] = 0 ViewId[ layer_id_in_nalu[ i ] ] = ScalabilityId[ i ]
[0000] DependencyId [ layer_id_in_nalu[ i ] ] = ScalabilityId[ i ]
[0001] }
[0236] In Section 2 of the initial version [3], it is stated that the corresponding depth view and texture components of a particular camera are derived from other depth views and textures in the NAL Unit Header Semantics section of the initial version [3] as follows, and can be distinguished by their extensibility-identifying view order index (ViewIdx) and depth flag (DepthFlag).
[0237] ViewIdx = layer_id >> 1 DepthFlag = layer_id % 2
[0238] Therefore, the individual view components (i.e., the texture and depth view components of a particular camera) must be packetized into NAL units by the individual values of distinguishable layer_id via the value of variable ViewIdx in the decoding process in, for example, the initial version 0 of section G.8.
[0239] The concept just outlined allows the use of the same value of the layer identifier in the NAL unit header (nuh_layer_id) for different views. Thus, the derivation of the identifiers ViewIdx and DepthFlag needs to be adapted to use the extended view identifier derived previously as follows:
[0240] ViewIdx = LayerId >> 1 DepthFlag = LayerId % 2
[0241] A generalized embodiment of the third form is described later with respect to FIG. 35, which shows a decoder 800 configured to decode a multi-layer video signal. The decoder can be implemented as outlined above with respect to FIGS. 2, 9 or 18. That is, examples for a more detailed description of the decoder 800 of FIG. 35 according to a particular embodiment can be obtained using those aspects and embodiments outlined above. To illustrate the possible overlap between the aspects and embodiments outlined above and the embodiment of FIG. 35, for example, the same reference signs are used for the multi-layer video signal 40 in FIG. 35. Regarding what the multiple layers of the multi-layer video signal 40 can do, reference is made to the description proposed with respect to the second aspect.
[0242] As shown in Figure 35, the multilayer video signal consists of a sequence of packets 804, each comprising a layer identification syntax element 806, which is embodied using the syntax element nuh_layer_id in the specific HEVC extension embodiment outlined above. The decoder 800 is configured to respond to signaling of the layer identification extension mechanism in the multilayer video signal 40, which itself may partially include a layer identification syntax element, as outlined further below. The signaling 808 of the layer identification extension mechanism is detected by the decoder 800 and, in response to the signaling 808, operates as follows for a given packet in a packet group 804 that includes a given packet input to the decoder 800 using the arrow 810. As illustrated using a switch 812 of the decoder 800 controlled via signaling 808 of the layer identification extension mechanism, the decoder 800 reads a layer identification extension from the multilayer data stream 40 at 814 for a predetermined packet 810, and uses this layer identification extension to determine the layer identification index of the current packet 810 at 816. If signaling 808 signals the deactivation of the layer identification extension mechanism, the layer identification extension is read at 814 and, as illustrated at 818, can be provided by the current packet 810 itself or can be located elsewhere in the data stream 40 in a manner that can be associated with the current packet 810. Thus, if signaling 808 of the layer identification extension mechanism signals the activation of the layer identification extension mechanism, the decoder 800 determines the layer identification index for the current packet 810 according to 814 and 816. However, if the signaling 808 of the layer identification extension mechanism signals the deactivation of the layer identification extension mechanism, the decoder 800 simply determines a predetermined layer identification index of the packet from the layer identification syntax element 806 of the current packet 810 820. In that case, the layer identification extension 818, i.e., its presence in signal 40, is unnecessary, i.e., it does not exist.
[0243] In an embodiment, the layer identification syntax element 806 contributes to the signaling of the layer identification extension mechanism 808 in packet detection: as far as each packet, such as the current packet 810, is concerned, whether the signaling of the layer identification extension mechanism 808 signals the activation or deactivation of the layer identification extension mechanism is determined by the decoder 800, at least in part, according to whether the layer identification syntax element 806 for each packet 810 assumes an escape value. For example, within a particular set of parameters 824, a high-level syntax element 822 provided by the data stream 40 can contribute to the signaling of the layer identification extension mechanism 808, i.e., the activation or deactivation of the layer identification extension mechanism, rather macroscopically or over a wider range. In particular, the decoder 800 can be configured to determine whether the signaling of the layer identification extension mechanism 808 signals the activation or deactivation of the layer identification extension mechanism, primarily for predetermined packets 810, according to the high-level syntax element 822. When the high-level syntax element takes the first state, the layer identification extension mechanism is signaled by the deactivating signaling 808. Referring to the embodiments outlined above, this relates to layer_id_ext_flag = 0, layer_id_mode_idc = 0, or layer_id_ext_len = 0. In other words, in the embodiments of the particular syntax described above, layer_id_ext_flag, layer_id_ext_idc, and layer_id_ext_len represent embodiments for the high-level syntax element 822, respectively.
[0244] For a specific packet, such as packet 810, this means that the decoder 800 decides that the signaling 808 of the level identification extension mechanism signals activation of the level identification extension mechanism for packet 810 if the high-level syntax element 822 takes a state different from the first state and the layer identification syntax element 806 of packet 810 takes an escape value. However, if the valid high-level syntax element 822 for packet 810 takes the first state or the layer identification element 806 of packet 810 takes a value different from the escape value, the decoder 800 decides to deactivate the layer identification extension mechanism that is signaled by the signaling 808.
[0245] Rather than simply having only two possible states, as outlined in the syntax embodiment above, the high-level syntax element 822 has multiple further states that the high-level syntax element 824 can take beyond the deactivated state, i.e., the first state. Depending on these further possible states, the determination 816 can change as indicated by the dashed line 824. For example, in the syntax embodiment above, the case of layer_id_mode_idc = 2 is shown to likely result in the decoder 800 concatenating a digit representing the layer identification syntax element 806 of packet 810 with a digit representing the layer identification extension in order to obtain the layer identification index of packet 810. In contrast, the embodiment case where layer_id_len≠0 is shown to likely result in a decoder 800 performing the following: the decoder 800 determines the length n of the layer identification extension 818 associated with packet 810 using a high-level syntax element to obtain a predetermined level identification index for the packet, and concatenates the digit representing the layer identification syntax element 806 of packet 810 with the n digit representing the level identification extension 818 of packet 810. Furthermore, the decision 816 may include adding the level identification extension 818 associated with packet 810 to a predetermined value that can accommodate a number exceeding the maximum number of states (less than the escape value) that the layer identification syntax element 806 can represent, for example, to obtain a predetermined layer identification index for packet 810.
[0246] However, as indicated by 808' in Figure 34, it is also possible to exclude the layer identification syntax element 806 of packet 810 from contributing to the signaling 808 of the layer identification extension mechanism, such that all the representable values / states of the syntax element 806 remain, and none of them are reserved as escape codes. In that case, the signaling 808' instructs the decoder 800 for each packet 810 whether or not a layer identification extension 818 exists, and therefore whether the determination of the layer identification index follows 814 and 816 or 820.
[0247] The encoder that fits into the decoder in Figure 35 then simply forms the data stream. The encoder decides whether or not to use an expansion mechanism, for example, according to the number of layers encoded in the data stream.
[0248] A fourth aspect of the present application relates to dimension-dependent direct dependency signaling.
[0249] In the current HEVC extensions ([2], [3], [4]), coding layers can utilize zero or more reference coding layers for data prediction. Each coding layer is identified by a unique nuh_layer_id value, which can be bijectively mapped to a layerIdInVps value. The layerIdInVps value is continuous, and when a layer with layerIdInVps equal to A is referenced by a layer with layerIdInVps B, the requirement for bitstream compatibility is that A is less than B.
[0250] For each coding layer in the bitstream, the reference coding layer is signaled in the video parameter set. Therefore, a binary mask is transmitted for each coding layer. For a coding layer with layerIdinVps value b, the mask (represented as direct_dependency_flag[ b ] ) consists of b-1 bits. When a layer with layerIdinVps equal to x is the reference layer of a layer with layerIdinVps equal to b, the x-th bit in the binary mask (represented as direct_dependency_flag[ b ][ x ] ) is equal to 1. Otherwise, when a layer with layerIdinVps equal to x is not the reference layer of a layer with layerIdinVps equal to B, the value of direct_dependency_flag[ b ][ x ] is equal to 0.
[0251] After parsing all direct_dependency_flags, a list is created for each coding layer containing the nuh_layer_id values of all reference layers, as identified by direct_dependency_flags.
[0252] Furthermore, information is signaled in the VPS that allows each layerIdinVps value to be mapped to a position in the scalability space of the T dimension. Each dimension t represents the type of scalability, which can be, for example, view scalability, spatial scalability, or distance image indication.
[0253] By signaling one bit for each possible dependency, the current design offers maximum flexibility. However, this flexibility comes with several drawbacks:
[0254] 1. This is a case of a common usage where a specific dependency structure is used for each scalability dimension. Furthermore, direct dependencies between dimensions are uncommon and will be rejected. An example of a typical layer setup is shown in Figure 36. Here, dimension 0 would be a view scalability dimension that utilizes a kind of hierarchical predictive structure. Dimension 1 would be a spatial scalability dimension using an IP structure. The direct_dependency_flags related to the setup are shown in Figure 37. The drawback of the current solution is that it requires an algorithmically complex analysis of direct_dependency_flags, making it indirect to identify this type of dimension-dependent dependency from the current VPS design. 2. Even when only one scalable dimension type is used, the same structure is generally used for subsets of layers. For example, in the case of view scalability only, the view would be mapped to a space spanned by horizontal and vertical camera positions. An example of this type of scenario is shown in Figure 36, where dimensions 0 and 1 are interpreted as the dimensions of the horizontal and vertical camera positions. While it is common practice to use one predictive structure for each camera position dimension, current VPS designs cannot take advantage of the resulting redundancy. Furthermore, current VPS designs do not directly indicate that dependencies are dimension-dependent. 3. The number of direct_dependency_flags is proportional to the square of the number of layers in the bitstream, and therefore, the current worst-case scenario with 64 layers requires approximately 64 * 63 / 2 = 2016 bits. Furthermore, when the maximum number of layers in the bitstream is expanded, this results in a drastically increasing number of bits.
[0255] The aforementioned drawbacks can be resolved by enabling clear signaling of dependencies for each dimension t in the dependency space of the T dimension.
[0256] Dimensional dependency direct dependency signaling offers the following benefits: 1. Dependencies for each dependency dimension are directly available in the bitstream, and complex parsing of direct_dependency_flags is not required. 2. The number of bits required for dependent signaling can be reduced.
[0257] In some embodiments, the dependency space may be the same as the scalability space, for example, as described in the current MV- and scalable drafts [2]. In other embodiments, the dependency space may be explicitly signaled, for example, a spanned space by camera position.
[0258] An example of dimension-dependent dependency signaling is given in Figure 38. It can be seen that the dependencies between dimensions can be directly derived from a binary mask, reducing the number of bits required.
[0259] In the following, it is assumed that each layerIdInVps value is bijectively mapped to a T-dimensional dependency space with dimensions 0, 1, 2, ..., (T-1). Therefore, each layer identifies its position in the corresponding dimensions 0, 1, 2, ..., (T-1), with d0, d1, d2, ..., d T-1 Related vector (d0,d1,d2,…,d T-1 )' has.
[0260] The basic idea is layer-dependent dimension-dependent signaling. Therefore, each dimension t ∈ {0, 1, 2 … (T-1)} and each position d in dimension t tFor a reference position set Ref(d t ) in dimension t is signaled. As described below, the reference position set is used to determine the direct dependencies between different layers:
[0261] A layer with position d in dimension t t and position d in dimension x x where x ∈ {0, 1, 2 … (T-1)} / {t} depends on d t,Ref when d t is an element in Ref(d t,Ref and position d in dimension x x where x ∈ {0, 1, 2 … (T-1)} / {t}.
[0262] In other specific embodiments, all dependencies are reversed, and thus the position in Ref(d t ) indicates the position of the layer in dimension t that depends on the layer at position d in dimension t t .
[0263] As far as the signaling and derivation of the dependency space are concerned, the signaling described below can be done, for example, in the VPS in the SEI message, in the SPS, or in other places in the bitstream.
[0264] Regarding the number of dimensions and the number of positions in the dimensions, the following is noted. The dependency space is defined by a specific number of dimensions and a specific number of positions in each dimension.
[0265] In a specific embodiment, the number of dimensions num_dims and the number of positions num_pos_minus1[t] in dimension t can be explicitly signaled, for example, as shown in FIG. 39.
[0266] In other embodiments, the value of num_dims or num_pos_minus1 can be fixed and not signaled in the bitstream. In other embodiments, the value of num_dims or num_pos_minus1 can be derived from other syntax elements present in the bitstream. More specifically, in the current HEVC extension design, the number of dimensions and the number of positions in a dimension can be equal to the number of scalability dimensions and the length of a scalability dimension, respectively.
[0267] Therefore, by NumScalabilityTypes and dimension_id_len_minus1[t] defined in [2]: num_dims = NumScalabilityTypes num_pos_minus1[ t ] = dimension_id_len_minus1[ t ]
[0268] In other embodiments, whether the value of num_dims or the value of num_pos_minus1 is explicitly signaled or derived from other syntax elements present in the bitstream can be signaled in the bitstream.
[0269] In other embodiments, the value of num_dims can be derived from other syntax elements present in the bitstream and augmented by additional signaling of the division of one or more dimensions or by signaling an additional dimension.
[0270] Regarding the mapping of layerIdInVps to positions in the dependency space, it is noteworthy that layers are mapped to the dependency space.
[0271] In a particular embodiment, the syntax element pos_in_dim[i][t] that identifies the position of a layer by the layerIdinVps value i in dimension t can, for example, be explicitly transmitted. This is illustrated in Figure 40.
[0272] In other embodiments, the value of pos_in_dim[i][t] is not signaled in the bitstream, but can be directly derived from the layerIdInVps value i, for example, as follows:
[0273] idx = i dimDiv
[0000] = 1 for ( t = 0; t < T 1 ; t++ ) dimDiv[ t + 1 ] = dimDiv[ t ] * ( num_pos_minus1[ t ] + 1 ) for ( t = T 1 ; t >= 0; t-- ) [ pos_in_dim[ i ][ t ] = idx / dimDiv[ t ] / / integer devision idx = idx pos_in_dim[ i ][ t ] * dimDiv[ t ] }
[0274] In particular, the above description may replace the current explicit signaling of the dimension_id[i][t] value for the current HEVC extension design.
[0275] In other embodiments, the value of pos_in_dim[i][t] is derived from other syntax elements in the bitstream. More specifically, in the current HEVC extension design, the value of pos_in_dim[i][t] can be derived, for example, from the dimension_id[i][t] value.
[0276] pos_in_dim[ i ][ t ] = dimension_id[ i ][ t ]
[0277] In other embodiments, pos_in_dim[i][t] can be signaled to be either explicitly signaled or derived from other syntax elements.
[0278] In another embodiment, in addition to the pos_in_dim[i][t] value derived from other syntax elements present in the bitstream, it is possible to signal whether the pos_in_dim[i][t] value is explicitly signaled.
[0279] The following methods are used for signaling and deriving dependencies.
[0280] The use of direct position dependency flags is the subject of the following embodiments. In these embodiments, reference positions are signaled by a flag, for example, pos_dependency_flag[ t ][ m ][ n ], which indicates whether position n in dimension t is included in the set of reference positions for position m in dimension t, as specified, for example in Figure 41.
[0281] In embodiments using a reference position set, the variables num_ref_pos[ t ][ m ], which specify the number of reference positions in dimension t for a given position m in dimension t, and ref_pos_set[ t ][ m ][ j ], which specify the j-th reference position in dimension t for a given position m in dimension t, can be derived as follows: for( t = 0; t <= num_dims; t++ ) for( m = 1; m <= num_pos_minus1[ t ]; m++ ) num_ref_pos[ t ][ m ] = 0 for( n = 0; n < m; n++ ) { if ( pos_dependency_flag[ t ][ m ][ n ] = = true ) { ref_pos_set[ t ][ m ][ num_ref_pos[ t ][ m ] ] = n num_ref_pos[ t ][ m ] ++ } }
[0282] In other embodiments, elements of a set of reference locations can be directly signaled, for example, as shown in Figure 42.
[0283] In embodiments using direct dependency flags, the direct dependency flag `directDependencyFlag[i][j]`, which identifies that a layer with `layerIdInVps` equal to `i` is dependent on a layer with `layerIdInVps` equal to `j`, would be derived from a set of reference locations. For example, it would be done as follows:
[0284] The function posVecToPosIdx(posVector) with a vector posVector as input derives an index posIdx related to the position posVector in the dependency space, as specified below:
[0285] for ( t = 0, posIdx = 0, offset = 1; t < num_dims; t++) [ posIdx = posIdx + offset * posVector[ t ] offset = offset * ( num_pos_minus1[ t ] + 1 ); }
[0286] The variable posIdxToLayerIdInVps[idx], which identifies the layerIdinVps value i dependent on the index idx derived from pos_in_dim[i], can be derived, for example, as follows:
[0287] for (i = 0; i < vps_max_layers_minus1; i++) posIdxToLayerIdInVps[ posVecToPosIdx( pos_in_dim[ i ] )] = i
[0288] The variable directDependencyFlag[i][j] is derived as follows:
[0289] for (i = 0; i <= vps_max_layers_minus1; i++) { for (k = 0; k < i; k++) directDependencyFlag[ i ][ k ] = 0 curPosVec = pos_in_dim[ i ] for (t = 0; t < num_dims; t++) { for (j = 0; j < num_ref_pos[ t ][ curPosVec[ t ] ]; j++) { refPosVec = curPosVec refPosVec[ t ] = ref_pos_set[ t ][ curPosVec[ t ] ][ j ] directDependencyFlag[ i ][ posIdxToLayerIdInVps[ posVecToPosIdx( refPosVec ) ] ] = 1 } } }
[0290] In an embodiment, the direct dependency flag `directDependencyFlag[i][j]`, which identifies that a layer with `layerIdInVps` equal to `i` is dependent on a layer with `layerIdInVps` equal to `j`, would be directly derived from the `pos_dependency_flag[t][m][n]` flag. For example, as specified below:
[0291] for (i = 1; i <= vps_max_layers_minus1; i++) { curPosVec = pos_in_dim[ i ]; for (j = 0; j < i; j++) { refPosVec = pos_in_dim[ j ] for (t = 0, nD = 0; t < num_dims; t++) if ( curPosVec[ t ] ! = refPosVec[ j ][ t ] ) { nD++ tD = t } if ( nD = = 1 ) directDependencyFlag[ i ][ j ] = pos_dependency_flag[ tD ][ curPosVec[ tD ] ][ refPosVec[ tD ] ] else directDependencyFlag[ i ][ j ] = 0 } }
[0292] In embodiments using a set of reference layers, the variable NumDirectRefLayers[i], which identifies the number of reference layers for a layer with layerIdInVps equal to i, and the variable RefLayerId[i][k], which identifies the value of layerIdInVps for the k-th reference layer, may be derived, for example, as follows:
[0293] for( i = 1; i <= vps_max_layers_minus1; i++ ) for( j = 0, NumDirectRefLayers[ i ] = 0; j < i; j++ ) if( directDependencyFlag[ i ][ j ] = = 1 ) RefLayerId[ i ][ NumDirectRefLayers[ i ]++ ] = layer_id_in_nuh[ j ]
[0294] In other embodiments, for example, as specified below, the reference layer can be derived directly from the reference location set without deriving the directDependencyFlag value:
[0295] for (i = 0; i <= vps_max_layers_minus1; i++) { NumDirectRefLayers[i] = 0 curPosVec = pos_in_dim[ i ] for (t = 0; t < num_dims; t++) [ for (j = 0; j < num_ref_pos[ t ][ curPosVec[ t ] ]; j++) { refPosVec = curPosVec refPosVec[ t ] = ref_pos_set[ t ][ curPosVec[ t ] ][ j ] m = posIdxToLayerIdInVps[ posVecToPosIdx( refPosVec ) ] RefLayerId[ i ][ NumDirectRefLayers[ i ]++ ] = layer_id_in_nuh[ m ] } }
[0296] In other embodiments, the reference layer may be derived directly from the pos_dependency_flag variable without deriving the ref_pos_set variable.
[0297] Thus, the diagram discussed above illustrates a data stream according to a fourth aspect, revealing a multilayer video data stream in which video material is encoded with LayerIdInVps in different levels of information, i.e., number, using interlayer prediction. The levels have a sequential order determined among them. For example, they follow the sequence 1…vps_max_layers_minus1. See, for example, Figure 40, where the number of layers in the multilayer video data stream is given by 900, given by vps_max_layers_minus1.
[0298] Through inter-layer prediction, the video material is encoded into a multi-layer video data stream such that no layer is dependent on any subsequent layer in sequential order. That is, using numbering from 1 to vps_max_layers_minus1, layer i can simply be dependent on layer j < i.
[0299] Through inter-layer prediction, each layer dependent on one or more other layers increases the amount of information encoded in the video material by one or more other layers. For example, this increase relates to spatial resolution, number of views, SNR accuracy, or other dimension types.
[0300] A multilayer video data stream, for example at the VPS level, comprises a first syntax structure. In the above embodiment, num_dims can be comprised by the first syntax structure, as shown in 902 in Figure 39. Thus, the first syntax structure defines a number M of dependent dimensions 904 and 906. In Figure 36, there are two, one horizontally and the other vertically. In this regard, see item 2 above: the number of dimensions is not necessarily equal to the number of different dimension types, in terms of how the level increases the amount of information: the number of dimensions can be greater, for example, by distinguishing between vertical and horizontal view shifts. M dependent dimensions 904 and 906 spanning the dependent space 908 are shown exemplarily in Figure 36.
[0301] TIFF2026062695000002.tif86170
[0302] A multilayer video data stream comprises a second syntax structure 914 for each dependent dimension i, for example at the VPS level. In the above embodiment, it includes pos_dependency_flag[ t ][ m ][ n ] or num_ref_pos[ t ][ m ] + ref_pos_set[ t ][ m ][ j ]. The second syntax structure 914 describes the dependency within the Ni rank level of the dependent dimension i for each dependent dimension i. The dependency is illustrated in Figure 36 by all horizontal arrows or all vertical arrows between rectangles 910.
[0303] In general, this means that for each dependency dimension, all of these dependencies run parallel to each of the dependency axes by dependencies parallel to each of the dependency dimensions other than the respective dependency dimension, and the dependencies between available points in the dependency space are defined in a restricted manner, so as to point from higher rank levels to lower rank levels. See Figure 36: all horizontal arrows between rectangles in the upper line of the rectangle are duplicated in the lower row of the rectangle, and the same applies to the vertical arrows with respect to the four vertical columns of the rectangle, by the rectangles corresponding to the available points and the arrows corresponding to their dependencies. Through this means, via bijective mapping, the second syntax structure simultaneously defines dependencies between layers.
[0304] A network entity like a decoder or a mane like an MME can read the first and second syntax structures of a data stream and determine the dependencies between layers based on the first and second syntax structures.
[0305] TIFF2026062695000003.tif61170
[0306] In this case, the network entity can select one of the levels; and the selected level discards packets of multilayer video data streams, such as NAL units, that belong to the layer that is independent as a dependency between layers, for example via nuh_layer_id.
[0307] While several embodiments have been described in the context of apparatus, it is clear that these embodiments also indicate a description of a corresponding method, where a block or device corresponds to a method step or a feature of a method step. Similarly, embodiments described in the context of a method step also indicate a description of a corresponding block, item, or feature of a corresponding apparatus. Some or all method steps can be performed by (or using) a hardware device, such as a microprocessor, a programmable computer, or an electronic circuit. In some embodiments, one or more of the most important method steps can be performed by such a device.
[0308] Embodiments of the present invention can be implemented in hardware or software, depending on the specific requirements for implementation. Implementations can be carried out using a digital storage medium, such as a floppy disk, DVD, Blu-ray, CD, ROM, PROM, EPROM, EEPROM, or flash memory, which has electronically readable control signals stored thereon and cooperates (or can cooperate) with a computer system programmable to perform each method. Therefore, the digital storage medium can be computer-readable.
[0309] Some embodiments of the present invention include a data carrier having an electronically readable control signal and capable of cooperating with a programmable computer system on which one of the methods described herein is performed.
[0310] Generally, embodiments of the present invention can be implemented as a computer program product having program code that is operable to perform one of the methods when the computer program product runs on a computer. The program code can be stored, for example, on a machine-readable carrier.
[0311] Other embodiments include a computer program stored in a machine-readable carrier for performing one of the methods described in this specification.
[0312] In other words, an embodiment of the method of the invention is a computer program having program code for performing one of the methods described in this specification when the computer program is running on a computer.
[0313] A further embodiment of the method of the invention is a data carrier (or digital storage medium or computer-readable medium) on which a computer program is recorded and which performs one of the methods described in this specification. The data carrier, digital storage medium or recording medium is usually tangible and / or fixed.
[0314] A further embodiment of the method of the invention is a data stream or sequence of signals representing a computer program for performing one of the methods described in this specification. The data stream or sequence of signals may be configured to be transmitted, for example, over a data communication connection, for example, over the Internet.
[0315] Further embodiments include processing means, such as a computer or a programmable logic device, configured or adapted to perform one of the methods described in this specification.
[0316] Further embodiments include a computer on which a computer program is installed that performs one of the methods described in this specification.
[0317] Further embodiments of the present invention include an apparatus or system configured to transfer (e.g., electronically or optically) a computer program that performs one of the methods described in this specification to a receiver. The receiver may be, for example, a computer, a mobile device, a memory device, etc. The apparatus or system may include, for example, a file server that transfers the computer program to the receiver.
[0318] In some embodiments, a programmable logic device (e.g., a field-programmable gate array) can be used to perform some or all of the functions of the methods described herein. In some embodiments, a field-programmable gate array can cooperate with a microprocessor to perform one of the methods described herein. Generally, the methods are preferably performed by any hardware device.
[0319] The apparatus described in this specification may be implemented using hardware devices, a computer, or a combination of hardware devices and a computer.
[0320] The methods described in this specification may be carried out using hardware devices, or using a computer, or using a combination of hardware devices and a computer.
[0321] The embodiments described above are merely illustrations of the principles of the present invention. Modifications and changes to the configurations and details described in this specification will be obvious to those skilled in the art. Therefore, the present invention is intended to be limited only by the scope of the claims and not by any specific details provided in the description and explanation of the embodiments in this specification.
[0322] [References] [1]B. Bross et al., "High Efficiency Video Coding (HEVC) text specification draft 10", JCTVC-L1003, Geneva, CH, 14-23 Jan. 2013 [2]G. Tech et al., "MV-HEVC Draft Text 3", JCT3V-C1004, Geneva, CH , 17-23 Jan. 2013 [3]G. Tech et al., "3D-HEVC Test Model 3", JCT3V-C1005, Geneva, CH , 17-23 Jan. 2013 [4]J. Chen et al., "SHVC Draft Text 1", JCT-VCL1008, Geneva, CH , 17-23 Jan. 2013 [5]WILBURN, Bennett, et al. High performance imaging using large camera arrays. ACM Transactions on Graphics, 2005, 24. Jg., Nr. 3, S. 765-776. [6]WILBURN, Bennett S., et al. Light field video camera. In: Electronic Imaging 2002. International Society for Optics and Photonics, 2001. S. 29-36. [7]HORIMAI, Hideyoshi, et al. Full-color 3D display system with 360 degree horizontal viewing angle. In: Proc. Int. Symposium of 3D and Contents. 2010. S. 7-10.
Claims
1. Multiple coded picture buffers (CPBs) and Processor and A decoder equipped with, The aforementioned processor, A multilayer encoded video bitstream is received, wherein each layer of the multilayer video bitstream comprises a sequence of network abstraction layer (NAL) units, and one or more NAL units of the sequence associated with the same time instance are grouped into decoding units (DUs) of the corresponding layers of the multilayer encoded video bitstream. The DU of each layer is stored in the respective CPB among the plurality of CPBs. A decoder configured such that each CPB stores one layer of DU.
2. The decoder according to claim 1, wherein the processor is configured to decode the multilayer encoded video bitstream by simultaneously removing one or more DUs of each layer associated with the same time instance from their respective CPBs.
3. The decoder according to claim 2, wherein each DU of each layer associated with the same time instance has the same nominal CPB rejection time.
4. The decoder according to claim 2, wherein the processor is a parallel processor configured to remove DUs associated with the same time instance from their respective CPBs in accordance with parallel processing.
5. The decoder according to claim 1, wherein the first of the plurality of CPBs is configured to store the DU of the base layer of the multilayer encoded video bitstream.
6. The decoder according to claim 5, wherein a second of the plurality of CPBs is configured to store the DU of the enhancement layer of the multilayer encoded video bitstream, which depends on the base layer.
7. The decoder according to claim 1, further comprising a multiplexer configured to transfer the DU of each layer to each of the plurality of CPBs.
8. The decoder according to claim 1, wherein the multilayer encoded bitstream comprises a plurality of access units, each access unit comprising NAL units or DUs of all layers of the multilayer encoded video bitstream interleaved within each access unit, wherein the NAL units or DUs within each access unit are associated with the same time instance.
9. Multiple coded picture buffers (CPBs) and Processor and An encoder comprising, The aforementioned processor, A multilayer encoded video bitstream encodes video, wherein each layer of the multilayer encoded video bitstream comprises a sequence of network abstraction layer (NAL) units, and one or more NAL units of the sequence associated with the same time instance are grouped into a decoding unit (DU) of the corresponding layer of the multilayer encoded video bitstream. The DU of each layer is stored in the respective CPB among the plurality of CPBs. An encoder configured such that each CPB stores one layer of DU.
10. The encoder according to claim 9, wherein each DU of each layer associated with the same time instance has the same nominal CPB rejection time.
11. The encoder according to claim 9, wherein the first of the plurality of CPBs is configured to store the DU of the base layer of the multilayer encoded video bitstream.
12. The encoder according to claim 9, wherein a second of the plurality of CPBs is configured to store the DU of the enhancement layer of the multilayer encoded video bitstream, which depends on the base layer.
13. The encoder according to claim 9, further comprising a multiplexer configured to transfer the DU of each layer to each of the plurality of CPBs.
14. The encoder according to claim 9, wherein the multilayer encoded bitstream comprises a plurality of access units, each access unit comprising NAL units or DUs of all layers of the multilayer encoded video bitstream interleaved within each access unit, wherein the NAL units or DUs within each access unit are associated with the same time instance.
15. A non-temporary computer-readable medium for storing a data stream associated with a video, wherein the data stream is: A multilayer encoded video bitstream, wherein each layer of the multilayer video bitstream comprises a sequence of network abstraction layer (NAL) units, and one or more NAL units of the sequence associated with the same time instance are grouped into a decoding unit (DU) of the corresponding layer of the multilayer encoded video bitstream. Includes, Here, the DU of each layer is stored in its respective CPB among several coded picture buffers (CPBs). Each CPB stores DU for one layer. Non-temporary computer-readable media.
16. The non-temporary computer-readable medium according to claim 15, wherein the multilayer encoded bitstream includes a base layer, and the DU of the base layer is stored in a first of the plurality of CPBs.
17. The non-temporary computer-readable medium according to claim 16, wherein the multilayer encoded video bitstream further comprises an enhancement layer dependent on the base layer, and the DU of the enhancement layer is stored in a second of the plurality of CPBs.
18. The non-temporary computer-readable medium according to claim 15, wherein the multilayer encoded video bitstream comprises a plurality of access units, each access unit comprising NAL units or DUs of all layers of the multilayer encoded video bitstream interleaved within each access unit, wherein the NAL units or DUs within each access unit are associated with the same time instance.