Highly scalable coding concept
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2014-01-04
- Publication Date
- 2026-08-11
Smart Images

Figure CN116320392B_ABST
Abstract
Description
[0001] This application is a divisional application, with its parent application having application number 201811477939.1, an application date of January 4, 2014, and an invention title of "High-Efficiency Scalable Coding Concept". Technical Field
[0002] This application relates to scalable coding concepts such as scalable video coding. Background Technology
[0003] The concept of scalable coding is known in this field. For example, in video coding, H.264 involves SVC (Scalable Video Coding) extensions, which allow additional enhancement layer data to be appended to the base layer encoded video data stream to improve the reconstruction quality of the base layer video in various aspects such as spatial resolution and signal-to-noise ratio (SNR). The recently finalized HEVC standard can also be extended via SVC profiles. HEVC differs from its predecessor H.264 in many ways, such as its suitability for parallel decoding / coding and low-latency transmission. Regarding the parallel encoding / decoding involved, HEVC supports WPP (Wavefront Parallel Processing) encoding / decoding and the concept of tile parallel processing. According to the WPP concept, individual pictures are divided into substreams in a row-wise manner. The encoding order within each substream is guided from left to right. The substreams have a defined decoding order, i.e., guided from the top substream to the bottom substream. Entropy coding of the substreams is performed using probabilistic adaptivity. For each substream, probability initialization is performed independently, or an initial self-adaptive state based on the probabilities used during entropy coding is established, ensuring that the immediately preceding substream reaches a specific position from the left-hand edge of the preceding substream at the end of the second CTB (Coding Tree Block). Spatial prediction is not restricted; that is, spatial prediction can cross boundaries between intermediate consecutive substreams. In this way, substreams can be encoded / decoded in parallel with the current encoding / decoding position forming the wavefront, with the wavefront advancing in a sloping manner from lower left to upper right and from left to right. Based on the slice concept, the image is segmented into slices, and spatial prediction across boundary lines is suppressed to allow for parallel processing of the encoding / decoding of these slices. In-loop filtering across slice boundaries is only permitted. To support low-latency processing, the slice concept is extended to allow slices to be switchable, re-initializing the entropy probabilities using the entropy probabilities saved during the substream processing—that is, the substream preceding the current slice's starting substream—and using continuously updated entropy probabilities until the immediately preceding slice ends. This measure makes WPP and the slice concept more suitable for low-latency processing.
[0004] However, the concept under discussion is more favorable for further improvements in scalable coding probabilities.
[0005] Therefore, the objective of this invention is to provide a concept that further improves the concept of scalable coding.
[0006] This objective is achieved through the subject matter of the pending independent claims. Summary of the Invention
[0007] The first aspect of this application relates to scalable video coding incorporating the concept of parallel processing. Parallel processing concepts such as WPP and slices allow for the parallel decoding of video images in spatial segments (e.g., in the form of substreams, slices, or slices) into which the images are subdivided. Just as spatial intra-frame image prediction limits the degree of parallelism of decoding layers that are interdependent with each other, inter-layer prediction, as spatial intra-frame image prediction, limits the degree of parallelism during single-layer image decoding, addressing this issue in different ways. For example, when slices are used as spatial segments, spatial intra-frame image prediction is limited by not being able to cross boundary lines. In the case of WPP substreams, their parallel processing is performed in an interleaved manner to produce an appropriate tilted processing wavefront. In the case of inter-layer prediction, decoding of dependent layers is performed based on the co-occurrence portion of the reference layer. Therefore, decoding of the spatial segments of dependent layers can begin first, at which point the co-occurrence portion of the reference layer has already been processed / decoded. Similar to the case of inter-layer prediction where different views are used as different layers, or due to upsampling from lower to higher layers, the area of the "co-occurrence portion" is magnified, allowing for "motion compensation." That is, it is easy for video decoders using inter-layer prediction scalable decoding and parallel decoding to derive the degree of parallelism of the parallel processing of interdependent layers from short-term syntax elements (LTS elements) about these related dependent layers, which define the subdivision of the images of these interdependent layers into their spatial segments. However, performing this operation stably is cumbersome and computationally complex. Furthermore, when performing this operation, the video decoder cannot properly schedule parallel-running decoding threads to decode multi-layer video data streams. Therefore, according to a first aspect of the invention, parallel decoding of interdependent layers of multi-layer video data streams is improved by introducing a long-term syntax element (LTS) structure that, when assuming specific values, guarantees that the video decoder subdivides the images of dependent layers into spatial segments such that the boundaries between spatial segments of the second layer overlap with each boundary of the spatial segments of the first layer within a predetermined time period larger than the time interval of the LTS elements. Through this measure, the video decoder can rely on the fact that the multi-layer video data stream is properly encoded to subdivide the images of interdependent layers into spatial segments without inadvertently reducing the feasible degree of parallelism between these interdependent layers. More precisely, within a predetermined time period, by developing constraints that allow the boundaries of spatial segments in different layers to overlap with each other in a signaling manner, the decoder can pre-schedule the distribution of spatial segments to the parallel processing threads of the video decoder. However, the long-term syntax element structure allows this guarantee to be disabled, thus allowing in other application scenarios or enabling high-end video decoders to perform parallel processing scheduling based solely on short-term syntax elements, i.e., without developing any guarantees regarding the relative positioning between the boundaries of spatial segments in interdependent layers. Long-term syntax elements can also be used for the purpose of determining opportunistic decoding.
[0008] Another aspect of this application relates to scalable coding, which involves inter-layer prediction based on upsampling from the base layer to the enhancement layer, combined with parallel processing of the interdependent layers. Specifically, this aspect relates to an interpolation method for performing upsampling from the base layer to the enhancement layer. Naturally, this interpolation method makes adjacent partitions of the base layer image interdependent. That is, the interpolation method makes the interpolation results at the outer circumference of each part of the upsampled base layer reference image interdependent with two pixels / cells within the co-occurring partition of the base layer image and pixels / cells in adjacent partitions. In other words, the region of the base layer image used as the reference for inter-layer prediction of the co-occurring part is predicted in the "blackened" or widened enhancement layer image. Incidentally, the interdependence caused by the interpolation method of inter-layer prediction adversely affects the degree of parallelism achievable during parallel processing of the interdependent layers. For example, according to a second aspect of the invention, a syntax element is introduced that informs the decoder to modify interpolation along the partitions of the base layer to avoid mixing pixels / cells of adjacent partitions of the base layer image, the partitions of the base layer image and their upsampling patterns depending on the spatial segments of the enhancement layer image or the spatial segments of the base layer and enhancement layer. By introducing this syntax element, the encoder can switch between two modes: if the interpolation method is restricted to partially self-contained base layer images, i.e., a restriction is enabled, then the parallelism obtained during parallel decoding of interdependent layers is maximized as the interpolation quality along the partition edges of the base layer image decreases slightly; however, without restricting the interpolation method, the parallelism decreases and the interpolation quality at the partition edges increases.
[0009] The third aspect of this application relates to scalable video coding with parallel decoding of interdependent layers and attempts to alleviate the burden on the decoder to perform parallel processing scheduling by introducing a long-term syntax element (LTF) structure, i.e., distributing spatial segments across parallel processing threads. The LTF structure allows the decoder to determine inter-layer offsets or inter-layer delays over a predetermined time period longer than the time interval, during which short-term syntax elements signal the size and location of spatial segments of the images of the interdependent layers, as well as the spatial sampling resolution of these images. By introducing an LTF element that signals the inter-layer offset, the video encoder can switch between two modes: In a first mode, the encoder guarantees the decoder a specific inter-layer offset corresponding to a specific degree of parallelism between the decoding of the interdependent layers and accordingly sets the short-term syntax elements within the predetermined time period such that the actual inter-layer offset is equal to or even lower than the guaranteed offset. In the other mode, no guarantee is provided to the decoder; therefore, the encoder does not need to set the short-term syntax elements to meet another criterion, such as, optionally, adapting the short-term syntax elements to the video content within the predetermined time period. Therefore, when the count is obeyed over the entire predetermined time period, and the first spatial segment of the enhancement layer image that is temporally co-aligned does not face any conflict, at least relative to the decoding of the first spatial segment of the enhancement layer image over the predetermined time period, the interlayer offset that explicitly signals in the data stream can be the count of the minimum base layer spatial segment that must be decoded.
[0010] The fourth aspect of this application relates to the signaling of each layer in the various NAL units to which scalable video coding and multi-layer video data streams belong, the positioning of these layers within the scalable space, and the meaning of the scalable dimension across the scalable space. This information is easily accessible and manageable for intermediate network entities involved in transmitting multi-layer video data streams, so that tasks can be easily performed by these entities. Based on the inventor's normal, the fourth aspect of this application demonstrates, in a typical application scenario, the expenditure of the type indicator field, the way the type indicator field is associated with the layer indicator fields within the NAL unit header, whereby the type indicator field is interpreted as follows: if the type indicator field has a first state, the mapping information in the common information NAL unit maps the possible values of the layer indicator field to the operation point, and associates the layer NAL unit with the operation point using the corresponding layer indicator field and mapping information. Similarly, the mapping between each layer and the scalable constellation can be adapted to varying degrees and allows for the implementation of most scalable spaces; however, this increases the disadvantage of managing the total overhead. If the type indicator field has a second state, the layer indicator field is divided into more than one part, and the operation point associated with the corresponding NAL unit is located by using the values of these parts as coordinates of vectors within the scalable space. This approach allows for a smaller number of scalable spaces in the mapping between layers and the scalable constellation, while simultaneously reducing the overall administrative overhead of the network entities. In both cases, the layer indicator field can be identical, independent of application requests; however, the way the layer indicator field guides the NAL units of the layer through the scalable space can be adapted to the current application and its specific details. The advantages of this adaptation overcompensate for the necessity of the additional overhead of the type indicator field.
[0011] The fifth aspect of this application relates to multi-layer video coding, i.e., scalable video coding, which allows different layers to use different codecs / standards. The possibility of allowing consecutive layers to use different codecs / standards enables latency scaling of existing video environments, thereby addressing multi-layer video streams subsequently scaled by further enhancement layers and, consequently, the use of new and potentially better codecs / standards. A network aggregation of codecs / standards for some enhancement layers can still process lower layers and feed back to the multi-codec decoder via the transport layer decoder. For each NAL unit of the multi-layer video stream, the transport layer decoder identifies the same codec associated with it and accordingly hands over the NAL unit of the multi-layer video stream to the multi-standard multi-layer decoder.
[0012] The sixth aspect of this application relates to multi-layer video coding, which subdivides both base layer and enhancement layer images into arrays of blocks. In this case, inter-layer offset signals can be effectively emitted by inserting syntax element structures into the multi-layer video data stream. These syntax element structures indicate the inter-layer offsets used for parallel decoding of base layer and enhancement layer images, unit by unit. That is, the sixth aspect of this application, based on the discovery of explicit transmission syntax element structures, indicates that the inter-layer offsets between the base layer and enhancement layer, unit by unit, are increased by mirroring the transmitted data. This significantly reduces the decoder's computational complexity compared to cases where the decoder derives the inter-layer offsets for parallel decoding of base layer and enhancement layer images based on other syntax elements, such as the block size of the base layer and enhancement layer blocks, and the sampling resolution of the base layer and enhancement layer images. When the syntax element structure is implemented as a long-term syntax element structure, the sixth aspect is closely related to the third aspect. Therefore, the inter-layer offset indicates a guarantee to the decoder: this guarantee also applies to a predetermined time period larger than the time interval, within which the short-term syntax element indicators in the multi-layer video data stream are referred to as necessary. These hints are used to determine the inter-layer offset by combining these syntax elements in a relatively complex manner.
[0013] Naturally, all of the above aspects can be combined in two, three, four, or all of them. Attached Figure Description
[0014] Preferred embodiments of this application will now be described with reference to the accompanying drawings, in which:
[0015] Figure 1 A video encoder is shown as a schematic embodiment for implementing any of the multilayer encoders further summarized with reference to the following figures;
[0016] Figure 2 The image shows the installation to... Figure 1 A schematic block diagram of the video decoder in the video encoder;
[0017] Figure 3 A schematic diagram showing the image subdivided into sub-streams for WPP processing is provided.
[0018] Figure 4 The video decoder according to an embodiment is illustrated schematically, employing inter-layer alignment of spatial segments of the base layer and enhancement layer according to this embodiment to reduce decoding processing;
[0019] Figure 5The diagram illustrates the subdivision of an image into code blocks and slices, where slices consist of an integer multiple of code blocks and the decoding order between the code blocks follows the subdivision of the image into slices.
[0020] Figure 6 The following diagram illustrates the implementation. Figure 4 Syntax examples of the implementation methods in the document;
[0021] Figure 7 This diagram illustrates how a pair of base layer and enhancement layer images are subdivided into distinct slices.
[0022] Figure 8 The combination is shown Figure 4 Another exemplary syntax applicable to the implementation methods described above;
[0023] Figure 9 The diagram shows an image that is subdivided into slices and an application of interpolation filters to perform upsampling for the sake of interlayer prediction;
[0024] Figure 10 A schematic block diagram of a multi-layer decoder is shown, which is configured to enable or disable upsampling interpolation in response to syntax elements in a multi-layer data stream.
[0025] Figure 11 A schematic diagram of a pair of base and enhancement layers using interlayer prediction from the base layer to the enhancement layer is shown, thereby utilizing upsampling to convert the base layer sampling resolution to the increased enhancement layer sampling resolution;
[0026] Figure 12 The display shows according to Figure 10 A schematic diagram of switchable upsampling and interpolation separation;
[0027] Figure 13 A schematic diagram showing the overlap of the base layer image and the enhancement layer image is shown, and the base layer image and the enhancement layer image are subdivided into WPP sub-streams;
[0028] Figure 14 The following diagram illustrates the implementation. Figure 10 Exemplary syntax for implementation methods in the document;
[0029] Figure 15A This diagram illustrates the spatial overlap and alignment of the base layer image and the enhancement layer image, with the base layer image and the enhancement layer image subdivided into different spatial segments.
[0030] Figure 15B It shows Figure 15A The diagram illustrates the overlap of the base layer image and the enhancement layer image, but it also shows another possibility for selecting the partition along which to perform upsampling interpolation separation;
[0031] Figure 16 A schematic block diagram of a video decoder according to an embodiment is shown. The video decoder according to this embodiment responds to the long-term syntax element structure within a multi-layer video data stream to derive, or not derive, a guarantee regarding the inter-layer offset between the base layer and the enhancement layer from the decoding therein.
[0032] Figure 17A A schematic diagram of a pair of base layer and enhancement layer images is shown, and the pair of base layer and enhancement layer images are subdivided into pieces to illustrate the implementation according to the embodiment. Figure 16 Interlayer offset signals are transmitted based on the long-term syntax element structure within the data.
[0033] Figure 17B A schematic diagram of a pair of base layer and enhancement layer images is shown, and the pair of base layer and enhancement layer images are subdivided into sub-streams for WPP processing, to implement the embodiment. Figure 16 An example of a long-term syntax element structure in the text is described;
[0034] Figure 17C A schematic diagram of a pair of base layer images and enhancement layer images is shown, and the pair of base layer and enhancement layer images are subdivided into slices for implementation according to or even further embodiments. Figure 16 An example of a long-term syntax element structure in the text is described;
[0035] Figure 18 A schematic diagram of an image subdivided into substreams for WPP processing is shown, and additionally, the wavefront results when the image is decoded / encoded in parallel using WPP according to an embodiment are indicated.
[0036] Figure 19 A graph showing the relationship between the minimum inter-layer decoding offset and the block size, as well as the sampling resolution ratio between the base layer and the enhancement layer, according to an embodiment of this application, is shown.
[0037] Figure 20 It shows the method for implementing according to Figure 16 An example syntax for long-term syntax element structure signaling;
[0038] Figure 21 The following diagram illustrates the implementation. Figure 16 Another embodiment of the syntax of the implementation method in;
[0039] Figure 22 The syntax of the NAL unit header is shown in an embodiment configured according to the HEVC class.
[0040] Figure 23A schematic block diagram of a network entity according to an embodiment is shown, which reduces scalable coding according to this embodiment by allowing switching between interpolations of different layer indicator fields;
[0041] Figure 24 A diagram illustrating how the display of the response type indicator field is toggled;
[0042] Figure 25 A schematic diagram showing the interpolation of the switchable layer indicator field according to the embodiment is shown in further detail;
[0043] Figure 26 It shows the method for implementing according to Figure 23 The indicator syntax for interpolating switchable layer indicator fields;
[0044] Figure 27 The display shows the relationship with Figure 26 A diagram illustrating the switching of layer indicator fields related to syntax in the code;
[0045] Figure 28 A block diagram of a transport stream decoder configured to simplify codecs other than the base layer codec by dropping enhancement layer NAL units is shown.
[0046] Figure 29 A block diagram of a transport stream decoder connected to a single standard multilayer decoder is shown, illustrating the behavior of the transport stream decoder according to an embodiment.
[0047] Figure 30 A transport stream decoder connected to a multi-standard multi-layer decoder is shown, as well as the behavior of the transport stream decoder according to an embodiment;
[0048] Figure 31 Another embodiment of the syntax for implementing the interpretable layer indicator field according to a further embodiment is shown;
[0049] Figure 32 The illustration shows an image displayed at any layer that is subdivided into blocks, where blocks indicate that the image is further subdivided into spatial segments;
[0050] Figure 33 This is a schematic diagram showing any layer of an image that is subdivided into blocks and slices;
[0051] Figure 34 A schematic diagram of an image subdivided into blocks and sub-streams is shown;
[0052] Figure 35A schematic block diagram of a video decoder according to an embodiment is shown. The video decoder according to this embodiment is configured to use a syntax element structure in a data stream to derive inter-layer offsets between inter-predictive processed images in a cell having a raster scan order defined therein.
[0053] Figure 36 The display shows Figure 34 A schematic diagram of possible modes of operation of the video decoder related to the syntax element structure within the data stream according to the implementation method;
[0054] Figure 37 The illustration shows a combination of further embodiments. Figure 35 A schematic diagram of the operation mode of the video decoder in the embodiment, wherein the inter-layer offset signaling according to the further embodiment can switch between different explicit signaling types, that is, signaling in different types of units;
[0055] Figure 38 The following are illustrated according to, or even further embodiments: Figure 35 A schematic diagram of the operation mode of the video decoder in the diagram, continuously measuring the inter-layer offset according to or even further implementations during the parallel decoding of base layer and enhancement layer images;
[0056] Figure 39 This illustrates the relationship between the rank of a specific block in an image based on the raster scan decoding order and the row and column indices based on the implementation method.
[0057] Figure 40 Different embodiments for relatively regularly subdividing base layer images and enhancement layer images into blocks are shown, along with the results produced by these different embodiments;
[0058] Figure 41 The following diagram illustrates the implementation. Figures 35 to 40 Examples of the syntax of any of the implementation methods in;
[0059] Figure 42 Alternatives were shown Figure 41 Another grammatical example of the syntax in [the document / the document];
[0060] Figure 43 A grammar example is shown, according to which the grammar example is shown. Figure 16 and Figure 35 The implementation method described herein can transmit signals in another part of a multi-layer data stream; and
[0061] Figure 44 A schematic block diagram of a transport layer decoder connected to a multi-layer multi-standard decoder according to an embodiment is shown. Detailed Implementation
[0062] First, as an overview, an embodiment of an encoder / decoder architecture is presented, which is adaptable to any subsequently proposed concept.
[0063] Figure 1 The overall structure of the encoder according to an embodiment is shown. The encoder 10 can be implemented to operate in a multi-threaded manner or not just a single-threaded manner. That is, the encoder 10 can, for example, be implemented using multiple CPU cores. In other words, the encoder 10 can support parallel processing, but it is not necessary for the encoder 10 to necessarily support parallel processing. A generated bitstream can also be generated / decoded using a single-threaded encoder / decoder. However, the encoding concept of this application can support parallel processing of the encoder, thereby effectively applying parallel processing without compromising compression efficiency. Regarding parallel processing capabilities, refer to the following references. Figure 2 The decoder described is similar to the valid declaration.
[0064] Encoder 10 is a video encoder, but it can also typically be an image encoder. Image 12 of video 14 is shown as input to encoder 10 at input 16. Image 12 illustrates a specific context, i.e., image content. However, simultaneously, encoder 10 also receives another image 15 at its input 16, which is related to both image 12 and image 15, which belong to different layers. For illustrative purposes only, image 12, belonging to layer 0, is shown, while image 15 belongs to layer 1. Figure 1 It is shown that layer 1 may involve a higher spatial resolution than layer 0, i.e., a higher number of image samples can be used to display the same scene. However, this is only for illustrative purposes, and alternatively, image 15 of layer 1 may have the same spatial resolution, and may be different, for example, may have different spatial resolutions in the field of view direction relative to layer 0, i.e., images 12 and 15 may be captured from different viewpoints.
[0065] Encoder 10 is a hybrid trained encoder, meaning that predictor 18 predicts images 12 and 15 and the predicted residuals 20 obtained by residual determiner 22 undergo transformations such as spectral analysis (DCT) and quantization by transformation / quantization module 24. The resulting transformed and quantized predicted residuals 26 are then entropy encoded in entropy encoder 28, such as arithmetic encoding or using variable length with context adaptability. Decoder can use a reconstructible version of the residuals, i.e., recovering the dequantized and reconverted residual signals 30 by reconversion / requantization module 31 and recombining the dequantized and reconverted residual signals 30 with the predicted signal 32 of predictor 18 by combiner 33, thereby producing reconstructed images 34 for images 12 and 15 respectively. However, encoder 10 is based on block operations. Therefore, the reconstructed signal 34 encounters discontinuities at the block boundary lines, and thus, filter 36 can be applied to the reconstructed signal 34 to generate reference images 38 for images 12 and 15 respectively. The predictor 18 then sequentially predicts coded images of different layers based on the reference image 18. However, as Figure 1 As shown by the dashed line, predictor 18 can also directly use the reconstructed signal 34 without filter 36 or intermediate type in other prediction modes such as spatial prediction mode.
[0066] Predictor 18 can select different prediction modes to predict specific blocks in image 12. Figure 1 The example shows a block 39 of image 12. A temporal prediction mode may exist, according to which image 12 is partitioned into blocks 39 representing any part of image 12, predicting blocks 39 based on previously encoded images such as image 12' at the same layer. A spatial prediction mode may also exist, according to which blocks 39 are predicted based on previously encoded portions (adjacent blocks 39) of the same image 12. Figure 1 The diagram also schematically shows block 41 of image 15, used to represent any other block that partitions image 15. For block 41, predictor 18 may support the prediction modes just discussed, namely, temporal prediction mode and spatial prediction mode. Furthermore, predictor 18 may provide an inter-layer prediction mode, according to which block 41 is predicted based on corresponding portions of lower-layer image 12. The term "corresponding portion" should indicate spatial correspondence, i.e., predicting a portion within image 12 in image 15 that displays the same scene as block 41.
[0067] Of course, the predictions of predictor 18 are not limited to image sampling. Prediction can also be applied to any encoding parameters, i.e., prediction mode, motion vectors predicted in time, disparity vectors predicted in multi-viewpoints, etc. Only residuals can be encoded in bitstream 40.
[0068] A specific syntax is used to compile the quantized residual data 26, transform the coefficient levels and other residual data, and encoding parameters, including, for example, the prediction modes and prediction parameters for the separate block 39 for image 12 and the separate block 41 for image 15 determined by predictor 18, and to entropy encode the syntax elements through entropy encoder 28. The resulting data stream 40, output by entropy encoder 28, forms the bit stream 40 output by encoder 10.
[0069] Figure 2 It shows the adaptation to Figure 1 The encoder's decoder, i.e., capable of decoding bit stream 40. Figure 2 The decoder in the diagram is represented by reference symbol 50 and includes an entropy decoder, a conversion / dequantization module 54, a combiner 56, a filter 58, and a predictor 60. The entropy decoder 52 receives the bitstream and performs entropy decoding to recover the residual data 62 and encoding parameters 64. The conversion / dequantization module 54 dequantizes and reconverts the residual data 62 and forwards the resulting residual signal to the combiner 56. The combiner 56 also receives a prediction signal 66 from the predictor 60, and then uses the encoding parameters 64 to form the prediction signal 66 based on the reconstructed signal 68 determined by the combiner 56 by combining the prediction signal 66 and the residual signal 65. This prediction mirrors the prediction ultimately selected by the predictor 18; that is, the same prediction mode can be used, and these modes are selected for individual blocks of images 12 and 15 and imported according to the prediction parameters. (See reference above.) Figure 1 As already explained, alternatively or additionally, the predictor 60 may use a filtered form of the reconstructed signal 68 or some intermediate form thereof. Similarly, the different layers of images ultimately reproduced and output at the output 70 of the decoder 50 may be determined based on the unfiltered form of the combined signal 68 or some filtered form thereof.
[0070] Based on the slice concept, image 12 is subdivided into slice 80 and image 15 into slice 82, respectively. The prediction of block 39 within slice 80 and the prediction of block 41 within slice 82, as the basis for spatial prediction, are respectively limited to using only data related to the same slices of the same images 12 and 15. That is, the spatial prediction of block 39 is limited to using previously encoded portions of the same slice; however, the temporal prediction mode is not limited to relying on information from previously encoded images such as image 12'. Similarly, the spatial prediction mode of block 41 is limited to using only previously encoded data of the same slice; however, the temporal prediction mode and inter-layer prediction mode are unrestricted. For illustrative purposes only, images 15 and 12 are subdivided into six slices, respectively. The subdivided slices can be selected individually within bitstream 40 for images 12', 12 and 15, 15', and signals can be transmitted. The number of slices in each image 12 and the number of slices in each image 15 can be any one, two, three, four, six, etc., where slice partitioning can be limited to simply regularly partitioning the rows and columns into slices. For completeness, it should be noted that the method of encoding slices independently is not limited to intra-frame prediction or spatial prediction, but can also include any prediction of coding parameters across slice boundaries, and the context selection in entropy coding can also be limited to data that depends only on the same slice. Therefore, the decoder can perform the operations just mentioned in parallel (i.e., slice-by-slice).
[0071] Alternatively or additionally, Figure 1 and Figure 2 The encoder and decoder in the code can use WPP concepts. See also Figure 3WPP substream 100 also represents spatially partitioning images 12 and 15 into WPP substreams. Unlike slices and slices, WPP substreams do not impose restrictions on prediction and context selection across WPP substream 100. WPP substream 100 is expanded line by line, such as across lines of LCU (Maximum Coding Unit) 101, i.e., the largest possible block of predictive coding patterns that can be transmitted individually in the bitstream, and only one compromise related to entropy coding is made in order to support parallel processing. Specifically, sequence 102 is defined along the WPP substream 100, which is exemplarily guided from top to bottom, and for each WPP substream 100, the probability evaluation of the symbol Arabic letters, i.e., the entropy probability, is not completely reset except for the first WPP substream in sequence 102, but is instead adopted or set to be equal to the probability generated after entropy coding / decoding. Before the preceding WPP substream reaches its second LCU, as shown by line 104, for each WPP substream located on the same side of images 12 and 15, it starts from the left-hand side, as indicated by arrow 106, according to the LCU order or the decoder order of the substream, and is guided to the other side in the LCU row direction. Therefore, by obeying a certain encoding delay between the sequences of WPP substreams of the same images 12 and 15 respectively, these WPP substreams 100 can be decoded / encoded in parallel, such that the portions of the corresponding images 12, 15 that are encoded / decoded in parallel (i.e., simultaneously) form a certain wavefront 108, which moves in a slanted manner from left to right on the images.
[0072] Briefly, sequence 102 also defines the raster scan order within the LCU, thus guiding row by row from the top left LCU 101 to the bottom right LCU from top to bottom. Each WPP substream can correspond to one LCU row. Briefly, returning to the reference slice, the latter is also constrained by alignment with the LCU boundaries. The substream can be segmented into one or more slices without being constrained by the LCU boundaries, as long as the boundary between two slices within the substream is involved. However, in the case of transitioning from one slice of the substream to the next slice of the substream, entropy probability is used. If it is a slice, all slices can be combined into one slice, or a slice can be segmented into one or more slices, without needing to be constrained by the domain LCU boundaries again, as long as the boundary between two slices within the slice is involved. If it is a slice, the order between LCUs is changed by first transecting the slices according to the slice order in the raster scan order before proceeding to the next slice in slice order.
[0073] As described so far, image 12 can be partitioned into slices or WPP substreams, and similarly, image 15 can be partitioned into slices or WPP substreams. Theoretically, a WPP substream partition / concept can be chosen for one of images 12 and 15, while a slice partition / concept can be chosen for the other of these two images. Alternatively, a constraint can be imposed on the bitstream, according to which the concept type (i.e., slice or WPP substream) of each layer must be the same. Another embodiment of spatial segments includes slicing. For transmission purposes, slicing is used to divide bitstream 40 into segments. Slices are packaged in NAL units, which are the smallest entities used for transmission. Each slice can be encoded / decoded independently. That is, as with context selection, any prediction across slice boundaries is prohibited. In summary, these are three embodiments of spatial segments: slices, slices, and WPP substreams. Furthermore, all three parallel concepts—slices, WPP substreams, and slices—can be combined; that is, image 12 or image 15 can be partitioned into slices, where each slice is divided into multiple WPP substreams. Furthermore, slicing can be used to partition the bitstream into multiple NAL units, for example, but not limited to, partitioning at slice or WPP boundary lines. If slices or WPP substreams are used and slices are additionally used to partition images 12 and 15, and there is a discrepancy between the slice partition and another WPP / slice partition, then the spatial segment should be limited to the smallest independently decodable segment of images 12 and 15. Alternatively, constraints can be imposed on the bitstream by utilizing concept combinations within images (12 or 15) and / or, if it is necessary to align the boundaries between concepts used differently.
[0074] Before discussing the concepts described above in this application, please refer again to Figure 1 and Figure 2 It should be noted that Figure 1 and Figure 2 The encoder and decoder box structures in the diagram are for illustrative purposes only and the structures can be different.
[0075] According to the first aspect, referred to as "piece boundary alignment," the long-term syntax element structure is used to send signals to ensure that, within a predetermined time period, such as the time period of a series of picture extensions, the second-layer picture 15 is subdivided such that the boundaries 84 between the spatial segments 82 of the second-layer picture overlap with each boundary 86 between the spatial segments 80 of the first layer. In shorter time intervals than the predetermined time period (e.g., on a per-picture basis), i.e., within the picture-interval interval, the decoder still periodically determines, based on the short-term syntax elements of the multi-layer video data stream 40, the actual subdivision of the first-layer picture 12 into spatial segments 80 and the actual subdivision of the second-layer picture 15 into spatial segments 82; however, knowledge of alignment has already helped in planning the allocation of parallel processing workloads. For example, Figure 1The solid line 84 in the diagram represents an embodiment where slice boundary line 84 is spatially aligned with slice boundary line 86 of layer 0. However, the aforementioned guarantee also allows slice partitions of layer 1 to be finer than those of layer 0, such that slice partitions of layer 1 further include additional slice boundary lines that do not spatially overlap with any of the slice boundary lines 86 of layer 0. In any case, knowledge of slice registration between layer 1 and layer 0 helps the decoder allocate available workload or processing power among spatial segments processed in parallel simultaneously. Without a long-term syntax element structure, the decoder would have to perform workload allocation in shorter time intervals, i.e., per picture, thus wasting computing power that would otherwise be used to perform workload allocation. On the other hand, there is “opportunistic decoding”: a decoder with multiple CPU cores can use knowledge of layer parallelism to decide whether to attempt to decode layers of higher complexity (i.e., higher spatial resolution or a higher number of layers). Bitstreams exceeding the capabilities of a single core can be decoded by utilizing all cores of the same decoder. This information is particularly useful if the profile and level indicator do not contain indications of minimum parallelism.
[0076] To better understand the aspects just outlined in this application, please refer to... Figure 4 , Figure 4 References are shown Figure 2 The video decoder 600 is achievable according to the specifications. As described above, the decoder 600 is configured to decode multi-layer video data streams into scenes, which are encoded in the layer hierarchy using inter-layer predictions from the first layer 0 to the second layer 1 as described above. The video decoder supports parallel decoding of multi-layer video data streams by subdividing the images of that layer into spatial segments, such as slices, WPP substreams, etc. In other words, the video decoder is capable of parallel decoding of multi-layer video data streams; thus, the video decoder 600 operates on images 12 of layer 0 and images 15 of layer 1, on a spatial segment basis.
[0077] As summarized above, for example, spatial segments can be slices, and video decoder 600 is configured to decode image 12 of layer 0 and image 15 of layer 1 using intra-frame picture spatial prediction. Video decoder 600 causes the intra-frame picture spatial prediction for each slice to be interrupted at its slice boundary. For example, individually, for each time frame 604 associated with each image 12 and 15, i.e., for each pair of images 12 and 15 belonging to a particular time frame 604, a signal is sent within data stream 40 based on a short-term syntax element: to subdivide images 12 and 15 into slices, such as in time intervals. As mentioned above, i.e., subdividing images 12 and 15 into slices can be limited to rows and columns that are subdivided into slices only in a rectangular regularity. Therefore, short-term syntax element 602 will individually set the number of rows and columns for slice subdivision of each image 12 and each image 15 for both layers. When decoding the inbound multi-layer video data stream 40, video decoder 600 is configured to apply spatial prediction, potentially temporal prediction. Optionally, the video decoder 600 performs entropy decoding on each slice separately. If probabilistic adaptability is used during the decoding of each slice, the video decoder 600 initializes the entropy probability of each slice separately to enable parallel entropy decoding of the slices. In addition to spatial prediction and optionally temporal prediction, the video decoder 600 supports inter-layer prediction whenever decoding of slices of image 15 in layer 1 is involved. As described above, inter-layer prediction can involve different parameters involved in decoding layer 1: inter-layer prediction can predict the prediction residuals of layer 1, such as transformation coefficients, prediction modes used in decoding layer 1, prediction parameters used in decoding layer 1, enhancement layer 1, and image sampling. For example, in the case where layers 0 and 1 involve different views of the same scene, by controlling the disparity vector prediction parameters of inter-layer prediction, inter-layer prediction predicts the portion within a slice of image 15 in layer 1 based on the decoded portion of image 12 in layer 0 (either the directly co-located portion or the portion slightly offset spatially from the directly co-located position).
[0078] If used Figure 4As indicated by reference numeral 606, the video decoder 600 responds to the long-term syntax element structure of the data stream 40 to process, to varying degrees, a predetermined time period 608 immediately following the long-term syntax element structure 606. The predetermined time period 608 comprises several time intervals, i.e., multiple time frames 604, within which short-term syntax elements 602 individually transmit signals to subdivide the image into pieces. It should be noted that 608 may involve the range (=time period) and SPS variation that could cause significant reinitialization in any way. The just-mentioned notes are valid for all embodiments relating to other aspects, as long-term characteristics are mentioned herein. Specifically, if the long-term syntax element structure 606 assumes values outside the first set of possible values, the video decoder 600 interprets this situation as guaranteeing that, within the predetermined time period, the image 15 of layer 1 is subdivided such that the boundaries between pieces of image 15 overlap with each boundary of a piece of image 12 of layer 0. In this scenario, the video decoder 600 still detects the short-term syntax element 602 to determine the subdivision of images 12 and 15 into their slices within the time interval 602 of the predetermined time period 608. However, by comparing each time-aligned pair of images 12 and 15, the video decoder 600 can rely on the fact that the boundary of the base layer slice of image 12 completely overlaps with the boundary of the enhancement layer slice of image 15; that is, the slice subdivision of image 15 locally corresponds to or represents the spatial refining of image 12 into slices. As described above, by correspondingly scheduling the parallel processing of slices of image 12 and slices of image 15 within the predetermined time period 608—that is, decoding time-aligned pairs of images 12 and 15 in parallel—the video decoder 600 can take advantage of the signaling that the long-term syntax element structure 606 assumes values outside the first set of possible values. For example, if the long-term syntax element structure assumes values outside the first set of possible values, the video decoder 600 can know that for a specific image 12 of layer 0, the first slice among the slices of image 12 in slice order locally coincides with the corresponding slice of the time-aligned enhancement layer image 15, or completely locally overlaps with the first slice of the time-aligned enhancement layer image 15 among the slices of the enhancement layer image 15 in slice order. Therefore, at least in the case of inter-layer prediction without parallax / motion compensation, because the aforementioned guarantee indicates to the video decoder 600 that the co-occurring portion of the base layer image 12 required for inter-layer prediction is available for the entire first slice of the enhancement layer image 15, the video decoder 600 can begin decoding the first slice of the enhancement layer image 15 as soon as possible after completing the decoding of the first slice of the time-aligned base layer image 12. Thus, the video decoder 600 can recognize / determine that the inter-layer offset or degree of parallelism between the base layer image 12 and the enhancement layer image 15 is equal to a slice of the base layer image 12.If the interlayer prediction contains a disparity vector with a non-zero vertical component and / or a disparity vector with a horizontal component, the horizontal component will shift the corresponding portion within the base layer image to the right, which may slightly increase the offset. The slice order between slices may be guided by the row raster scan order from the upper left corner to the lower right corner of images 12 and 15.
[0079] However, if the long-term syntax element structure assumes values outside the second set of possible values, which is significantly different from the first set of possible values, the video decoder 600 does not take advantage of any guarantees. Instead, it plans and schedules the parallel decoding of slices of images 12 and 15 based on the short-term plan using the short-term syntax element 602, and potentially, for at least some of the time-aligned pairs of images 12 and 15, it plans and schedules the parallel decoding of slices of the base layer and enhancement layer. However, in this case, the video decoder 600 determines the minimum inter-layer offset or inter-layer spatial processing offset, i.e., the degree of parallelism between layers 0 and 1, in the parallel decoding between layers 0 and 1 based on the short-term (cumbersome procedure). At least for a subset of the set of possible values of the short-term syntax element, there are boundaries between spatial segments of images in the second layer that do not overlap with any of the boundaries of spatial segments in the first layer. However, according to a further subset of the set of possible values of the short-term syntax element, there are boundaries between spatial segments of images in the second layer that overlap with each boundary of a spatial segment in the first layer. If the long-term syntax element indicates slice boundary alignment between the base layer and enhancement layer, only the latter subset is used.
[0080] Alternatively, the video decoder 600 may use or employ the following facts: the long syntax element structure assumes a value outside the first set of possible values to perform an experiment (i.e., attempt to fully decode layer 1), and if the long syntax element structure 606 assumes a value outside the second set of possible values, then the execution of this experiment is suppressed. In this case, especially for battery-powered devices, valuable computational power is saved when the result or success rate of timely (i.e., real-time) decoding of enhancement layer 1 is speculative. It is worth mentioning that refraining can also be selected based on the layer indicator of the volume referred to in the fourth aspect below.
[0081] Although the above uses slices as spatial segments to describe the example Figure 4However, it is clear that the video decoder 600 can combine other spatial segments, such as substreams or slices, to leverage the advantages of the Long Term Syntax Element (LTS) structure and the guarantees derived therefrom. In the former case, the video decoder 600 will use intra-frame picture spatial prediction to decode images 12 and 15 of the layer, and will decode the spatial segments of the images of the first layer 12 in parallel, and will support intra-frame picture spatial prediction across the boundaries of the spatial segments of the images of the first layer, and will be subject to the decoding latency between the decoding of these spatial segments (i.e., substreams). As mentioned above, a substream can correspond to a horizontal stripe of the corresponding image, that is, vertically subdividing the corresponding image. When decoding each substream, the video decoder 600 can use a decoding order that is generally left-guided to right, and the decoding order defined between the substreams of the image can be top-guided to bottom. According to the typical spatial prediction concept used, spatial prediction is performed from the top adjacent decoded portion and the left-hand decoded portion of the current image, subject to a specific decoding latency between the subsequent substreams, thereby allowing parallel decoding of substreams. For example, the decoding latency can be measured within a cell of the LCU. This operation can be performed in image 12 of layer 0 and image 15 of layer 1. Therefore, the parallelism in decoding the video data stream can individually include both parallelism within images 12 and 15, but it can also include parallel decoding of substreams of images 12 and 15 belonging to different layers of a time frame 604. Whenever optional entropy decoding of a substream is involved, the parallelism can include the adaptation of entropy probabilities during the decoding process of the corresponding substream. The first substream in the substream sequence of each image 12 or image 15 can be initialized with entropy probabilities independently of other substreams. Based on the instantaneous adaptive entropy probability of the immediately preceding substream in the decoding sequence of the same image, such as by adopting the entropy probability used in decoding the immediately preceding substream that reaches a specific distance to the left-hand side of the corresponding preceding substream, such as after decoding two LCUs of the immediately preceding substream, any subsequent substream can be initialized with entropy probabilities.
[0082] Even when wavefront parallel processing of substreams is employed, the video decoder 600 can leverage the advantages of the long-term syntax element structure 606: if a guarantee signal is issued through this syntax element structure 606, the video decoder 600 can rely on the fact that all boundaries between consecutive / adjacent substreams of the base layer image 12 overlap with the corresponding boundaries between adjacent / consecutive substreams of the time-aligned enhancement layer image 15 within a predetermined time period 608. That is, the base layer substreams are locally consistent with the corresponding enhancement layer substreams of the time-aligned enhancement layer image 15, or the base layer substreams precisely correspond to two or more substreams of the time-aligned enhancement layer image. Therefore, if the guarantee applies, the decoder 600 knows that after completing the decoding of the first substream of the base layer image 12, the decoding of the first substream of the time-aligned enhancement layer image 15 can begin as soon as possible.
[0083] As described above, unlike slice subdivision, the short-term syntax element 602 can be selected such that it defines the positions of substreams in pictures 12 and 15 related to subdividing these pictures into coding blocks such as LCUs. Therefore, a substream can be a collective term for one or more lines of coding blocks. Similar to slice subdivision, the time interval 604 can be as follows: the short-term syntax element 602 issues a signal to subdivide pictures 12 and 15 into substreams based on each picture (i.e., based on each picture frame 604). If the long-term syntax element structure 606 does not provide this guarantee, however, the video decoder 600 may attempt to decode substreams of different layers of a common event frame in parallel; however, to achieve this, the video decoder 600 needs to check the short-term syntax element 602.
[0084] If slices are used as spatial segments, the video decoder 600 can perform a speculative experiment to decode enhancement layer 1 based on the values assumed by the long-term syntax element structure 606.
[0085] It should be noted that, if in Figure 1 As shown, whenever the video encoder is involved, the long-term syntax element structure 606 can be inserted and set into the data stream 40, and a decision can be made on whether to grant the corresponding video encoder a guarantee to the decoder 600. If granted, the encoding probability is constrained when the short-term syntax element 602 is set to conform to the boundary alignment guarantee within the corresponding predetermined time period 608. If not granted, the encoder remains free to set the short-term syntax element 602 according to its own habits within the time period 608. When slices are used as spatial segments, the encoder conforms to the following constraints: spatial predictions cannot cross boundary lines, and optional entropy coding of the slices of pictures 12 and 15 is performed for each slice in a self-contained manner. For example, the entropy probability is reinitialized for slices that are not dependent on other slices (for each slice). If it is a substream, the entropy probability initialization of the substream is re-performed for any first substream of the corresponding pictures 12, 15, i.e., regardless of any other substream, and whenever any second substream and subsequent substreams are involved, the entropy probability is adopted when adapting to the middle position of the immediately preceding substream. Perform spatial prediction without any restrictions on the intersection of substreams.
[0086] By reference Figure 4 The alignment concept can be introduced into the currently envisioned extended HEVC standard in the manner described below. For now, the description to be presented below should also be interpreted as based on the above references. Figure 4 The proposed description includes possible implementation details.
[0087] HEVC allows the CTB of the encoded base layer image to be divided into rectangular regions called patches by a grid with vertical and horizontal boundaries, and these patches can be processed individually (except for in-loop filtering). At the patch boundaries, in-loop filtering can be turned off to make it completely independent.
[0088] Similar to the image boundary lines, the interdependence between resolution and prediction breaks at the slice boundary lines, where, if configured accordingly, in-loop filters can cross the boundary lines to reduce boundary line artifacts. Therefore, the processing of individual slices does not entirely depend on other slices within the image or on a wide range of filter configurations. The limitation of the setup is that all CTBs of a slice should belong to the same slice, or all CTBs of a slice should belong to the same slice. From Figure 1 As can be seen, the slices force the CTB scanning order to follow the slice order; that is, all CTBs belonging to the first slice (e.g., the top left slice) are traversed before continuing with CTBs belonging to the second slice (e.g., the top right slice). The structure is defined within each slice row and each slice column of the raster that makes up the image by the number and size of the CTBs. The structure can change based on each frame or remain unchanged within the encoded video sequence.
[0089] Figure 5 The image shows an example of dividing the CTB within it into nine slices. The thick black lines represent the boundary lines, and the numbers indicate the scan order of the CTBs, which also reflects the slice order.
[0090] The HEVC extended enhancement layer slices can be decoded by decoding all slices covering the corresponding image regions in the underlying layer bitstream.
[0091] The following section describes the permitted uses Figure 4 The concepts in the middle make it easier to access the constraints, signaling, and decoding processes of the base layer information.
[0092] The simplest case of patch-level parallelism is that the patch boundaries in the base layer and enhancement layer are aligned. For SNR scalability, this means the boundaries are precisely in the same location. For spatial scalability, this means that for every two enhancement layer pixels belonging to the same patch, the corresponding base layer pixels also belong to the same patch, and for every two base layer pixels belonging to the same patch, the corresponding enhancement layer pixels also belong to the same patch.
[0093] HEVC is characterized by, and Figure 4 The short signaling corresponding to 602 in the image uses the following set of image parameters: column_width_minus1[i] and row_height_minus1[i] from [1] to indicate the dimensions and structure of the slices within the image based on each image. Figure 6 An example syntax is shown.
[0094] A further feature of HEVC is the constraint signaling that guarantees specific settings for HEVC-coded video sequences, such as indicating a fixed tile structure in a single-layer HEVC-coded video sequence (cp.tiles_fixed_structure_flag in the VUI syntax given below). This further constraint on tiles in scalable coded video sequences benefits decoder initialization and operation. To allow the decoder to begin decoding enhancement layer image regions associated with active layer tiles after completing the base layer tiles, it is not necessary to enforce full alignment. This may help allow decoding of more tiles in enhancement layers than in base layers, especially in terms of spatial scalability. For example, with respect to both spatial scalability factors, the enhancement layer image region contains four times the number of pixels compared to the corresponding base layer image region. Therefore, for each base layer tile, it may be possible to allow decoding of four tiles in the enhancement layer. See Figure 7 , Figure 7 An embodiment with spatially scalable aligned sheet boundary lines is shown. All vertical boundary lines in the base layer and enhancement layer are aligned. In the enhancement layer, additional sheets (horizontal boundary lines) are used to allow parallelism in each enhancement layer using the same number of cells as when the base layer is sheet-partitioned.
[0095] Therefore, we define patch boundary alignment in the following way: only each base layer boundary has a corresponding boundary in the enhancement layer; however, the reverse is not possible. Specifically, this means that for every two enhancement layer pixels belonging to the same patch, the corresponding base layer pixels also belong to the same patch.
[0096] Signaling 606 helps initialize the parallel decoder environment; otherwise, it would have to acquire information by parsing multiple sets of parameters. Moreover, for example, in terms of bitstream constraint forms... Figure 4 The concept in the document ensures that the restrictions are effective for the complete encoded video sequence.
[0097] If the base layer's sheet boundary lines are a subset of the enhancement layer's sheet boundary lines, then possible implementations may allow the signaling of the base layer's sheet boundary lines to be stored in the enhancement layer.
[0098] Precisely, signals regarding slice alignment can be emitted from a bitstream easily accessible to the decoder.
[0099] In specific implementations, signaling can be achieved by using flags in the VUI parameters of the enhancement layer SPS, such as... Figure 8 The following is given:
[0100] When present, `tiles_fixed_structure_flag` equal to 1 indicates that each set of picture parameters (i.e., those active in the encoded video sequence) has the same value for the syntax elements `num_tile_columns_minus1`, `num_tile_rows_minus1`, `uniform_spacing_flag`, `column_width_minus1[i]`, `row_height_minus1[i]`, and `loop_filter_across_tiles_enabled_flag`. `tiles_fixed_structure_flag` equal to 0 indicates whether tile syntax elements in different sets of picture parameters can or cannot have the same value. When no `tiles_fixed_structure_flag` syntax element exists, it is inferred to be equal to 0.
[0101] It should be noted that the signaling of tiles_fixed_structure_flag equal to 1 guarantees to the decoder that each picture in the encoded video sequence has the same number of tiles distributed in the same way, which may help with workload distribution in the case of multi-threaded decoding.
[0102] tile_boundaries_aligned_flag corresponds to Figure 4 Structure 606 in the table. If `tile_boundaries_aligned_flag` is equal to 1, it indicates that all tile boundary lines of the corresponding base layer image have the corresponding tile boundary lines in the given enhancement layer. A `tile_boundaries_aligned_flag` equal to 0 indicates that there are no restrictions on the tile configuration between the corresponding base layer and the given enhancement layer.
[0103] It should be noted that the long-term syntax element structure guarantees that, within a predetermined time period, the minimum number of spatial segments 82 into which the second-level image 15 is subdivided (e.g., the image sequence) is n times the minimum number of spatial segments 80 into which the first-level image 12 is subdivided, or that each spatial segment of image 12 consists precisely of n spatial segments of time-aligned image 15, where n depends on the value of the long-term syntax element structure. Figure 7In the case of n equals 3. The decoder still periodically determines, based on short-term syntax elements of the multi-layer video data stream 40, whether to actually subdivide the first layer's image 12 into spatial segments 80 and the second layer's image 15 into spatial segments 82, within time intervals shorter than the predetermined time interval. However, again, the decoder can employ this guarantee to perform workload distribution more efficiently. On the other hand, there is "opportunistic decoding": a decoder with multiple CPU cores can employ the guarantee as an indication of the parallelism of the layers and decide accordingly to decode layers of higher complexity (e.g., higher spatial resolution or a higher number of layers). Bitstreams exceeding the capabilities of a single core can be decoded by utilizing all cores of the same decoder. This information is particularly helpful if the configuration file and layer indicator do not contain indications about minimum parallelism.
[0104] The second aspect discussed and proposed below concerns the concept known as “restricted inter-tile upsampling”: In the case of spatially scalable multi-layer video, upsampling filters 36 are adjusted using syntax elements in the bitstream (e.g., exemplarily, independent_tile_upsampling_idc). If upsampling filtering is performed in layer 0 across spatial segment boundaries 86, the latency satisfied by parallel decoding / encoding of spatial segment 82 of layer 1 relative to the encoding / decoding of spatial segment 80 of layer 0 increases due to the combination of upsampling filters, and thus makes the information of adjacent spatial segments of layer 0 interdependent, used as prediction reference 38 for inter-layer prediction of block 41 of layer 1. See, for example, Figure 9Images 12 and 15 are shown in an overlapping configuration, where the two images are registered to each other based on spatial correspondence and their dimensions are consistent, meaning that portions showing the same scene overlap each other. Images 12 and 15 are exemplarily shown as being divided into 6 spatial segments (such as slices) and 12 spatial segments, respectively. A filter kernel 200 is exemplarily shown moving on the upper-left slice of image 12 to obtain its upsampling configuration, which serves as the basis for inter-layer prediction of any block within a slice of image 15 (spatially overlapping with the upper-left slice). In some intermediate instances, such as 202, kernel 200 overlaps with adjacent slices of image 12. Therefore, the sampled value in the middle of kernel 200 at position 202 of the upsampling configuration depends on the sampling of the upper-left slice of image 12 and the sampling of its right-hand slice. If the upsampling configuration of image 12 is used as the basis for inter-layer prediction, the inter-layer offset increases during parallel processing of slices of the layer. Therefore, this constraint can help increase the amount of parallelism between different layers and correspondingly reduce the overall coding latency. Naturally, grammatical elements can also be long-term grammatical elements that are valid for image sequences. This restriction can be achieved by one of the following methods: filtering the overlapping portion of kernel 200 at overlapping position 202, for example, having a central tendency for sampled values to lie within the non-dashed portion of kernel 200; extrapolating the non-dashed portion to the dashed portion using a linear function or other function, etc.
[0105] To make the following points clearer, please refer to... Figure 10 , Figure 10 A decoder 610 is shown that receives spatially scalable bitstreams 40 (images encoded as spatially scalable bitstreams 40) from different spatial layers corresponding to image 12 in layer 0 and image 15 in layer 1. For at least one of these spatial layers, the decoder 610 is configured to decode that spatial layer within a spatial segment. As described above, these spatial segments can be slices, bitstreams, or slices. Similarly, the decoder 610 can be configured to perform parallel decoding of that spatial segment of image 12 or image 15. That is, the base layer image 12 can be subdivided into spatial segments such as slices and / or bitstreams and / or slices, and / or the enhancement layer image 15 can be subdivided into slices and / or bitstreams and / or slices. Refer above for details regarding parallel processing. Figure 4 The description in [the original text] has been moved to [the new text]. Figure 10Decoder 610. That is, for example, when decoding the base layer image 12, if base layers 12 and 15 are part of a layered video, decoder 610 uses spatial prediction and optionally temporal prediction. If it is a slice, spatial prediction is limited to not crossing slice boundaries and spatial prediction is applicable to entropy decoding performed entirely on the slice alone (if entropy decoding is used). Spatial prediction is applied to enhancement layer image 15 and additionally supports inter-layer prediction. As described above, inter-layer prediction is not only about the prediction parameters of the enhancement layer, based on the corresponding prediction parameters used when decoding the base layer, but also about the prediction derived from the reconstructed sampling of the base layer image, which is located in a co-location relative to the portion of enhancement layer image 15 currently predicted using inter-layer prediction. However, since bitstream 40 can be a spatially scalable bitstream, decoder 610 can upsample any co-location portion of the base layer image 12 that forms the basis of the inter-layer prediction of the current processing portion of enhancement layer image 15 to indicate a higher spatial resolution of image 15 relative to image 12. For example, see Figure 11 , Figure 11 Reference symbol 612 is used to denote the currently predicted portion of enhancement layer image 15. Reference symbol 614 is used to denote the co-occurring portion of base layer image 12. Due to the higher spatial resolution of enhancement layer image 15, the number of sampling locations within portion 612 (denoted by dots) is shown as higher than the number of sampling locations within portion 614 (also denoted by dots). Therefore, decoder 610 uses interpolation to upsample the reconstructed form of portion 614 of base layer image 12. Thus, Figure 10 The decoder 610 corresponds to the syntax element 616 in the spatially scalable bitstream 40.
[0106] Specifically, refer to Figure 12 This provides a more detailed explanation of the responsiveness just mentioned. Figure 12 A portion 614 within the base layer image 12 (i.e., a reference portion of its non-upsampled form) and its corresponding upsampled form, indicated by reference symbol 618, are shown. As just mentioned, for example, the form 618 used for interlayer prediction can be obtained from the base layer image 12 via interpolation 620 by copying the sampled values of the corresponding upsampled form in portion 618 into portion 612 of the enhancement layer 15. However, interpolation 620 depends on the syntax element 616 just mentioned. The way interpolation 620 changes according to syntax element 616 relates to the region along the boundary between adjacent partitions of the partitions of the base layer image 12 and its upsampled form. Specifically, the partitions depend on the aforementioned spatial segments into which at least one of images 12 and 15 is subdivided. Figure 11In the diagram, a partition within the base layer image 12 is shown using dashed line 622. For example, as summarized in more detail below, partition 622 may correspond to a logical AND or logical OR combination of the spatial overlap of the spatial segments of images 12 and 15, or spatially coincide with a partition defined by the spatial segment of the enhancement layer image 15. In any case, decoder 610 performs interpolation 620 according to syntax element 616, regardless of or taking into account partition 622. When considering partitions, decoder 610 performs interpolation 620 such that all samples within upsampling portion 618 originate only from, depend on, or are influenced by samples from one of the partitions of partition 622, and are independent of any other partition of partition 622. For example, if partition 622 is a local AND combination or a local OR combination of the spatial segments of images 12 and 15, then all samples in interpolation portion 618 originate only from one partition of base layer image 12. However, if syntax element 616 guides decoder 610 to be insensitive to partition 622, it is possible that different samples within interpolation portion 618 originate from adjacent partitions of partition 622.
[0107] For example, 612 illustrates the following situation: interpolation 620 is performed using filter kernel 200, and in order to obtain Figure 12 The interpolated samples enclosed in the middle overlap the boundaries between two adjacent partitions of kernel 624 and partition 622. In this case, decoder 610 responds to syntax element 616 to properly fill filter kernel 624, i.e., by applying filter kernel 624 fully to the corresponding included sample of base layer image 12, or by using a fallback rule (such as...). Figure 12 The hash used (as shown) fills the segments of filter kernel 624 that extend into adjacent partitions (excluding portions 614 and 618, respectively), according to a backoff rule, filling corresponding segments that are not dependent on the base layer samples of the base layer image 12. For example, the hash portion of filter kernel 624 is filled using some average measurement of the sampled values of the non-hash portion or by some extrapolation. Alternatively, the sampled values of the base layer image 12 that overlap with the hash portion are filled with a predetermined value such as 0. Typically, decoder 610 can handle partition boundaries that separate the portion containing portion 614 from adjacent partitions (such as the outer edge of image 12 itself), and can use the same backoff rule as in interpolation 620 used, for example, near the outer circumference of image 12 or when performing upsampling / interpolation.
[0108] According to one embodiment of this application, partition 622 is selected such that it is consistent with the spatial segmentation of the base layer image and is independent of the arbitrary spatial segmentation of the enhancement layer image. Therefore, since portions such as portion 614 of the base layer image 12 do not require the decoder 610 to decode adjacent partitions / spatial segments before performing inter-layer prediction of portion 612 of the enhancement layer image 15, the inter-layer offset between the base layer image 12 and the enhancement layer image 15 can be reduced by the decoder 610.
[0109] Alternatively, decoder 610 may be configured to determine that partition 622 is locally consistent with the spatial segmentation of image 15. Another alternative, decoder 610 may be configured to select partition 622 consisting only of the boundaries of the spatial segments of images 12 and 15, partition 622 being spatially consistent, i.e., with local AND corresponding to the boundaries of images 12 and 15; in other words, only these boundaries of image 15 are subdivided into spatial segments corresponding to the boundaries between partitions of partition 622, i.e., spatially corresponding to the corresponding boundaries of the base layer image 12 are subdivided into spatial segments.
[0110] Syntax element 616 can guide decoder 610 to not only ignore partition 622 in interpolation 620, but also distinguish between different ways of selecting partition 622, which is also feasible and will be summarized in more detail below. For example, see Figure 9 In this example, slices are used as spatial segments. For instance, if syntax element 616 signals decoder 610 to perform interpolation 620 separately for partition 622, decoder 610 can use the boundary of base layer image 12 as the boundary of partition 622 because partition 622 corresponds to the refined subdivision of enhancement layer image 15 into slices. Therefore, in order to begin decoding the second slice in the highest slice row of enhancement layer image 15, decoder 610 does not need to wait for the decoding of the second slice in the highest row of base layer image 12 to be completed, because "interpolation separation" prohibits arbitrary mixing of reconstruction samples of the first two slices in the highest row of base layer image 12. If the sampling completely subdivides enhancement layer image 15 into slices, in order to determine partition 622, it is also... Figure 9 Interpolation separation is performed at the dashed lines in the diagram, and decoder 610 can begin decoding the upper left slice of enhancement layer image 15 even earlier; that is, decoder 610 manages the decoding of the corresponding co-located sub-section of the first slice of base layer image 12 as quickly as possible. In view of this, it should be noted that even when decoding slices, decoder 610 can use a certain decoding order, for example, a raster scan order that guides the decoder from the upper left corner to the lower right corner of the corresponding slice in a row-like manner.
[0111] That is, according to the second aspect, the encoder forming bitstream 40 can select between two modes via syntax element 616: if syntax element 616 is set and inserted into bitstream 40, the decoder 610 is insensitive to partition 622, and better inter-layer prediction can be achieved due to better interpolation; however, the degree of parallelism achievable when decoding images 12 and 15 in parallel is reduced, i.e., the minimum inter-layer offset required is increased. In the other mode, syntax element 616 guides the decoder 610 to consider partition 622 for the purpose of inter-layer prediction when performing interpolation 620, and therefore, the quality of inter-layer prediction is reduced, which is beneficial for increasing the degree of parallelism and reducing the minimum inter-layer decoding offset when decoding images 12 and 15 in parallel.
[0112] Although the description of the second aspect of this application focuses primarily on the concept of slice subdivision or slice parallel processing, it is clear that when using WPP bitstreams, it is also advantageous to use syntax element 616 to control interpolation 620. For example, see Figure 13 , Figure 13 The following scenario is illustrated: the base layer image 12 is exemplarily subdivided into two bitstreams, wherein the co-occurring portions of the enhancement layer image 15 are each subdivided into two bitstreams. If interpolation separation is applied to the response syntax element 616, the decoder 610 can begin decoding the first (i.e., highest) bitstream of the enhancement layer image 15. Once the decoder 610 has decoded the first bitstream of the base layer image 12, it sufficiently covers the upper-left portion of the response of the first enhancement layer bitstream of image 15. And because interpolation separation makes any interlayer prediction independent of any reconstructed portion of the base layer bitstream of image 12 that spatially overlaps with the second enhancement layer bitstream, this also applies even to these portions of the first enhancement layer bitstream of image 15 that overlap with the boundary of the second bitstream of image 15.
[0113] Below is a detailed implementation of the switchable-limit interlayer upsampling outlined above. It should be noted, for example, as... Figure 4 In the case of images 12 and 15 being time-aligned paired images and videos, syntax element 616 can signal or toggle restrictions within each time frame. Furthermore, it should be noted again that the decoder according to the embodiments of this application can be compared with the above-referenced... Figure 4 as well as Figure 10 The provided description and functionality are consistent. Therefore, it should be noted that the above reference... Figure 4 The signaling provided regarding the short grammatical elements and the location and orientation of the spatial segments in Figures 12 and 15 should be considered equally applicable to the reference. Figures 10 to 13 The described implementation method. Finally, it should be noted that if Figure 10The decoder in the image decoder is an image decoder that decodes images in layer 0 and layer 1, so the second aspect is also advantageous. The temporal component is optional.
[0114] The following describes the implementation of restricted inter-layer upsampling in HEVC. For spatial scalability, the enhancement layer image is predicted using the upsampled base layer image. In this process, the predicted value for each pixel location in the enhancement layer is calculated using multiple pixel values (typically in the horizontal and vertical directions) corresponding to the base layer image region. If pixels from different base layer slices are used, the enhancement layer slice cannot be decoded solely from base layer slice information covering the same image region as the enhancement layer slice. Signaling restricted inter-layer upsampling as a bitstream constraint ensures that the decoder conforms to this constraint by making the spatial partitioning signals emitted from all parameter sets of the encoded video sequence comply with it, thereby simplifying the initialization and operation of the parallel inter-layer decoder.
[0115] Figure 10 The concept can be implemented as a mechanism to disallow the use of neighboring cell information about upsampling that is not included in the base layer corresponding to the enhancement layer. In the bitstream, this involves signaling whether the decoder is allowed to use cells outside the corresponding image region of the base layer located at all enhancement layer boundaries.
[0116] Alternatively, in the bitstream, the decision on whether the decoder is allowed to emit signals from pixels outside the corresponding image region of the base layer located at the enhancement layer boundary is made only for the enhancement layer boundary line corresponding to the base layer boundary line.
[0117] In a specific implementation, upsampling of the base layer at the edge of the image is performed, as if the base layer were located on the edge of the image. Here, no adjacent pixels can be obtained.
[0118] In specific implementation methods, such as Figure 14 As shown in the document, signaling can be achieved by using the markers in the image parameter set of the enhancement layer.
[0119] `independent_tile_upsampling_idc` corresponds to syntax element 612. A non-zero `independent_tile_upsampling_idc` restricts upsampling filters from crossing tile boundary lines. If `independent_tile_upsampling_idc` equals 2, any base layer samples located outside the image region corresponding to the enhancement layer tile cannot be used for upsampling. If `independent_tile_upsampling_idc` equals 1, this restriction applies only to enhancement layer tile boundary lines aligned with the base layer tile boundary lines. A `independent_tile_upsampling_idc` of 0 does not imply this restriction.
[0120] at last, Figure 15A Images 12 and 15, which are two slices overlapping in a spatial correspondence manner, are shown exemplarily to illustrate... Figure 14 An example of syntax element 612: An independent_tile_upsampling_idc equal to 2 restricts the upsampling filter from crossing any enhancement layer boundary line. See single-dotted line 400. If independent_tile_upsampling_idc equals 1, this restriction applies only to enhancement layer boundary lines aligned with the base layer boundary line. See double-dotted line 402. An independent_tile_upsampling_idc equal to 0 does not imply this restriction.
[0121] As an alternative to mode independent_tile_upsampling_idc=2, or as an additional mode such as independent_tile_upsampling_idc=3, the upsampling filter may be restricted from crossing any tile boundary lines; that is, it cannot cross these base layers or these enhancement layers. See Figure 15B Line 404 in the middle.
[0122] That is, as referenced above. Figure 9 As explained, according to this pattern, an upsampling filter is applied at boundaries 400, 402, or 404.
[0123] Beyond the next aspect of this application, it should be noted that, for example, in [the following text is missing from the original] Figure 2 The interpolation method 620 discussed above is performed in predictor 60 to obtain inter-layer prediction results. Similar to the encoder end (e.g., within predictor 18), since the encoder performs the same prediction at the encoding end, information 620 is executed according to the settings of syntax element 616. For example, at the encoding end, a decision on whether to set a syntax element can be made based on the application scenario. For example, if latency is a high priority, the syntax element can be set to restrict inter-layer upsampling, and in other application scenarios, better prediction and increased compression ratio may be more important, making it more advisable to set syntax element 612 to not restrict inter-layer upsampling.
[0124] The minimum coding delay or offset between the encoded spatial segments of consecutive layers, mentioned earlier, is the subject of the next section and can be referred to as a "layer decoding delay indicator." The decoder can determine the minimum decoding delay or offset between the encoded spatial segments of image 15 and those of image 12 based on short-term syntax elements. However, according to the next concept, the inter-layer delay or offset is signaled in advance within a predetermined time period using the previous syntax element structure. Again, this helps the decoder perform workload distribution within the parallel decoding of bitstream 40. As a measure for the "delay" or "offset," spatial segments can be used; that is, the offset can be expressed within the units of a spatial segment (slices, slices, or CTB rows of WPP).
[0125] To describe the following aspects in more detail, references are largely made to [the relevant sources]. Figure 4 Consistent Figure 16 .therefore, Figure 16 The same reference notation is used, and references are made to elements referred to as these common elements (if possible), as mentioned above. Figure 4 The proposed description also applies Figure 16 It should also be mentioned that, in addition to the functions set below, Figure 16 The video decoder 640 shown can be integrated Figure 4 The functionality described in long-term grammar element 606 is referenced in this application. Figure 16 This aspect also utilizes a long-term syntax element structure, namely 642, which is also inserted into the bit stream 40 to reference or relate to a predetermined time period 608. In other words, although the video decoder 640 can respond... Figure 4 The syntax element structures 606 and 642 are mentioned, however, only the latter's functionality, further summarized below, is relevant. Figure 16The decoder 640 is of particular importance, where the functionality related to the syntax element structure 606 and the existence of the syntax element structure 606 in the bitstream 40 are optional for the video decoder 640. However, the description above with reference to the video decoder 600 also applies to the video decoder 640. That is, the video decoder 640 is capable of decoding multi-layer video data stream 40, i.e., using inter-layer predictions from the first layer (layer 0) to the second layer (layer 1) to encode the scene in the layer hierarchy. The video decoder 40 supports parallel decoding of multi-layer video data streams into spatial segments, i.e., partitioning the images of a layer into spatial segments by sequentially traversing the spatial segments in a temporally overlapping manner, with an inter-layer offset between the spatial segments of the first layer image traversed and the spatial segments of the second layer image traversed. It should be noted that this means that a spatial segment can be a slice, substream, or piece; however, even a mixture of the segment units just mentioned is possible. In fact, the definition of "spatial segment" can differ when combining the concept of a slice with the concept of a piece and / or the concept of a substream.
[0126] In any case, regarding images 12 and 15 of the common time frame 604, Figure 16 The video decoder 640 is capable of decoding spatial segments of image 12 in parallel with spatial segments of image 15 on the other hand (i.e., by temporal overlap). Naturally, therefore, due to inter-layer prediction, the video decoder 640 needs to conform to a certain minimum decoding offset between the two layers, and the current decoding portion of enhancement layer 1 in image 15 must belong to the temporally aligned decoded portion of image 12 in layer 0.
[0127] in the case of Figure 16 The video decoder 640 uses the long-term syntax element structure 642 to predetermine the interlayer offset for a predetermined time period 608.
[0128] Combination Figure 16 In the implementation method described, the interlayer offset is related to one aspect. Figure 12 The "distance" of the first spatial segment is a scalar measurement of the time-aligned image 15 on the other hand. Preferably, the spatial "distance" is measured. Moreover, for this to be meaningful, the inter-layer offset determined based on the long-term syntax element structure 642 should be valid for the entire decoding process of the first spatial segment of image 12. That is, all necessary reference portions of image 12 used for inter-layer prediction can be used for decoding the entire first spatial segment of image 15, provided that the first "inter-layer offset" spatial segment of the base layer image 12 has been decoded previously.
[0129] As described above, the "current decoding portion" within image 15 traverses image 15 in a specific predetermined manner; that is, if slice-parallel processing is used, it follows the slice order described above, and in the case of using the WPP concept of bitstream, it is done in the form of a tilted wavefront. The same applies to the spatial segments of the base layer image 12. Before the first spatial segment of image 15 undergoes initial decoding, the interlayer offset determines the traversed portion of image 12 that has already been processed.
[0130] For a more detailed description, please refer to [reference needed]. Figure 17A and Figure 17B . Figure 17A The interlayer offsets related to the concept of slice, as determined from the long-term grammar element structure 642, are described in more detail. Figure 17B The inter-layer offsets related to WPP, determined based on the Long Term Syntax Element Structure 642, are described in more detail below. This will be discussed in conjunction with... Figure 17C The concept of inter-layer offset signaling using Long Syntax Element Structure 642 is shown to be unrestricted by the slice and / or WPP concepts used. More precisely, the interpretation of inter-layer offsets based on Long Syntax Element 642 is made feasible by defining that only the picture is subdivided into slices that are decodeable in a self-contained manner (i.e., entropy decoding and spatial intra-slice picture prediction are performed "in-slice" or entropy decoding and spatial intra-slice picture prediction are independent of adjacent slices).
[0131] Figure 17A Two time-aligned images, 12 and 15, are shown, both subdivided into slices. From the description of the slice concept presented above, it becomes clear that, typically, there is an arbitrary fixed order within the slice of image 12 or image 15 that is being decoded. More precisely, the slices can be decoded in any order. However, combining... Figure 16 In the implementation method described, at least relative to the sheet order of the base layer image 12, the sheet order 644 is defined as a row raster scan order guiding from the upper left sheet in a regular arrangement of sheets to the lower right sheet. According to... Figure 17A In the implementation described, the inter-layer offset signal emitted by the long-term syntax element structure 642 indicates the number of slices that have been decoded according to the slice order 644 of the base layer image 12, allowing the decoder 640 to begin decoding the first slice of the enhancement layer image 15. To determine the "first slice" among the slices of the enhancement layer image 15, the first slice of the enhancement layer image 15 can be fixedly defined as the uppermost slice of the enhancement layer image 15. From the first slice to the enhancement layer image 15, the video decoder 640 can employ a slice order for traversing the enhancement layer slices of image 15 according to the slice subdivision of image 12. Figure 17AIn some cases, such as when image 12 is subdivided into two rows and three columns, and image 15 is subdivided into four rows and two columns, it is advantageous for decoder 640 to select the order in which to traverse the enhancement layer slices, first traversing the left-hand slices of the first two rows, then traversing the right-hand slices of the first two rows, and then repeating the traversal for the lower rows of enhancement layer image 15, as indicated by arrow 646. However, according to alternative embodiments effective for all aspects described herein, the slice decoding order between enhancement layer slices of image 15 is fixed and independent of the subdivision of the base layer image. In summary, if only the signal layer coding offset is used as a trigger for starting / commencing decoding of the enhancement layer image, the recorder is unnecessary. Figure 17A The dashed lines indicate the local location in image 12 corresponding to the first piece of the enhancement layer image 15. Figure 17A It can become clear in the middle, in Figure 17A In the exemplary case, the interlayer offset determined by the long-term syntax element structure 642 can be "2" because the first two slices of image 12 must be decoded before the decoder 640 begins decoding the first slice of enhancement layer image 15. In this case, the co-location positions required for interlayer prediction can only be used in the base layer image 12.
[0132] That is, in Figure 17A In this case, the video decoder 640 will determine from the long syntax element structure 642 that the interlayer offset between the base layer of image 12 and the first layer of the enhancement layer of image 15 is two base layers: before the video decoder 640 starts decoding the first layer of the enhancement layer of image 15 in slice order 646, it has to wait for the first two base layers to be decoded in slice order 644.
[0133] Figure 17B An exemplary case involves subdividing two time-aligned images 12 and 15 into substreams, namely, in Figure 12 In the case of Figure 12, the subdivision is divided into two substreams, and in the case of Figure 15, it is divided into four substreams. For example, the substreams can be consistent with the rows and columns of the coded blocks that regularly subdivide Figures 12 and 15 in the manner described above, that is, in such a way that each substream corresponds to a row of such a coded block. In any case, as described above, due to WPP processing, a decoding order is defined between the substreams of Figure 12 and Figure 15, with decoding order 648 and decoding order 650 guided from top to bottom. Figure 17ASimilar to the case in [the previous example], decoder 640 is configured to determine from long syntax element structure 642 the number of pilot substreams that have been decoded before the first substream of image 15 begins decoding. In this case, long syntax element structure 642 will emit a signal with an interlayer offset of 1, because fully decoding the first substream of the base layer image 12 is sufficient to provide the necessary basis for any interlayer prediction of the first substream of the enhancement layer image 15.
[0134] Figure 17C Two time-aligned images, 12 and 15, both subdivided into slices, are shown. Again, the slice order or decoding order is defined between the slices of image 12 and between the slices of image 15, respectively, with order 652 and order 654 guiding from top to bottom. Figure 17C In the exemplary case, the boundaries between slices within image 12 on one side and slices within image 15 on the other side locally correspond to each other. Therefore, based on the "blurring" introduced by the inter-layer prediction from the base layer image 12 to the enhancement layer image 15, the Long Term Syntax Element structure 642 will signal that the inter-layer prediction offset is equal to 1 or equal to 2. Specifically, for example, due to the magnification of the corresponding co-located reference portion in image 12 for the inter-layer prediction portion of image 15, for example, as referenced above... Figure 9 As discussed, due to the disparity compensation vector or due to the upsampling interpolation filter kernel, the first two slices of image 12 may have already been decoded in slice order 652 before decoder 640 can begin decoding the first slice of enhancement layer image 15. However, for example, since the sampling resolution between images 12 and 15 is equal to each other and images 12 and 15 are related to the same field of view, such that no disparity compensation occurs, if the blurring option of interlayer prediction is turned off or not applied, the long-term syntax element structure can be set to 1 by the encoder so that decoder 640 can begin decoding the first slice of enhancement layer image 15 as soon as possible before the first slice of base layer image 12 is fully decoded.
[0135] therefore, Figures 16 to 17CThe description illustrates how the use of the long-term syntax element structure 642 helps the encoder support arbitrary parallel decoding of time-aligned images 12 and 15 scheduled by the decoder. Specifically, the long-term syntax element structure informs the decoder, based on the inter-layer offset, that the inter-layer offset is valid for the entire predetermined time period 608 and is related to the number of spatial segments of the base layer image 12, which has already been decoded beyond the first spatial segment of the time-aligned image 15. It should be noted that the video decoder 640 is capable of determining the inter-layer offset signaled by the long-term syntax element structure 642 (or even by itself) based on the detection / evaluation of the short-term syntax element 602 and further grammatical uniformity regarding potential options related to inter-layer prediction, which can turn on or off the just-summarized blur in inter-layer prediction from the base layer to the enhancement layer. However, the video decoder 640 must detect multiple syntax elements to derive the same information provided by the long-term syntax element structure 642, and the video decoder 640 can only derive the same information about the long predetermined time period 608 based on the short-term, rather than in advance.
[0136] Similar to aspects 1 and 2, possible ways to introduce the delay indication aspect into HEVC are described below.
[0137] First, refer to Figure 18 This describes how WPP is currently implemented in HEVC. That is, this description also forms the basis for alternative implementations of WPP processing in any of the above embodiments.
[0138] In the base layer, wavefront parallel processing allows for parallel processing of code tree block (CTB) rows. Predictions across CTB rows do not break interdependencies. Regarding entropy coding, from Figure 18 As can be seen, WPP alters the CABAC interdependencies of the top-left CTB in the corresponding upper-layer CTB row. Once the entropy decoding of the corresponding top-right CTB is completed, entropy encoding of the CTBs in subsequent rows can begin.
[0139] In the enhancement layer, once the CTB containing the corresponding image region is fully decoded and available, the decoding of the CTB can begin.
[0140] Figure 16 The decoding delay or offset in the context is merely a concept that can end in the signaling that prompts the decoder to initialize and manipulate the layered bitstream of slices, WPPs, or slices that utilize parallelism.
[0141] When spatial scalability is used, decoding of the enhancement layer CTB can only begin when the base layer CTB covering the corresponding image region is available. When spatial scalability is used to parallelize WPP with layered bitstreams, the layers can differ in terms of image size. For example, the published calls for scalability extensions for HEVC[1] and additional maximum CTB sizes specify image size scaling factors of 1.5 and 2 between layers. For example, the main HEVC profile supports image sampling of 16, 32, and 64. Regarding quality scalability, the image size scaling factor is usually constant, but the maximum CTB size between layers can still differ.
[0142] The ratio between the maximum CTB size of each layer and the image size scaling factor affects the layer decoding latency. That is, relative to the decoding of the base layer's CTB line, the CTB lines offset before the first CTB line of the enhancement layer can be decoded. Figure 19 The report presents exemplary parameter values for image size scaling factor and CTB size, as well as the ratio of the CTB to the corresponding image regions in the two layers, introduced by the layer decoding delay in the CTB row aspect.
[0143] Regarding quality scalability between layers, typically, the image size scaling factor between layers is equal to 1, while the maximum CTB size of the corresponding layers can still differ and affect the layer decoding latency.
[0144] Syntax element structure 642 provides decoder hints in the bitstream, i.e., signals about the layer decoding delay of independent spatial or quality enhancement layers when WPP processing across spatial enhancement layers is parallelized.
[0145] The implementation uses the image size scaling factor and the maximum CTB size scaling factor between the corresponding layers to determine the signal layer decoding delay.
[0146] Depending on the type of scalability between the independent base layer bitstream and the independent enhancement layer bitstream, the factors affecting layer decoding latency can be different.
[0147] In terms of viewpoint scalability, layers represent camera views and predictions are performed between camera views from various angles using an inter-layer prediction mechanism. This prediction utilizes a motion compensation mechanism to compensate for different camera positions within the camera setup. In this case, layer decoding latency is further limited by the maximum or actual motion vector in the vertical direction, compared to cases of spatial or mass scalability.
[0148] Syntax element structure 642 describes the decoder hints in the bitstream, namely, signaling the layer decoding delay of individual camera views when WPP processing across polyphonic camera views is made parallel.
[0149] The implementation uses an image size scaling factor, a maximum CTB size scaling factor, and the maximum motion vector length between corresponding layers in the vertical direction to determine the signal layer decoding delay.
[0150] In the VUI syntax or SPS associated with the enhancement layer, or in the compilation within the VPS extended syntax, when using WPP, the implementation signals the layer decoding latency regarding spatial scalability, quality scalability, or multi-view scalability in terms of spatial segments (i.e., lines of CTB).
[0151] slices and slices
[0152] Parallel processing using other partitioning techniques similar to slices or slicing can also benefit from hints within the bitstream indicating decoding latency based on dividing the image into spatial segments (i.e., slices or slices). Enhancement layer decoding processing requires information from the base layer (e.g., reconstructing image data).
[0153] Syntax element structure 642 describes the decoder hints within the bitstream, i.e., signals about the layer decoding delay of slices and / or segments.
[0154] Possible embodiments of the present invention use spatial segments as units to express the introduced layer processing delay, depending on the type of parallel technology used in the encoded video sequence.
[0155] Figure 20 The syntax in the document provides an exemplary implementation of the indication of min_spatial_segments_delay in the VUI parameters of the enhancement layer SPS for the parallel tool WPP, slices, and slices (an embodiment of syntax element structure 642).
[0156] min_spatial_segment_delay describes the decoding delay of the current layer relative to the corresponding base layer in terms of spatial segments, introduced by encoding interdependencies.
[0157] Based on the value of min_spatial_segment_delay, apply the following:
[0158] If min_spatial_segment_delay equals 0, no signal is emitted regarding the minimum delay limit between layer decoding.
[0159] Otherwise, (min_spatial_segment_delay is not equal to 0), bitstream consistency is specified to be true only if one of the following conditions is true:
[0160] • In the set of picture parameters activated within the encoded video sequence, tiles_enabled_flag is equal to 0 and entropy_coding_sync_enabled_flag is equal to 0 (i.e., neither slices nor WPP are used in the video sequence), and all base layer resources for decoding the first slice of the current layer in bit order are available when the first min_spatial_segment_delay slice of the base layer is fully decoded in bit stream order.
[0161] • In the set of individual picture parameters activated within the encoded video sequence, tiles_enabled_flag equals 1 and entropy_coding_sync_enabled_flag equals 0 (i.e., slices are used in the video sequence), and all base layer resources for decoding the first slice of the current layer in bitstream order are available when the first min_spatial_segment_delay slice covering the same image region is fully decoded.
[0162] • In each set of picture parameters activated within the encoded video sequence, tiles_enabled_flag is equal to 0 and entropy_coding_sync_enabled_flag is equal to 1 (i.e., WPP is used in the encoded video sequence), and all base layer resources for decoding processing of the first CTB line in the current layer are available when the first min_spatial_segment_delayCTB line of the base layer is completed.
[0163] Another exemplary implementation is as reported in [4] and as... Figure 21 The indication of min_spatial_segments_delay in the extended VPS extended syntax is shown.
[0164] min_spatial_segment_delay describes the decoding delay of layer [i] relative to the corresponding base layer in terms of spatial segments, introduced by encoding interdependencies.
[0165] Based on the value of min_spatial_segment_delay, the following applies: If min_spatial_segment_delay equals 0, a signal is emitted indicating the minimum delay limit between layer decoding.
[0166] Otherwise (min_spatial_segment_delay is not equal to 0), bitstream consistency is guaranteed to be true only if one of the following conditions is met:
[0167] • In the set of picture parameters activated within the encoded video sequence, tiles_enabled_flag is equal to 0 and entropy_coding_sync_enabled_flag is equal to 0 (i.e., neither slices nor WPP are used in the video sequence), and all base layer resources used for decoding the first slice of the current layer in bit order are available when the first min_spatial_segment_delay slice of the base layer is fully decoded in bit order.
[0168] • In the set of individual picture parameters activated within the encoded video sequence, tiles_enabled_flag is equal to 1 and entropy_coding_sync_enabled_flag is equal to 0 (i.e., slices are used in the video sequence), and all base layer resources for decoding the first slice of the current layer in bit order are available when the first min_spatial_segment_delay slice covering the same image region is fully decoded.
[0169] • In each set of picture parameters activated within the encoded video sequence, tiles_enabled_flag is equal to 0 and entropy_coding_sync_enabled_flag is equal to 1 (i.e., WPP is used in the encoded video sequence), and all base layer resources for decoding processing of the first CTB line in the current layer are available when the first min_spatial_segment_delayCTB line of the base layer is completed.
[0170] The various prediction modes supported by the encoder and decoder, the constraints imposed on these modes, and the contextual derivation with entropy encoding / decoding have been described above, enabling parallel processing concepts such as slices and / or WPP concepts. It has also been mentioned above that the encoder and decoder can operate on a block-based basis. For example, the prediction modes described above can be selected based on blocks, i.e., at a finer granularity than the image itself. Before proceeding to describe another aspect of this application, the relationship between slices, flakes, WPP bitstreams, and the blocks just mentioned will be explained.
[0171] Figure 32An image is shown, which could be a layer 0 image such as layer 12 or a layer 1 image (such as image 15). This image is regularly subdivided into an array of blocks 90. These blocks 90 are sometimes referred to as maximum coding blocks (LCBs), maximum coding units (LCUs), coding tree blocks (CTBs), etc. Subdividing the image into blocks 90 can form a base or coarse-grained level by which the aforementioned prediction and residual coding are performed, and this coarse-grained level (i.e., the size of block 90) can be set by signaling the encoder for layers 0 and 1 respectively. For example, multiple trees such as quadtree subdivision can be used, and signals can be emitted within the data stream to subdivide each block 90 into prediction blocks, residual blocks, and / or coding blocks respectively. Specifically, the coding block can be a leaf block of a recursive multi-tree subdivision of block 90 and can emit some prediction-related decisions (such as prediction modes) at the granularity of the coding block, and the prediction block and the residual block can be separate leaf blocks of the recursive multi-tree subdivision in the coding block. For example, prediction parameters such as motion vectors (if it is inter-temporal prediction) and disparity vectors (if it is inter-layer prediction) are encoded at the granularity of the prediction block, and the prediction residuals are encoded at the granularity of the residual blocks.
[0172] The raster scan encoding / decoding order 92 can be defined within block 90. The encoding / decoding order 92 restricts the availability of adjacent portions for spatial prediction purposes: only portions of the image according to the encoding / decoding order 92 advance to a current position such as block 90 or a smaller block thereof. The current position, related to the currently predicted syntax element, is available for spatial prediction within the current image. The encoding / decoding order 92 traverses all blocks 90 of the image within each layer, and then continues traversing the blocks of the next image in the corresponding layer according to the image encoding / decoding order. The image encoding / decoding order does not necessarily have to follow the temporal reproduction order of the images. Within each block 90, the encoding / decoding order 92 is refined into scans between smaller blocks, such as encoded blocks.
[0173] Regarding the blocks 90 and smaller blocks just summarized, each image is further subdivided into one or more slices according to the encoding / decoding order 92 mentioned earlier. Figure 32 Slices 94a and 94b, as exemplarily shown, seamlessly cover their respective images. The boundary or contact surface 96 between consecutive slices 94a and 94b in an image may or may not be aligned with the boundary of adjacent block 90. More precisely, Figure 32 The right-hand side of the image shows consecutive slices 94a and 94b that may border each other at the boundary of smaller blocks such as coding blocks, i.e., subdivided leaf blocks of one of blocks 90.
[0174] Because slices 94a and 94b of the image can form a minimum unit, within which this portion of the data stream encoded from the image can be packaged into a packet, i.e., a NAL unit. For example, another possible property of slices described above is the restriction on slices determined by the prediction and entropy context across slice boundaries. Slices with this restriction can be called “normal” slices. In addition to normal slices, “dependent slices” can also exist, as summarized in more detail below.
[0175] If the concept of segmentation is applied to an image, the encoding / decoding order 92 defined between the array 90 of blocks can be changed. Figure 33 This situation is illustrated in the example, where an image is exemplarily divided into four slices 82a to 82d. (See also...) Figure 33 As shown, the slice itself is defined as regularly subdividing the image into blocks 90. That is, each of slices 82a to 82d consists of an n×m array of blocks 90, with n set individually for each row of slices and m set individually for each column of slices. Following the encoding / decoding order 92, before proceeding to the next slice 82b, etc., the blocks 90 in the first slice are first scanned according to the raster scan order, wherein slices 82a to 82d themselves are scanned according to the raster scan order.
[0176] Based on the WPP stream partitioning concept, the image in one or more rows of blocks 90 is subdivided into WPP substreams 98a to 98d according to the encoding / decoding order 92. For example... Figure 34 As shown, for example, each WPP substream can cover a complete line of block 90.
[0177] However, the slice concept can also be combined with the WPP substream concept. In this case, for example, each WPP substream can cover one row of blocks 90 within each slice.
[0178] It is even possible to use image slice partitions and / or WPP substream partitions together. Regarding slices, each of the one or more slices into which the image is subdivided, according to the encoding / decoding order 92, can consist of only one complete slice, or more than one complete slice, or only a sub-part of a slice. Therefore, the slice constituting the smallest unit for parallelism can include normal slices on one hand and dependent slices on the other: normal slices impose the aforementioned constraints and entropy context on the prediction, while dependent slices do not. Dependent slices starting at image boundaries use the entropy context generated from the entropy decoding block 90 in the immediately preceding block 90, with the encoding / decoding order 92 unfolding sequentially in rows, and dependent slices starting elsewhere can use the entropy encoding context generated from entropy encoding / decoding the immediately preceding slice to its end. In this way, each of the WPP substreams 98a to 98b can consist of one or more dependent slices.
[0179] That is, the encoding / decoding order 92 defined between blocks 90 linearly guides from the first side (here, exemplarily, the left side) of the corresponding picture to the opposite side (exemplarily, the right side), and then steps down / to the next line of block 90. Therefore, the encoded / decoded portions of the current picture, mainly arranged to the left and top of the current encoding / decoding portions such as the current block 90, are available. Due to prediction interruptions and entropy context derivation across boundary lines, slices of a picture can be processed in parallel. Even, the encoding / decoding of slices of a picture can begin simultaneously. The constraint originates from the aforementioned in-loop filtering, in which case the constraint allows for crossing boundary lines. Furthermore, the encoding / decoding of the starting WPP substream is performed in an interleaved manner from top to bottom. Measuring the intra-frame picture delay between consecutive WPP substreams in block 90 is two blocks 90.
[0180] However, it is even advantageous to parallelize the encoding / decoding of images 12 and 15, i.e., time slots between different layers. Clearly, the encoding / decoding of image 15, which depends on the base layer, must be delayed relative to the encoding / decoding of the base layer to ensure that a usable "spatial correspondence" portion of the base layer already exists. These ideas are valid even if no parallelism in the encoding / decoding within either image 12 or image 15 is used independently. Even if a slice is used to cover both images 12 and 15 separately, and no slices or WPP substream processing is used, the encoding / decoding of images 12 and 15 can be parallelized. The signaling described below (i.e., the sixth aspect) expresses the possibility of delaying the decoding / encoding of layers, even in this case, regardless of whether slices or WPP processing are used for any image of the layer.
[0181] From the above information regarding the minimum coding delay between consecutive layers, it becomes clear that the decoder can determine the minimum decoding delay based on short-term syntax elements. However, if the aforementioned long-term syntax elements are used to signal the inter-layer time delay in advance within a predetermined time period, the decoder can plan to use the provided guarantees in the future and can more easily perform workload distribution within the parallel decoding of bitstream 40.
[0182] The aspect of this application described below (i.e., the sixth aspect of this application) relates in part to aspect 3, in that it involves explicit signaling of any inter-layer offset. However, in contrast to the sixth aspect of this application, the syntactic element structure for explicitly issuing signaling of inter-layer offsets does not need to be based on long-term signaling related to short-term syntactic elements (from which inter-layer offsets can be further derived). More precisely, the sixth aspect of this application adopts another finding: when describing Figures 17A to 17CIt becomes clear that if the base layer image and enhancement layer image are subdivided into blocks according to the time-defined raster scan decoding order, then the interlayer offset signal between the base layer and enhancement layer can be effectively and clearly emitted by measuring the interlayer offset within the cell of the base layer block. In conjunction with aspects further described below, the base layer block within the cell from which the interlayer offset signal is emitted is not limited to a spatial segment. More specifically, given this, other coded blocks can be used. Therefore, when referring to... Figure 34 In describing embodiments of the sixth aspect of this application, references are used extensively. Figures 16 to 17C The reference numerals used, and the descriptions of the features mentioned above with reference to the latter, should also apply to the embodiments further described below, in order to avoid unnecessary repetition. In addition, references... Figure 32 and Figure 33 The description is based on the possibility of coexistence between coded blocks on one side and spatial segments on the other side.
[0183] therefore, Figure 35 A video decoder 720 configured to receive a multi-layer video data stream 40 is shown, which encodes the scene in the layer hierarchy using inter-layer prediction from a portion of the first layer to a co-located portion of the second layer. Similar to the above figure, layers 0 and 1 juxtaposed as a representative embodiment are shown exemplarily. Figure 35 The image exemplarily illustrates two time-aligned images 12 and 15 of two layers. Image 12 of the base layer 0 is subdivided into a series of first blocks 722, and the image of the enhancement layer 1 is subdivided into a series of second blocks 724. A raster scan decoding order 726 is defined between blocks 722, and similarly, a raster scan decoding order 728 is defined between blocks 724.
[0184] The video data stream 40 includes a syntax element structure indicating the interlayer offsets for sequentially traversing the first block 722 and the second block 724 in a temporally overlapping manner to decode images 12 and 15 in parallel, as well as the interlayer offsets between the traversed first block 722 and second block 724 measured in the cells of the base block 722. The video decoder 720 is configured to respond to the syntax element structure 730. Specifically, the video decoder determines the interlayer offsets from the latter.
[0185] and Figure 16Consistent with the implementation described above, the syntax element structure 730 can indicate inter-layer offsets as a guarantee for successful parallel decoding of time-aligned images 12 and 15 over a predetermined time period longer than a short time interval, during which optional syntax element signal images 12 and 15 are subdivided into blocks 722 and 724, respectively. However, this is not mandatory. More precisely, the signaling of inter-layer offsets via the syntax element structure 730 can be implemented within different ranges of data stream 40, such as each pair of time-aligned base layer images 12 and enhancement layer images 15, for example, within the same interval as the signaling action regarding size and subdivision into blocks 722 and 724.
[0186] In further accordance with the above implementation, the decoder 720 can use the layer offset explicitly emitted by the syntax element structure 730 as a measurement of the offset relative to the traversal of the first block 722 when the second block 724 begins to traverse, respectively, during the parallel decoding of images 12 and 15. In other words, the video decoder 720 can be configured to derive a count value from the syntax element structure 730, and is configured to calculate the decoded block 722 of the base layer image 12 according to the raster scan decoding order 726, only after reaching the count of the decoded block 722 that explicitly emits the minimum count signal through the syntax element structure 730, while allowing the decoding of the sequence of blocks 724 of the enhancement layer image 15 to begin according to the decoding order 728. Therefore, the video decoder 720 does not need to detect arbitrarily high-complexity and distributed portions of the video data stream 40, which would otherwise allow the video decoder 720 to calculate the actual minimum inter-layer offset between the starting decoding block 722 on one side and the block 724 on the other side in other ways.
[0187] However, interestingly, according to Figure 35 In the implementation described above, blocks 722 and 724 do not necessarily represent spatial segments that have undergone any parallel processing. More precisely, blocks 722 and 724 can be common coded blocks, each representing the content of images 12 and 15 encoded into video data stream 40. For example, blocks 722 and 724 can be root blocks, where images 12 and 15 are regularly subdivided (i.e., by rows and columns), and then the root blocks are further subdivided individually using a recursive multi-tree approach as described above with reference to 32. For example, the generating leaf blocks representing the root blocks of images 12 and 15 are subdivided into coded blocks, emitting signals within video data stream 40 for the prediction mode selected by reference to FIG. 15 among spatial prediction, temporal prediction, and inter-layer prediction.
[0188] To be more detailed Figure 35 The implementation method is described below, with reference to Figure 36 .like Figure 36As shown, the video decoder 720 can use a counter 732 to count the number of decoded blocks 722 of the base layer image 12, starting from the first block 722 of image 12 according to the raster scan decoding order 726. The comparator 734 of the decoder 720 compares the steadily increasing count output by the counter 732 with the explicit signal value of the syntax element structure 730 obtained from the video data stream 40. If the count satisfies a predetermined relationship with the value indicated by the syntax element structure 730, such as once the count of the counter 732 reaches or equals the value indicated by the syntax element structure 730, the comparator is activated or causes the enhancement layer image 15 to begin decoding, i.e., causes the first block 724 of the enhancement layer image 15 to begin decoding according to the raster scan decoding order 728.
[0189] The syntax of grammar element structure 730, as described below in more detail, will be further explained below. Figure 35 The syntax element structure and Figure 16 The unification of syntactic element structure and its objectives is feasible. That is, as follows: Syntactic element structure 730 can have a set of possible values, i.e., a set of possible values. (See reference...) Figure 16 As exemplarily mentioned, values outside the non-explicit set of possible inter-layer offsets may cause the video decoder 720 to discard the values of the syntax element structure 730 and not specify parallel decoding of images 12 and 15, or to determine arbitrary inter-layer offsets based on the short-term syntax element 602. If the syntax element structure 730 assumes values outside the second set of possible values, this will cause the video decoder 720 to perform relative to... Figure 36 The actions already summarized, for example, according to this action, the value of the syntax element structure 730 will explicitly signal the inter-layer offset within the cell of the base layer block 722. However, when assuming through the syntax element structure 730, there may be another subset of possible values for the syntax element structure 730, causing the video decoder 720 to perform as described above. Figure 16 The described action is to determine the interlayer offset between the base layer image 12 and the enhancement layer image 15 by interpreting the latter as the interlayer offset of the cells of the spatial segment consisting of blocks 722 and 724, which can but do not necessarily have to be integer multiples of each other.
[0190] refer to Figure 37 It shows that Figure 35 The implementation methods and Figure 16 The possibilities for combining the implementation methods just mentioned. For example... Figure 37As shown, the video decoder can inspect the syntax element structure 730 to determine whether the syntax element structure 730 has a value in a first subset 736, a second subset 738, or a third subset 740 outside the set of possible values 742. Based on the results of this investigation or inspection, the decoder 720 cannot derive any guarantees outside the syntax element structure 730 and cannot derive any explicit signaling regarding slice offsets from the syntax element structure, or perform derivation of inter-slice offsets from the syntax element structure 730, i.e., perform the derivation within a cell of a spatial segment or a cell of a block. If it is the second subset 738, no derivation / guarantee occurs; if it is subset 736, the derivation of inter-slice offsets occurs within a cell of a spatial segment; and if the syntax element 730 assumes a value outside the third subset 740, the inter-slice offsets are derived within a cell of a block. In a specific grammatical embodiment further summarized below, the grammatical element structure includes two flags, namely, ctb_delay_enabled_flag and min_spatial_segment_delay, wherein ctb_delay_enabled_flag = 0 and min_spatial_segment_delay ≠ 0 corresponds to the case of subset 736, min_spatial_segment_delay = 0 corresponds to the second subset 738, and ctb_delay_enabled_flag = 1 and min_spatial_segment_delay ≠ 0 corresponds to the third subset 740.
[0191] Finally, refer to Figure 38 This illustrates that decoder 720 can be configured to decode any interlayer offset of the signal emitted through syntax element structure 730. This interlayer offset is not merely an interlayer offset relative to the first block or spatial segment of the starting enhancement layer picture 15, but a continuous interlayer offset. When conforming to continuous interlayer offsets, collision-free parallel decoding of pictures 12 and 15 is produced respectively. However, as... Figure 38As shown, counter 732 still counts the number of decoded blocks 722 of the base layer image 12, while additional counter 744 similarly counts the decoded blocks 724 of the enhancement layer image 15 according to the decoding order 728. Subtractor 746 forms the difference between the two counts, i.e., s and t-1, i.e., calculates s–t+1. Comparator 734 compares this difference with the inter-layer offset value derived from the syntax element structure 730, and once the two values (i.e., the interpolation between the derived inter-layer offset and the count) have a predetermined relationship, such as the difference being equal to or greater than the derived inter-layer offset, block t is decoded according to the decoding order 728 between the enhancement layer blocks 724. In this way, a continuity check is established between the decoded blocks 722 of the base layer image 12 on one hand and the blocks 724 of the enhancement layer image 15 on the other.
[0192] Obviously, according to Figure 38 The continuous testing can also be applied to spatial segments. More generally, it has also been applied to... Figure 38 and Figure 36 The description is transferred to the spatial segment and this statement also applies. Figure 16 In the implementation method described above, the syntax element structure 642 can be used as... Figure 36 and Figure 38 The relevant syntax element structure is shown at 730. In other words, at least when slices are used as spatial segments, there is also a defined raster scan decoding order between them, so that the discussion relative to the coding block can be... Figure 36 and Figure 38 Changes in the process can be easily transferred to the traversal and decoding of the slice.
[0193] In brief, summarizing the implementation of the sixth aspect and related descriptions, the syntax element structure 730 can be inserted into the bitstream by the video encoder to provide the decoder with explicit indications of how to control the parallel decoding of the base layer images and enhancement layer images relative to each other. The interlayer offset explicitly signaled by the syntax element structure can be either activated or deactivated. If activated, this indication can be within a cell such as a block (CTB), or within one of the block cells and spatial segment cells signaled via more precise signaling. Due to the raster scan order between the base layer blocks of one side and the enhancement layer blocks of the other, for example, both guided line by line from the top left corner to the bottom right corner of each image 12 / 15 from top to bottom, the explicitly signaled interlayer offset can be interpreted simply as a "trigger" for starting / commencing decoding of the first block of enhancement layer image 15, or as a continuous "safe distance" between the current decoding block of the base layer image 12 of one side and the current decoding block of the enhancement layer image 15 of the other side, i.e., a trigger for determining the decoding of each block of enhancement layer image 15. The description proposed in the sixth aspect can be transferred to the description and implementation of the third aspect, in that, at least whenever the description involves a slice as a spatial segment, it involves the interpretation and checking of the compliance of the signal layer offset, which can be done using... Figure 36 and Figure 38 The implementation method in the text, that is, can be achieved through communication with... Figure 36 and Figure 38 The description in the text corresponds to the method of controlling the traversal of the decoded chips in the base layer image and enhancement layer image according to the decoding order of the raster scan chips.
[0194] Therefore, a "delay" spatial segment can be used as a measure, that is, the delay can be expressed within the cells (slices, slices, or CTB rows) of the spatial segment, or the delay / offset can be measured within the cells of block 90.
[0195] The High Efficiency Video Coding (HEVC) standard can be extended to meet the sixth aspect below. Specifically, if reference data is available, parallel decoding of individual layers (or views) is permitted. The minimum delay (specifically, layer decoding delay) between the decoding of the base layer coding tree block (CTB) and the corresponding independent enhancement layer CTB is determined at the granularity of parallel tools such as slices. Wavefronts, slices, or motion compensation vectors are applicable (e.g., in stereoscopic or multi-view video coding).
[0196] Figure 20 This demonstrates a layer decoding delay indicator implemented by enhancing the semantics of the SequenceParameterSetSyntax and the syntax element min_spatial_segment_delay.
[0197] min_spatial_segment_delay describes the decoding delay of the current layer relative to the corresponding base layer in terms of spatial segments, introduced by encoding interdependencies.
[0198] The following mechanism can be implemented in HEVC high-order syntax to allow optional expression of layer decoding delays between independent base layers and independent enhancement layers based on multiple vertical and horizontal CTBs, independent of potential parallel techniques.
[0199] Signals can be emitted using exponential flags (e.g., ctb_delay_enabled_flag): layer decoding delays (such as signals emitted using the second syntax element) are expressed as specific CTB addresses in the encoded image.
[0200] from Figure 39 As can be seen and based on the following conditions, the CTB address in the raster scan sequence clearly defines the horizontal and vertical positions within the image used to express the delay.
[0201] CTB coordinates = (CTB address % PicWidthInCTBs, CTB address / PicWidthInCTBs)
[0202] PicWidthInCTBs describes the width of the image within a CTB cell.
[0203] Figure 39 The following is an example. The CTB address within the image (e.g., 7) defines the horizontal CTB column and the vertical CTB row, for example, the tuple (2,1).
[0204] If the flag is enabled, when decoding the CTB in the current independent layer, the value of another syntax element (cp.min_spatial_segment_delay) is interpreted as an offset of the CTB address relative to the co-located CTB in the base layer image.
[0205] like Figure 40 The details shown and described below demonstrate that the co-located CTB can be calculated based on the size of the CTB in the two corresponding layers and the width of the images in the two corresponding layers.
[0206] Figure 40 This includes three embodiments, from left to right, illustrating various settings for the CTB size and image size of the corresponding base and enhancement layers, regardless of image aspect ratio. Bold outlined boxes in the base layer image mark the image regions within the enhancement layer's CTB size and their co-located image regions within the corresponding base layer's CTB layout.
[0207] In the enhancement layer Sequence Parameter Set Syntax and by Figure 41The implementation of this optional CTB base layer decoding delay indication is given in the semantics of the syntax element min_spatial_segment_delay.
[0208] A ctb_based_delay_enabled_flag value of 1 indicates the delay for signaling via min_spatial_segment_delay in the CTB unit. A ctb_based_delay_enabled_flag value indicates the delay for signaling via min_spatial_segment_delay in the CTB unit.
[0209] min_spatial_segment_delay describes the decoding delay of the current layer relative to the corresponding base layer in terms of spatial segments, introduced by encoding interdependencies.
[0210] Based on the value of min_spatial_segment_delay, the following applies: if min_spatial_segment_delay equals 0, no signal is emitted regarding the minimum delay between layers during decoding.
[0211] Otherwise (min_spatial_segment_delay is not equal to 0), and if ctb_based_delay_enabled_flag is equal to 1, then the following conditions are specified for bitstream consistency to be true:
[0212] ·and CtbSizeY A PicWidthInCtbsY A and ctbAddrRs A It refers to CtbSizeY and PicWidthInCtbsY of base layer A, as well as CtbAddress and CtbSizeY of Ctb in base layer A according to the raster scan order. B PicWidthInCtbsY B and ctbAddrRs B It is the CtbSizeY and PicWidthInCtbsY of the independent layer / view B, and the CtbAddress of the independent layer B according to the raster scan order, and the CtbScalingFactor BA 、CtbRow BA (ctbAddrRs) and CtbCol BA (ctbAddrRs) is determined as follows:
[0213] CtbScalingFactorBA =(PicWidthInCtbsY A / PicWidthInCtbsY B )
[0214] CtbRow BA (ctbAddrRs)=Ceil((Floor(ctbAddrRs / PicWidthInCtbsY B )+1)*CtbScalingFactor BA )-1
[0215] CtbCol BA (ctbAddrRs)=Ceil(((ctbAddrRs%PicWidthInCtbsY B )+1)*CtbScalingFactor BA )–1
[0216] When using ctbAddrRs of the current enhancement layer / view B B When decoding CTB, if it has an equal value of PicWidthInCtbsY A *CtbRow BA (ctbAddrRs B )+CtbCol BA (ctbAddrRs B )+ctbAddrRs of min_spatial_segment_delay A When the base layer CTB is fully decoded, all necessary base layer resources are available.
[0217] Otherwise (min_spatial_segment_delay is not equal to 0 and ctb_based_delay_enabled is equal to 0), bitstream consistency is specified to be true only if one of the following conditions is true:
[0218] • In the set of picture parameters activated within the encoded video sequence, tiles_enabled_flag is equal to 0 and entropy_coding_sync_enabled_flag is equal to 0 (i.e., neither slices nor WPP are used in the video sequence), and all base layer resources for decoding the first slice of the current layer in bitstream order are available when the first min_spatial_segment_delay slice of the base layer is fully decoded in bitstream order.
[0219] • In each set of picture parameters activated within the encoded video sequence, tiles_enabled_flag equals 1 and entropy_coding_sync_enabled_flag equals 0 (i.e., slices are used in the video sequence), and all base layer resources for decoding the first slice of the current layer in bitstream order are available when the first min_spatial_segment_delay slice covering the same image region is fully decoded.
[0220] • In each set of picture parameters activated within the encoded video sequence, tiles_enabled_flag is equal to 0 and entropy_coding_sync_enabled_flag is equal to 1 (i.e., WPP is used in the encoded video sequence), and all base layer resources for decoding processing of the first CTB line in the current layer are available when the first min_spatial_segment_delayCTB line of the base layer is completed.
[0221] Alternatively, as in the previous implementation, a signal can be issued to use the inter-layer offset as a worst-case delay of the ctb_based_delay_enabled_flag, rather than the start delay of the first slice / piece / CTB row. The worst-case delay provides a guarantee that all necessary underlying layer resources are available when the co-located spatial segment with the signal offset is fully decoded during the decoding of spatial segments of individual images.
[0222] exist Figure 42 The implementation method for the syntax is shown in the figure.
[0223] min_spatial_segment_delay describes the decoding delay of the current layer relative to the corresponding base layer in terms of spatial segments, introduced by encoding interdependencies.
[0224] The following applies based on the value of min_spatial_segment_delay: If min_spatial_segment_delay equals 0, a signal is issued regarding any limit on the minimum delay between layers during decoding.
[0225] Otherwise (min_spatial_segment_delay is not equal to 0), bitstream consistency is specified to be true only if one of the following conditions is true:
[0226] • In the set of picture parameters activated within the encoded video sequence, tiles_enabled_flag is equal to 0 and entropy_coding_sync_enabled_flag is equal to 0 (i.e., neither slices nor WPP are used in the video sequence), and when the first slice segment C (which is immediately following the (min_spatial_segment_delay-1) slice in the base layer in bitstream order) is fully decoded, all base layer resources for decoding processing of any slice segment A in the current layer are available in bitstream order. This slice segment C is located in bitstream order after the last slice segment B, which contains at least a portion of the same image region relative to slice A in the current layer.
[0227] • In the set of individual picture parameters activated within the encoded video sequence, tiles_enabled_flag equals 1 and entropy_coding_sync_enabled_flag equals 0 (i.e., slices are used in the video sequence), and when the first slice C (which follows the (min_spatial_segment_delay-1) slice in bitstream order immediately after the last slice B containing at least a portion of the same image region relative to slice A) is fully decoded, all base layer resources for decoding processing of any slice A in the current layer are available in bitstream order.
[0228] • In the set of individual picture parameters activated within the encoded video sequence, tiles_enabled_flag is equal to 0 and entropy_coding_sync_enabled_flag is equal to 1 (i.e., WPP is used in the encoded video sequence), and all base layer resources are available for decoding processing of any CTB line A in the current layer when the first CTB line C (immediately following (min_spatial_segment_delay-1) CTB line) of the base layer, which is located in bitstream order and covers at least a portion of the same picture region relative to the CTB line A of the enhancement layer, is fully decoded.
[0229] Signaling based on sub-coded video sequence, such as images or min_spatial_segment_delay, is also feasible. For example... Figure 20 As given, with respect to the associated NAL unit, the range of the SEI message is smaller than the coded video sequence in the time domain, and the range of the SEI message is limited by its position in the bitstream or by an exponent. Figure 43 One implementation method is given in Layer_decoding_delay_SEI.
[0230] The syntax can be changed relative to the implementation described above to reflect the scope of the SEI message and its syntax elements.
[0231] The above-described implementation can be slightly modified. In the above embodiment, the syntax element structure includes `min_spatial_segment_delay` and `ctb_based_delay_enabled_flag`, and `min_spatial_segment_delay` measures the inter-layer coding offset within a unit of a spatial segment or CTB according to the spatial segment / CTB decoding order in a one-dimensional or scalar manner based on `ctb_based_delay_enabled_flag`. However, since the number of CTBs of the base layer image is usually equal to the number of spatial segments such as slices or bitstreams of the base layer image, in a slightly different implementation, if `ctb_based_delay_enabled_flag` indicates the inter-layer offset based on CTB, the inter-layer offset cannot be determined solely based on `min_spatial_segment_delay`. Instead, more precisely, in this case, the subsequent syntax element is interpreted as indicating the location of the CTB of the base layer image in the water dimension, and the complete decoding of the decoder can be used as a trigger to start the decoding of the enhancement layer image. Naturally, alternatively, `min_spatial_segment_delay` can be interpreted as indicating the CTB of the base layer image along the vertical dimension. Depending on the `ctb_based_delay_enabled_flag`, i.e., if its indication is based on the CTB, further syntax elements are sent in the data stream for positioning the base layer image's CTB in other dimensions, acting on the trigger just mentioned.
[0232] That is, the following syntax fragments used for signaling can be used, i.e., can be used as syntax element structures:
[0233] The indices i and j can indicate the layer IDs of the base layer and the enhancement layer, respectively.
[0234] min_spatial_segment_offset_plus1[i][j] ue(v) if(min_spatial_segment_offset_plus1[i][j]>0){ ctu_based_offset_enabled_flag[i][j] u(1) if(ctu_based_offset_enabled_flag[i][j]) min_horizontal_ctu_offset_plus1[i][j] ue(v)
[0235] The semantics of the above syntactic elements can be described as follows:
[0236] The following rules apply: in each image of the j-th direct reference layer of layer i (i.e., inter-layer predictions not used for decoding any image of layer i), min_spatial_segment_offset_plus1[i][j] indicates the spatial region either on its own or together with min_horizontal_ctu_offset_plus1[i][j]. The value of min_spatial_segment_offset_plus1[i][j] should be in the range of 0 to refPicWidthInCtbsY[i][j]*refPicHeightInCtbsY[i][j] (inclusive). If it does not exist, it is presumed that the value of min_spatial_segment_offset_plus1[i][j] is equal to 0.
[0237] Within a CTU cell, in each image of the j-th direct reference layer at layer j (i.e., inter-layer prediction not used for decoding any image in layer i), a spatial region is defined by ctu_based_offset_enabled_flag[i][j] equal to 1, jointly indicated by min_spatial_segment_offset_plus1[i][j] and min_horizontal_ctu_offset_plus1[i][j]. Within a slice segment cell, or within a CTU row, in each image of the j-th direct reference layer at layer i (i.e., inter-layer prediction not used for decoding any image in layer i), a spatial region is defined solely by ctu_based_offset_enabled_flag[i][j] equal to 0, indicated by min_spatial_segment_offset_plus1[i]. When it does not exist, the value of ctu_based_offset_enabled_flag[i] is presumably equal to 0.
[0238] The following stipulation applies: In each image of the j-th direct reference layer of layer i (i.e., inter-layer prediction not used for decoding any image in layer i), when ctu_based_offset_enabled_flag[i][j] equals 1, min_horizontal_ctu_offset_plus1[i][j] and min_spatial_segment_offset_plus1[i][j] jointly indicate the spatial region. The value of min_horizontal_ctu_offset_plus1[i][j] should be in the range of 0 to refPicWidthInCtbsY[i][j] (inclusive).
[0239] When ctu_based_offset_enabled_flag[i][j] equals 1, the derivation of the variable minHorizontalCtbOffset[i][j] is as follows:
[0240] minHorizontalCtbOffset[i][j]=(min_horizontal_ctu_offset_plus1[i][j]>0)? (min_horizontal_ctu_offset_plus1[i][j]–1):(refPicWidthInCtbsY[i][j]-1)
[0241] Set the variable curPicWidthInSamples L [i], curPicHeightInSamples L [i], curCtbLog2SizeY[i], curPicWidthInCtbsY[i], and curPicHeightInCtbsY[i] are set to be equal to PicWidthInSamples of the i-th layer, respectively. L PicHeightInSamples L ,CtbLog2SizeY,PicWidthInCtbsY andPicHeightInCtbs.
[0242] The variable refPicWidthInSamples L [i][j], refPicHeightInSamples L [i][j], refCtbLog2SizeY[i][j], refPicWidthInCtbsY[i][j], and refPicHeightInCtbsY[i][j] are set to be equal to PicWidthInSamples of the j-th direct reference layer in the i-th layer, respectively. L PicHeightInSamplesL, CtbLog2SizeY, PicWidthInCtbsY, and PicHeightInCtbsY.
[0243] Set the variables curScaledRefLayerLeftOffset[i][j], curScaledRefLayerTopOffset[i][j], curScaledRefLayerRightOffset[i][j], and curScaledRefLayerBottomOffset[i][j] to be equal to the scaled_ref_layer_left_offset[j]<<1, scaled_ref_layer_top_offset[j]<<1, scaled_ref_layer_right_offset[j]<<1, and scaled_ref_layer_bottom_offset[j]<<1, respectively, for the j-th direct reference layer of the i-th layer.
[0244] colCtbAddr[i][j] represents the raster scan address of the co-located CTU, and the raster scan address is equal to ctbAddr in the j-th layer image. In the j-th direct reference layer image of the i-th layer, the derivation of variable colCtbAddr[i][j] is as follows:
[0245] The variables (xP, yP) specify the location of the top-left luminance sample of the CTU, and the raster scan address is equal to ctbAddr, which is related to the top-left luminance sample in the image of the i-th layer.
[0246] xP=(ctbAddr%curPicWidthInCtbsY[i])< <curCtbLog2SizeY
[0247] yP=(ctbAddr / curPicWidthInCtbsY[i])< <curCtbLog2SizeY
[0248] The derivation of variables scaleFactorX[i][j] and scaleFactorY[i][j] is as follows:
[0249] curScaledRefLayerPicWidthInSamplesL[i][j]=curPicWidthInSamplesL[i]–curScaledRefLayerLeftOffset[i][j]–curScaledRefLayerRightOffset[i][j]
[0250] curScaledRefLayerPicHeightInSamplesL[i][j]=curPicHeightInSamplesL[i]–curScaledRefLayerTopOffset[i][j]–curScaledRefLayerBottomOffset[i][j]
[0251] scaleFactorX[i][j]=((refPicWidthInSamplesL[i][j]<<16)+(curScaledRefLayerPicWidthInSamplesL[i][j]>>1)) / curScaledRefLayerPicWidthInSamplesL[i][j]
[0252] scaleFactorY[i][j]=((refPicHeightInSamplesL[i][j]<<16)+(curScaledRefLayerPicHeightInSamplesL>>1)) / curScaledRefLayerPicHeightInSamplesL[i][j]
[0253] The variables (xCol[i][j], yColxCol[i][j]) define the co-located luminance sampling location in the image of the j-th direct reference layer for the luminance sampling location (xP, yP) in the i-th layer. The derivation of the variables (xCol[i][j], yColxCol[i][j]) is as follows:
[0254] xCol[i][j]=Clip3(0,(refPicWidthInSamplesL[i][j]–1),((xP-curScaledRefLayerLeftOffset[i][j])*scaleFactorX[i][j]+(1<<15))>>16))
[0255] yCol[i][j]=Clip3(0,(refPicHeightInSamplesL[i][j]–1),((yP-curScaledRefLayerT opOffset[i][j])*scaleFactorY[i][j]+(1<<15))>>16))
[0256] The derivation of variable colCtbAddr[i][j] is as follows:
[0257] xColCtb[i][j]=xCol[i][j]>>refCtbLog2SizeY[i][j]
[0258] yColCtb[i][j]=yCol[i][j]>>refCtbLog2SizeY[i][j]
[0259] colCtbAddr[i][j]=xColCtb[i][j]+(yColCtb[i][j]*refPicWidthInCtbsY[i][j])
[0260] When min_spatial_segment_offset_plus1[i][j] is greater than 0, the following conditions must be applied to ensure bitstream consistency:
[0261] If ctu_based_offset_enabled_flag[i][j] equals 0, then only one of the following will be used:
[0262] • In each PSS of the image assigned to the j-th direct reference layer of layer i, tiles_enabled_flag is equal to 0 and entropy_coding_sync_enabled_flag is equal to 0, and the following applies:
[0263] • Make slice segment A any slice segment of the image at layer i, and make ctbAddr the raster scan address of the last CTU in slice segment A. Make slice segment B a slice segment belonging to the same access unit as slice A (belonging to the j-th direct reference layer of layer i) and containing a CTU with raster scan address colCtbAddr[i][j]. Make slice segment C a slice segment located in the same image as slice segment B and immediately following slice segment B in the decoding order, and in the decoding order, the min_spatial_segment_offset_plus1[i]–1 slice segment exists between slice segment B and this slice segment. When slice segment C exists, the syntax elements of slice segment A are constrained such that any sample or syntax element value in slice segment C, or any slice segment or the same image immediately following C in the decoding order, can be used for inter-layer prediction during the decoding processing of any sample of slice segment A.
[0264] • In each PPS of the image assigned to the j-th direct reference layer of layer i, tiles_enabled_flag is equal to 1 and entropy_coding_sync_enabled_flag is equal to 0, and the following applies:
[0265] • Make slice A any slice in any picture picA of layer i, and make ctbAddr the raster scan address of the last CTU in slice A. Make slice B a slice in picture picB that belongs to the same access unit as picA and is in the j-th direct reference layer of layer i, and contains a CTU with raster scan address colCtbAddr[i][j]. Make slice C a slice that is also in picB and immediately follows slice B in decoding order, and in decoding order, the min_spatial_segment_offset_plus1[i]–1 slice exists between slice B and this slice. When slice segment C exists, the syntax elements of slice A are constrained such that any sample or syntax element value in slice C, or any slice or the same picture immediately following C in decoding order, can be used for inter-layer prediction during decoding of any sample in slice A.
[0266] • In each PPS of the image assigned to the j-th direct reference layer of layer i, tiles_enabled_flag is equal to 0 and entropy_coding_sync_enabled_flag is equal to 1, and the following applies:
[0267] • Make CTU line A any CTU line in any picture picA of layer i, and make ctbAddr the raster scan address of the last CTU in CTU line A. Make CTU line B a CTU line in picture picB that belongs to the same access unit as picA and to the j-th direct reference layer of layer i, and contains a CTU with raster scan address colCtbAddr[i][j]. Make CTU line C also located in picB and immediately following CTU line B, and in decoding order, min_spatial_segment_offset_plus1[i]-1 CTU line exists between CTU line B and this CTU line. When CTU line C exists, the syntax elements of CTU line A are constrained such that any sample or syntax element in CTU line C, or the line of the same picture immediately following C in decoding order, can be used for inter-layer prediction during decoding processing of any sample in CTU line A.
[0268] Otherwise (ctu_based_offset_enabled_flag[i][j] equals 1), and the following applies:
[0269] The derivation of variable refCtbAddr[i][j] is as follows:
[0270] xOffset[i][j]=((xColCtb[i][j]+minHorizontalCtbOffset[i][j])>(refPicWidthInCt bsY[i][j]-1))? (refPicWidthInCtbsY[i][j]-1–xColCtb[i][j]):(minHorizontalCtbOf fset[i][j])
[0271] yOffset[i][j]=(min_spatial_segment_offset_plus1[i][j]–1)*refPicWidthInCtbsY[i][j]
[0272] refCtbAddr[i][j]=colCtbAddr[i][j]+xOffset[i][j]+yOffset[i][j]
[0273] • Make CTUA any CTU in any image picA of layer i, and make ctbAddr the raster scan address ctbAddr of CTUA. Make CTUB the CTU in the image that belongs to the same access unit as picA and to the j-th direct reference layer of layer i, and has a raster scan address larger than refCtbAddr[i][j]. When CTUB exists, the syntax elements of CTUA are constrained such that any sample or syntax element value in CTUB is used for inter-layer prediction during decoding of any sample within CTUA.
[0274] In summary, the proposed implementation uses a flag to switch between CTB-based and spatial segment-based interlayer offset indications. The flag switches between a CTB-based indication and a spatial segment-based indication. Thus, alternative CTB-based indications can use unconditionally transmitted syntax elements, i.e., regardless of whether the interlayer offset indication is CTB-based or spatial segment-based. Specifically, if a signal for a spatial segment-based indication is issued, the syntax element indicating the interlayer offset within the spatial segment cell serves as a component such as a horizontal or vertical component for the positioning of the "trigger CTB" in the base layer image. Based on the syntax element switching between CTB-based and spatial segment-based interlayer offset indications, no further syntax elements are transmitted. Specifically, if a CTB-based indication signal is issued, further syntax elements are transmitted. In this case, the subsequent syntax element indicates a missing dimension in the location of the "trigger CTB". Therefore, the decoder can use two syntax elements to identify the "trigger CTB" according to the row and column arrangement of the CTBs in the base layer image. Once the CTB is decoded, the decoder can begin decoding the enhancement layer image. Nevertheless, the indication of either inter-layer offset can be completely switched by using one of the representative states of the first syntax element (i.e., min_spatial_segment_delay). Due to the predetermined decoding order between CTBs, in the case of CTB-based inter-layer offset indication, the decoder can still convert the horizontal and vertical components of the trigger CTB's location into the number of CTBs in the base layer image. Before starting the decoding of the first CTB of the enhancement layer image, the base layer image must be fully decoded so that the decoder can use the data. Figure 36 The implementation method described herein controls the compliance of inter-layer offsets with CTB indications.
[0275] Further aspects of the invention will be described in more detail below. A fourth aspect addresses the following problem: preferably, all partitioned network entities receiving the bitstream from the encoder are able to readily distinguish the various layers transmitted in a multi-layered data stream. For example, intermediate network entities may be interested in excluding specific information layers, such as layers with sampling resolutions exceeding a certain resolution threshold, from further transmission. The following description provides an overview of the current state of envisioned extensions to HEVC.
[0276] The Video Parameter Set (VPS) of HEVC[1] provides access to higher-order encoded bitstreams and contains information that is important for processing bitstreams in intermediate or end devices. The upcoming scalable and multi-view extensions of HEVC will further benefit from the VPS extension, which provides syntax for scalable bitstreams. One of the main tasks of the VPS extension is to provide a unified solution for interpreting nuh_reserved_zero_6bits in the NAL unit header. nuh_reserved_zero_6bits are intended to be re-signed as layer_id and used as a common layer identifier in scalable video coding scenarios. The following table shows the layer_id syntax elements in the NAL unit header and in [1] and as shown in [2]. Figure 2 The NAL unit header is shown below.
[0277] Two common solutions were considered in the design process [5]. The first solution, in the VPS model, is to signal the value of a single identifier in the header of the NAL unit to a potential plurality of scalable identifiers. The second solution, in the VPS extension, is to signal the individual bits (or bit blocks) of a single identifier in the header of the NAL unit to a specific scalable identifier.
[0278] The current VPS extension syntax reported in [4] is designed to use a mapping solution, however, it already contains all the syntax elements necessary for both solutions, namely, the two syntax elements that indicate the type of scalability (cp.scalability_map) and the number of layers for each scalability dimension (cp.dimension_id_len_minus1).
[0279] The mapping solution introduces further syntax elements into the VPS extended syntax, namely, the actual value of the scalable identifier encoded as u(v), which is optionally layer_id if the encoder chooses to sparsely assign the value of layer_id in a non-contiguous form.
[0280] In many very similar scalable scenarios, such as two or three spatial layers, two or three views, it is not always necessary to use all 63 values of the 6-bit layer identifier in the NAL unit header. For these scalable scenarios, the approach of assigning individual bits of the layer identifier in the NAL unit header to specific scalable dimensions has two advantages over mapping-based solutions.
[0281] • Regarding scalability dimension identifiers, the interpretation of layer identifier values in the NAL cell header does not require indirection or lookup methods.
[0282] • It does not require sending the VPS extension syntax element needed for the mapping solution, which is an important part of the VSP extension bits used for scalable signaling.
[0283] • Intermediate devices do not need to store mapping tables for each video bitstream.
[0284] Based on the fourth aspect of the concept described above, there may be hints in the higher-level syntax of HEVC indicating whether to use a mapping solution or a partitioning solution.
[0285] According to one implementation, based on a hint, a portion of the mapping related to the syntax element (cp.vps_nuh_layer_id_present_flag, layer_id_in_nuh[i], and dimension_id[i][j]) is sent or omitted, and a syntax element is signaled regarding the scalability type (cp.scalability_mask) and the amount of layers for each scalability (cp.dimension_id_len_minus1), and the syntax element is interpreted based on a hint that is information about the partition or mapping of the scalability identifier in the NAL unit header.
[0286] refer to Figure 23 An implementation method corresponding to or employing the concept of the fourth aspect of the present invention is proposed. Figure 23 A network entity is illustrated. This network entity can be any of the video decoders discussed above, or it can be an intermediate network entity generated between the encoder and decoder. Typically, reference numeral 680 is used to denote a network entity. The network entity is used to process multi-layer video data streams 682, such as any of the data streams 40 described above. In the case where network entity 680 is a video decoder, this processing will involve decoding multi-layer video data streams 628. In the case of an intermediate network entity, for example, the processing may involve forwarding the video data streams.
[0287] The scene is encoded into a multi-layer video data stream within each layer, such that the scene is encoded into different operational points in a scalable space expanded through a scalable dimension. The multi-layer video data stream is composed of first NAL units, each associated with a layer, and second NAL units are distributed within the first NAL units and present common information within the multi-layer video data stream. In other words, the first NAL unit 684 may carry one or more slices of a video image, and the "image" corresponds to any layer of the multi-layer video data stream 682.
[0288] In the above embodiments, only two layers are discussed for ease of description. Naturally, the number of layers can be greater than two, and even the classification of information regarding the participation of a layer in any preceding layer can differ between layers. In addition to the first two NAL units 684, NAL units 686 are shown scattered among the NAL units 684; however, unlike the first NAL units 684, their transmission can be performed via a separate channel. The second NAL unit presents common information in the multi-layer video data stream in a manner described in more detail below.
[0289] To describe in more detail the relationship between the first NAL unit on one side and the layer settings of the data stream 682 on the other side, refer to... Figure 24 . Figure 24 The first NAL unit 684, representing all first NAL units 684, is shown. The header 684 includes a layer indicator field 690. In addition to the header 688, the NAL unit 684 includes payload data 692 concerning the slice data discussed above with reference to other figures, i.e., data concerning the video content encoded using inter-layer prediction. Figure 24 The layer configuration (i.e., 694) is also shown. Specifically, Figure 24 The layer setting 694 shown should represent all possible layers, which can be represented and distinguished from each other by the layer indicator field 690 in the NAL unit 684. That is, the correlation function between one aspect of setting 694 and the possible values of layer indicator field 690 should be assumed to be a one-to-one mapping. Figure 24 In the diagram, small circles are used to exemplarily illustrate the various layers in setup 694, each circle bearing a specific number marked on it. While these markings suggest a defined order among the layers in setup 694, it should be noted that the following discussion will reveal that the arrangement or classification of layers in setup 694 is not solely derived from the layer indicator field 690. More precisely, given this, network entity 680 needs to detect the type indicator field 696 within the scattered second NAL units 686. However, this will be described later.
[0290] In other words, until now, Figure 24Each element of setting 694 represents only one possible state of the layer indicator field 690 in NAL unit 684. Layers in setting 694 can be distinguished by the layer indicator field 690; however, for network entity 680 without the additional information provided by the second NAL unit 686, the semantic meaning and order between these layers do not become clear. In practice, however, the layers in setting 694 form the nodes of a tree, and the branches between the trees correspond to specific scalability dimensions or axes. For example, a layer is the base layer and corresponds to the root of the tree. Each branch connects two layers, i.e., two nodes in the tree, describing how a particular layer contributes to another layer, i.e., providing additional information classification using inter-layer predictions. This information classification corresponds to the scalability dimension and can be, for example, an increase in spatial resolution, an increase in SNR, etc. For example, for simplification... Figure 24 A two-dimensional scalable space 698 is shown, which is extended by two scalability dimensions 700 and 702 corresponding to spatial scalability and SNR scalability. Figure 24 An exemplary tree machine with layers is shown in the diagram, extended through layer 698. Figure 24 It also shows that not all possible layers of configuration 694 can be presented in data stream 682. For example, in Figure 24 In the exemplary case, only five layers are actually used in data flow 682.
[0291] For example, Figure 22 The layer indicator field can have 6 bits, thus distinguishing between settings 694 and 2. 6 =64 possible states or possible layers. These possible values or layers from setting 694 can be adjusted via the second NAL unit 686 in a manner described in more detail below, to the mapping of operation points in the scalability space 698. Figure 24 The mapping is represented by reference symbol 704. An "operation point" should represent the location of at least the actual rendering layer within setting 694 within the scalable space 698. For example, the origin of the scalable space 698 can be associated with the base layer or the root, while each branch along either axis 700 or axis 702 of the scalable space 698 can have a fixed length of 1. Therefore, the vector pointing to the operation point in the scalable space 698 can have integer coordinate values.
[0292] In summary, the proposed description states that the multi-layer video data stream 682 provides information about the video content or context across multiple layers. The layers are arranged in a tree, with each layer connected to the tree via branches. Starting with the base layer forming the root, the next layer contributes to the reconstruction of video content information about a certain context, which can be interpreted as a scalability dimension. Therefore, each layer is either a root layer or connected to the latter via a specific path through a branch, and requires NAL units 684 belonging to layers arranged along that path to reconstruct the video content of the corresponding layer. Naturally, preferably, if mapping 704 is performed such that any "contributing" layer relative to the end of the corresponding branch leading from the root has a higher value for the layer indicator field 690 than the layer indicator field value of layers located near the end of the corresponding branch.
[0293] Figure 25 The meaning of type indicator field 696 is shown in more detail. Figure 25 A layer indicator field 690, which is a fixed-bit field, is shown. In any case, the length of field 690 is independent of the value of the type indicator field 696. However, if the type indicator field has a first state, the layer indicator field 690 is treated as a whole; that is, all n bits are processed together to distinguish its possible values. Preferably, if the type indicator field 696 assumes a first state, an n-bit integer is derived from the layer indicator field 690 by the network entity 680. In the case where the type indicator field 696 assumes a first state, the network entity 680 performs a mapping 704 from the possible values of the m-bit field 690 to the operation point using the mapping information passed within the second NAL unit 686. Figure 25As shown, for example, mapping information 708 includes a table of possible values for each actually used layer indicator field 690, with vectors pointing to the associated operation points of the corresponding possible values. That is, if the type indicator field 696 assumes a first state, network entity 680 can derive mapping information 708 from the second NAL unit 686 and can perform a lookup of mapping information or label 708 for each layer indicator field 690 to find the associated vector to locate the corresponding / associated operation point in space 698. The number of dimensions p of the vectors associated with the possible values in mapping information 708 can be a default setting or can be signaled within a data stream such as the second NAL unit 686. Subsequently, it is shown that the following signals can be emitted to convey information about mapping information 708: vps-_max_layers_minus1 determines the number M possible values M of the M-bit field 690 actually used. num_dimensions_minus1 limits the dimensions. The latter two variables can be sent using the syntax element scalability_mask. Then, the table itself can be signaled via a pair of possible values for a face (i.e., layer_id_in_nuh) and a p-dimensional vector (i.e., dimension_id[i][j]). Mapping 704 then directs to the p-dimensional vector, that is, a vector mapped via mapping information 708 to a corresponding vector 710, which points to the operation point within space 698 associated with the layer of the NAL cell having the layer indicator field 690.
[0294] However, if the type indicator field 696 assumes a second state, the mapping 704 is performed to a different degree. Specifically, in this case, the mapping is performed by dividing the layer indicator field 690 into more than one part, that is, by dividing the sequence of m-bit fields 690 into n consecutive subsequences of bits. Each part thus obtained is used as the coordinates x1…x1 of the n-dimensional vector 712. n Conversely, coordinates x1…x n This refers to the operation point associated with the layer within the scalable space 698 that includes the NAL unit containing the layer indicator field 690. For example, this is achieved by forming the first portion beyond the first (most important) m1 bits of field 690, the second portion beyond the subsequent m2 (less important) bits of field 690, and so on, up to m1+…+m… n =m, the nth part, divides the m bits of the layer indicator field 690 into n parts. The bits of each part can be directly treated as integer values.
[0295] If the type indicator field assumes a second state, the dimension n can be a default setting or signaled via the data stream. If p is determined using the specific implementation described herein after deriving n based on the same syntax elements, and if the type indicator field 696 assumes a first state, i.e., based on scalability_mask, then signaling is also exemplarily provided via the syntax element dimension_id_len_minus1: the bit length of the portion into which the layer indicator field 690 is subdivided (i.e., m1,…,m). n However, again, this segmentation can be achieved by default (without explicit sending).
[0296] It should be noted that parsable syntactic structures such as `scalability_mask` (unrelated to type indicator field 696) can indicate the number of dimensions and semantic meaning of the scalable space. However, because, for example, the maximum number of dimensions of the scalable space available in the mapping case (first state of the type indicator field) is higher than the maximum number available in the component interpretation case (first state of the type indicator field), the allowed states of a syntactic element may be limited to an appropriate subset of the possible states available to the syntactic element, relative to the mapping case, if type indicator field 696 indicates a component interpretation of the layer indicator field. The encoder only conforms to this restriction accordingly.
[0297] about Figures 23 to 25 The following are exemplary uses of the implementation methods described in the text:
[0298] Large-scale multi-party meeting
[0299] In large-scale RTP-based session service scenarios, such as web conferencing, the transmission of video between multiple parties is adapted in a multipoint control unit (MCU) that knows the set of parameters for the corresponding video bitstreams. Each party provides a thumbs-up bitstream and two bitstreams with enhanced spatial resolution (e.g., 720p and 4K) for the speaker. The MCU makes decisions about which streams to provide to which party. Therefore, the ease of parsing scalability parameters significantly reduces the burden on the MCU. Compared to mapping-based solutions for scalable signaling, partitioning-based solutions require fewer computational and memory resources.
[0300] Transmission system
[0301] In transmission systems such as RTP or MPEG2-TS, the mapping of scalability-related codec information to corresponding elements benefits from less complex bit-preservation mechanisms such as partitioning, compared to mapping-based solutions. Instead of using mapping-based scalability signaling, transmission systems more precisely address mapping indirectly and generate dedicated scalability identifiers for each scalability dimension, thus explicitly signaling the scalability dimensions in solutions such as partitioning.
[0302] Figure 26 The embodiments in [4] propose a possible implementation of syntax markers in VPS extension syntax that allows switching between a mapping-based solution and a partition-based solution for scalable signaling in extensions for HEVC based on VPS extension syntax as reported in [4].
[0303] A `dedicated_scalability_ids_flag` with a value of 1 indicates that the bits of the `layer_id` field in the NAL cell header are segmented according to the value of `dimension_id_len_minus1[]`, and the bits of the `layer_id` field belong to the corresponding scalability dimension that is signaled in the `scalability_mask`. A `dedicated_scalability_ids_flag` with a value of 1 indicates that no signal is emitted regarding the syntax elements `vps_nuh_layer_id_present_flag`, `layer_id_in_nuh[i]`, and `dimension_id[i][j]`. A `dedicated_scalability_ids_flag` with a value of 1 indicates that the derivation of the variable describing the scalability identifier of the corresponding NAL cell uses only the corresponding associated bits of the scalability identifier in the NAL cell header (`cp.layer_id`), for example, in the following form:
[0304] DependencyId = layer_id && 0x07
[0305] QualityId = layer_id && 0x38
[0306] Given a signal about `layer_id` emitted in the NAL unit header, a `dedicated_scalability_ids_flag` with a value of 0 indicates a signal about the syntax elements `vps_nuh_layer_id_present_flag`, `layer_id_in_nuh[i]`, and `dimension_id[i][j]`. The bits of `layer_id` in the NAL unit header are not associated with a specific scalability dimension, but are mapped to the scalability identifier in the VPS extension. A `dedicated_scalability_ids_flag` with a value of 0 indicates that the derivation of the variable describing the scalability identifier of the corresponding NAL unit uses the syntax element `dimension_id[i][j]`, for example, in the following form:
[0307] if(layer_id==layer_id_in_nuh[0]){
[0308] DependencyId=dimension_id[0][0]
[0309] QualityId=dimension_id[0][1]
[0310] When `dedicated_scalability_ids_flag` equals 0, `dimension_id_len_minus1[i]` indicates the length of bits in `dimension_id[i][j]`. When `dedicated_scalability_ids_flag` equals 1, `dimension_id_len_minus1[i]` indicates:
[0311] As indicated by scalability_mask, the number of bits in the layer_id in the NAL cell header associated with the i-th scalability dimension.
[0312] Figure 27 The flowchart further illustrates possible embodiments of the invention. The scalable identifier is derived directly from the bits of layer_id via masked bitcopy, or signaled in the VPS via association with a specific value of layer_id.
[0313] Figure 31Another syntax embodiment is illustrated below. Here, the type indicator field is signaled via “splitting_flag”, and the layer indicator field is referred to as nuh_layer_id. Based on “splitting_flag”, the operation point of the corresponding NAL cell layer within the scalable space is derived from nuh_layer_id using either a mapping concept or a partitioning concept. A partitioning concept is exemplarily signaled via a splitting_flag equal to 1. The scalable identifier, i.e., the vector component of the scalable dimension of the scalable space, can then be derived from the nuh_layer_id syntax element in the NAL cell header via bitmask copying. Figure 25 The corresponding bitmask for the i-th component (i-th scalability dimension) of vector 712 in the NAL unit is defined as follows. Specifically, a splitting_flag equal to 1 indicates the absence of a dimension_id[i][j] syntax element (i.e., no mapping information 708), and the binary identifier of the nuh_layer_id value in the NAL unit header is split into a series of NumScalabilityTypes, i.e., n segments, with a length x in bits according to the value of dimension_id_len_minus1[j]. 1…n Furthermore, the value of dimension_id[LayerIdxInVps[nuh_layer_id]][j] is inferred from the NumScalabilityType section of field 690, that is, the component x of vector 712. 1…n By issuing a series of flags, `scalability_mask_flag`, signals are sent regarding the syntactic meaning and number of scalable axes in the scalability space, thereby indicating to each of the predetermined number of scalable types (exemplarily fixed here) whether the corresponding scalability type belongs to any of the scalability dimensions of the scalability space 698. Specifically, network entity 680 can, according to... Figure 31 The for loop in the code derives the scalability space from the sequence marked with scalability_mask_flag, that is, the syntactic meaning and number of scalability axes NumScalabilityType:
[0314] for(i=0,NumScalabilityTypes=0;i<16;i++){
[0315] scalability_mask_flag[i]
[0316] NumScalabilityTypes+=scalability_mask_flag[i]
[0317] }
[0318] Here, a scalability_mask_flag[i] equal to 1 indicates the existence of the i-th scalability dimension, and a scalability_mask_flag[i] equal to 0 indicates the absence of the i-th scalability dimension. Here, i = 1 can represent multiple views (view scalability), i = 2 can represent spatial / SNR scalability, and i = 0 can represent the addition of depth mapping information. Other scalability dimension types may also exist; naturally, the example just summarized is merely illustrative. Assuming the length of nuh_layer_id is, for example, 6, partitioning or mask copying can be performed as follows:
[0319] The variable dimBitOffset[0] is set to 0, and the derivation of dimBitOffset[j] for j in the range of 1 to NumScalabilityTypes–1 (inclusive) is as follows:
[0320]
[0321] The value of dimension_id_len_minus1[NumScalabilityTypes-1] is presumed to be equal to 5-dimBitOffset[NumScalabilityTypes-1].
[0322] The value of dimBitOffset[NumScalabilityTypes] is set to 6.
[0323] Bitstream consistency is specified, that is, when NumScalabilityTypes is greater than 0, dimBitOffset[NumScalabilityTypes-1] should be less than 6.
[0324] For j ranging from 0 to NumScalabilityTypes–1 (inclusive), it is speculated that dimension_id[i][j] equals ((nuh_layer_id&((1<<dimBitOffset[j+1])-1))> >dimBitOffset[j]).
[0325] In summary, the following syntax elements relate each first NAL unit to an operation point within its scalability space:
[0326] 1) nuh_layer_id, i.e., the layer indicator field 690
[0327] 2) The sequence of scalability_mask_flag, that is, information on the number and meaning of scalable axes 700 and 702 in display space 698, and part of the x from this field. i The number n
[0328] 3) Each part of field 690 for each axis x i The dimension_id_len_minus1, i.e., the bit length (except for one, since it can be inferred that the remaining packet includes all the remaining bits of field 690 and 706).
[0329] 4) In addition, according to Figure 31 An optional implementation in the code sends vps_max_layers_minus1, which indicates that possible... The sequence of possible values for vps_max_layers_minus1 in some of the used / actual layers and the partition layer indicator field, i.e., the sequence of layer_id_in_nuh[i], thereby limiting the order of possible operation points.
[0330] If `splitting_flag` equals 0, then the mapping concept is used. Therefore, mapping information 708 is sent using the following information:
[0331] 1) The sequence of scalability_mask_flag and the number p of the components of the M vector in table 708, i.e., the information on the number and meaning of the scalable axes 700 and 702 in display space 698.
[0332] 2) The components x of the vector dimension_id[i][j] in Table 708 j The dimension_id_len_minus1, i.e., the bit length, is the length of each of the 698 axes in the space.
[0333] 3) Optionally, layer_id_in_nuh[i] is used as an index to a list of M vector dimensions_id[i][j].
[0334] 4) Optionally, although Figure 31 Not shown in the diagram, however, sending vps_max_layers_minus1 indicates the possible ∑ i 2 dimension_id_len_minus1[i] The number of used / actual layers M in some
[0335] Therefore, if splitting_flag equals 0, vector 710 is derived inherently from the partition without explicit signaling, i.e., without the signaling dimension_id[i][j] that is inferred.
[0336] Therefore, according to the fourth aspect, namely, the concept of "switchable interpretation of NAL unit layer identifiers" in bitstream 40 may include NAL units, namely, VPSNAL units, which include a type indicator field 300. Through the type indicator field 300, it is possible to switch between a mapping concept and a bit segmentation concept to interpret the layer indicator field 302 in the "normal" NAL unit. Thus, the same bit positions of field 302 are used in both modes; however, when a bitstream change occurs between modes, a signal is sent to transmit bit interpretation and interpretation instruction information, i.e., mapping information or segmentation and semantic information. Although this necessitates the additional transmission of the type indicator field, this concept together promotes more efficient bitstream transmission as an advantage of the mapping concept, and the bit segmentation concept can be adopted as needed, because both concepts can be applied to different multi-layer data to varying degrees, such as according to the number of layers.
[0337] The fifth aspect of this application relates to a multi-standard, multi-layer video decoder interface. The concepts set forth herein describe the interface between transport layer decoders and scalable video decoders (such as MPEG transport streams or RTP) that support different video coding standards in different layers (e.g., H.264 / AVC in the base layer and HEVC in the enhancement layer).
[0338] Scalable video bitstreams consist of layers: a base layer and one or more enhancement layers. The base layer contains an independent decoded video signal, and one or more enhancement layers can be decoded in combination with the base layer (and potentially other enhancement layers), providing higher temporal resolution (temporal scalability), spatial resolution (spatial scalability), quality (SNR scalability), higher bit depth (bit depth scalability), and video signal or other camera views (multi-view scalability).
[0339] Existing scalable video coding standards, such as H.264 / AVCSVC, define the base and enhancement layers within the same standard. The base and enhancement layers are designed so that the scalable bitstream has the same base format as the non-scalable bitstream. If the scalable bitstream is input into a non-scalable decoder, the packet type can still be identified and unknown packets can be discarded.
[0340] HEVC is the first video coding standard that allows the use of different video coding standards (e.g., H.264 / AVC) at the base layer. The packet formats of the two standards are different; therefore, the base layer decoder cannot understand enhancement layer packets. On the other hand, the enhancement layer decoder can understand the enhancement layer packet format, but cannot understand the base layer packet format.
[0341] In audio / video systems, the transport layer is used to combine several audio streams with video streams and provide metadata such as timing and stream type.
[0342] In existing multi-layer transport layer decoders, the access units of the base layer and enhancement layer are multiplexed into a single video data stream (e.g., an Annex B Byte stream for H.264 / AVC). This video stream is then input to the video decoder.
[0343] In all cases, if the base layer and enhancement layer use different video coding standards, the base layer packets and enhancement layer packets cannot be combined into a single bitstream.
[0344] According to the implementation of the fifth aspect, the transport layer decoder distinguishes the following cases:
[0345] 1. The output video decoder can only decode the basic layer.
[0346] 2. The output video decoder can decode both the base layer and the enhancement layer, and encode both layers using the same video coding standard.
[0347] 3. The output video decoder can decode the base layer and enhancement layer, and encode the base layer and enhancement layer using different video coding standards.
[0348] In case 1, the transport layer decoder uses the following behavior:
[0349] It extracts packets containing the base layer from the transport layer and inputs the packets into a single-standard, single-layer video decoder in the format specified in the video coding standard.
[0350] In a specific implementation, the transport stream decoder extracts H.264 / AVC NAL units from the MPEG-2 transport stream only those streams with the assigned stream type "AVC video stream or AVC video sub-stream that conforms to one or more profiles defined in Annex A of ITU-TRec.H.264|ISO / IEC 14496-10" and inputs the H.264 / AVC NAL units into the H.264 / AVC video decoder in the byte stream format defined in Annex B of the H.264 / AVC specification. The transport stream decoder discards NAL units belonging to streams whose stream type is not equal to "AVC video stream or AVC video sub-stream that conforms to one or more profiles defined in Annex A of ITU-TRec.H.264|ISO / IEC 14496-10". Figure 28 Examples of specific implementation methods are given.
[0351] In case 2, the transport layer decoder uses the following behavior:
[0352] Packets from the base and enhancement layers are extracted from the transport layer and fed into a single-standard multilayer video decoder in the format specified in the video coding standard.
[0353] In a specific implementation, H.264 / AVC NAL units can be extracted from the MPEG-2 transport stream by selecting a base layer stream with an assigned stream type "AVC video stream or AVC video sub-stream conforming to one or more profiles defined in Annex A of ITU-T Rec.H.264|ISO / IEC 14496-10" and one or more enhancement layers with an assigned stream type "SVC video sub-stream of AVC video stream conforming to one or more profiles defined in Annex G of ITU-T Rec.H.264|ISO / IEC 14496-10". The NAL units from different layers are multiplexed into a byte stream format defined in Annex B of the H.264 / AVC specification and input to the H.264 / AVC SVC video decoder. Figure 29 Examples of specific implementation methods are given.
[0354] In case 3, the transport layer decoder uses the following behavior:
[0355] Packets from the base and enhancement layers are extracted from the transport layer. These packets are processed in a multi-standard, multi-layer video decoder using one of the methods described in the following sections.
[0356] Interface A
[0357] If the enhancement layer standard packet format allows carrying base layer packets, then the base layer packets are encapsulated using the enhancement layer format. This means adding headers to each base layer packet that can be understood by the enhancement layer standard video decoder and allowing the enhancement layer video decoder to recognize the packets as base layers of different video coding standards.
[0358] In a specific embodiment of the invention, H.264 / AVC NAL units are used as the payload of HEVC NAL units; that is, an HEVC NAL unit header is added before the H.264 / AVC NAL unit header. The payload is identified as an H.264 / AVC NAL unit using fields in the HEVC NAL unit header (e.g., nal_unit_type). The bitstream can then be input to a video decoder in HEVC Annex B byte stream format. Figure 30 Examples of specific implementation methods are given.
[0359] Interface B
[0360] Different layers of a scalable bitstream use different channels. The video coding standard is determined in the decoder through channel selection.
[0361] In a specific embodiment of the invention, the two layers are switched in two separate channels. The first channel is used only for H.264 / AVC base layer packets (or AnnexB byte streams), while the second channel is used only for HEVC enhancement layer packets.
[0362] Interface C
[0363] Metadata fields indicating the type of video encoding standard are associated with the individual packets passed from the transport stream decoder to the multi-standard, multi-layer video decoder. Other metadata, such as timing data, can be signaled in the same way.
[0364] In a specific implementation, each base layer NAL unit is identified as an H.264 / AVC NAL unit through associated metadata fields, and each enhancement layer NAL unit is identified as an HEVC NAL unit through associated metadata fields.
[0365] Therefore, the concept of the fifth aspect can be described as a "multi-standard multi-layer video decoder interface," which provides possibilities for how to synthesize bitstreams from different codecs.
[0366] Therefore, according to the fifth aspect of this application, the transport layer decoder can be configured as follows (refer to the following). Figure 44 In summary. Typically, reference numeral 770 is used. Figure 44The transport layer decoder shown is illustrated. The transport layer decoder 770 is configured to pass the inbound multi-layer video data stream 40 through a scene encoded in the layer, thereby decoding it via a multi-standard multi-layer decoder 772 connected to the output interface of the transport layer decoder 770. The multi-layer video data stream 40 is composed of NAL units, which have been summarized above with reference to various other aspects of this application, the description of which can be transferred to... Figure 44 The implementation method is as follows: Each NAL unit is associated with a layer. Each layer is associated with a different codec (i.e., a different standard). For each layer, the NAL unit associated with the corresponding layer is encoded using the same codec, i.e., one NAL unit associated with the corresponding layer.
[0367] For each NAL unit, the transport layer decoder 770 is configured to identify the same codec associated with it and hand over the NAL units of the multi-layer video data stream 40 to a multi-standard multi-layer decoder, which uses inter-layer prediction between layers associated with different codecs to decode the multi-layer video data stream.
[0368] As described above, each NAL unit can be associated with a layer in the multi-layer video data stream 40 via a specific layer indicator field summarized above with reference to the fourth aspect of this application. Some or all of the NAL units can carry content-related data, i.e., one or more slices. By collecting all NAL units for a specific set of layers, the decoder 772 uses the amount of information granted by that set of layers to encode the video content or scene decoding of the data stream 40. For information on layer correlation, options with more than one scalability dimension, etc., refer to the description in the fourth aspect of this application.
[0369] The multi-layer multi-standard decoder 772 is capable of handling different codecs / standards. Examples of different standards, namely H.264 and HEVC, have been presented above; however, other standards can also be mixed. Different codecs / standards are not limited to hybrid codecs. More precisely, different kinds of mixed codecs can also be used. The inter-layer prediction used by the multi-layer multi-standard decoder 772 may involve prediction parameters used in different layers or may involve image sampling of various time-aligned layers. This has been described above with reference to other aspects and implementations.
[0370] The transport layer decoder 770 can be configured to perform handover to NAL units belonging to each layer of the encoder, and the multi-layer multi-standard decoder 772 can only handle this handover. That is, the handover performed by the transport layer decoder 770 may depend on the identifier of the transport layer decoder 770 of the encoder associated with each NAL unit. Specifically, the transport layer decoder 770 may perform the following on each NAL unit:
[0371] For example, the layer associated with the currently detected NAL cell can be identified by detecting the layer indicator field in the NAL cell header of the NAL cell.
[0372] On one hand, based on the correlation between the layers of data stream 40 and the correlation of the codecs / standards that the transport layer decoder 770 uses to detect the corresponding high-order syntax of data stream 40, the transport layer decoder 40 determines whether the currently detected NAL unit satisfies the following two criteria: the NAL unit layer belongs to a subset of the layers forwarded to decoder 772, determined by the operation points of the currently detected NAL unit layer in the scalable space, and external instructions regarding which operation points in the scalable space are allowed to be forwarded to the multi-layer multi-standard decoder 772 and which are not allowed to be forwarded to the multi-layer multi-standard decoder 772. Furthermore, the transport layer decoder 770 checks whether the codec of the currently detected NAL unit layer belongs to the set of codecs / standards that the multi-layer multi-standard decoder 772 can process.
[0373] If the check shows that the currently detected NAL unit meets both criteria, the transport layer decoder 770 forwards the current NAL unit to the decoder 772 for decoding.
[0374] For the transport layer decoder 770, there are different possibilities for determining the aforementioned correlations between the layers contained in the data stream 40 on one hand, and the codecs / standards based on those correlations on the other hand. For example, as described above, referring to "Interface B," different channels can be used to transmit the data stream 40 (i.e., NAL units of each layer of one codec / standard in one channel) and NAL units of each layer encoded according to another codec / standard on another channel. In this way, the transport layer decoder 770 can deduce the aforementioned correlations between the layers on one hand and the codecs / standards on the other hand by distinguishing the individual channels. For example, for each NAL unit of the data stream 40, the transport layer decoder 770 determines the channel to which the corresponding NAL unit arrives to identify the codec / standard associated with the corresponding NAL unit or the layer of the corresponding NAL unit.
[0375] Alternatively, the transport layer decoder 770 may forward NAL units belonging to different layers of different codecs / standards to the multi-layer multi-standard decoder 772 in a manner according to the corresponding codec / standard, such that NAL units belonging to one layer of one codec / standard are sent to the decoder 772 through one channel, and NAL units belonging to different layers of different codecs / standards are forwarded to the multi-layer multi-standard decoder 772 through another channel.
[0376] The basic transport layer can provide "different channels." That is, for ease of understanding, this is achieved by distinguishing between channels created by different transport layers. Figure 44 The different channel identifiers provided by the underlying transport layer (not shown) enable the differentiation between different channels.
[0377] Another possibility for transferring data stream 40 to multi-layer multi-standard decoder 772 is that transport layer decoder 770 encapsulates NAL units identified as associated with a layer (associated with any codec other than the predetermined codec) using a NAL unit type indicator field of a predetermined codec, with the NAL unit type indicator field set to indicate the state of the codec for the corresponding layer. This means, for example, that the predetermined codec can be any codec for any enhancement layer of data stream 40. For example, the base layer codec (i.e., the codec associated with the base layer of data stream 40) can be different from the predetermined codec (e.g., HEVC). Therefore, when data stream 40 is passed through multi-layer multi-standard decoder 722, transport layer decoder 770 can convert data stream 40 into a data stream conforming to the predetermined codec. In view of this, transport layer decoder 770 encapsulates each NAL unit belonging to a layer not encoded using the predetermined codec using a NAL unit type indicator field of the predetermined codec and sets the NAL unit type indicator field within the NAL unit type indicator field to indicate the state of the codec for the corresponding actual layer. For example, a base layer NAL cell is a base layer H.264 NAL cell that is appropriately encapsulated using HEVC and a NAL cell header (with a NAL cell type indicator set to indicate the H.264 status). Therefore, the multi-layer multi-standard 772 will receive HEVC that matches the data stream.
[0378] Naturally, as described in Reference Interface C, alternatively, the transport layer decoder 770 can provide metadata indicating the layer associated with the corresponding NAL unit and its associated codec for each NAL unit of the inbound data stream 40. Thus, the 40 NAL units of the data stream will be forwarded to the decoder 772 in this extended manner.
[0379] Using the alternative method just described, it is easy to extend the content encoded into a data stream by further layers; however, using another encoder, such as a new encoder, will further encode the layers without requiring modification to the existing parts of codec 40. The multi-layer, multi-standard decoder can then handle the new codec (i.e., the newly added one), and can process existing mixed data streams using the layers encoded with the new codec.
[0380] Therefore, the above presents the concept of parallel / low-latency video coding for HEVC scalable bitstreams.
[0381] The High Efficiency Video Coding (HEVC) standard [1] was originally characterized by two dedicated parallel tools that allow parallel processing at both the encoder and decoder ends: slice and wavefront parallel processing (WPP). These tools allow parallel processing within each picture when compared to HEVC-coded video, which is not characterized by parallel processing within each picture, with the goal of improving processing time while minimizing the loss of coding efficiency.
[0382] In scalable[2] HEVC bitstreams or multi-view[3] HEVC bitstreams, the base layer or base view image is used to predict the enhancement layer or related view image. In the above description, the term "layer" is used to also cover the concept of a view.
[0383] The above implementation describes a scalable video decoder capable of starting the decoding of enhancement layer images before the decoding of the associated base layer images is completed. Image region decoding is performed based on high-order parallel tools used in each layer. The base layer decoder and enhancement layer decoder can operate in parallel with each other and also in parallel with the actual layers. The amount of parallelism within each layer between the base and enhancement layers can be different. Furthermore, signaling specifying appropriate settings for the parallel decoding environment of the specific bitstream is described.
[0384] As a general note, the following should be observed: The above embodiments describe the decoder and correspondingly design the encoder according to various aspects. It should be noted that, whenever these aspects are involved, they all likely share commonalities: the decoder and encoder support WPP and / or slice-parallel processing and correspondingly describe details related to them. These details should be considered applicable simultaneously to any other aspect and its corresponding description, guiding new implementations of these other aspects or supplementing the descriptions of implementations of these other aspects, regardless of the corresponding aspects described using terms such as "part," "spatial segment," etc. (rather than slices / substreams of a more general representation of a picture's parallel processing segment) (the corresponding description is transferred to that corresponding aspect). The same applies to the details regarding the encoding / prediction parameters and descriptions of possible ways of subdividing the picture: all aspects can be implemented to produce a decoder / encoder that uses the subdivision of the LCU / CTB by determining slices and / or substreams within a unit of the LCU / CTB. Thus, through any of these aspects, the LCU / CTB can be further subdivided into coding blocks using the recursive multi-tree subdivision described above with reference to subsets of the various aspects and their implementations. In addition, or alternatively, the implementation of all aspects may adopt the concept of slices, referring to the relationship between slices and sub-streams / slices described in these aspects.
[0385] Although some aspects have been described in the context of the apparatus, it is clear that these aspects also represent a description of the corresponding method, where a block or device corresponds to a method step or a feature of a method step. Similarly, aspects described in the context of method steps also represent a description of a corresponding block or an item or feature of a corresponding apparatus. Some or all of the method steps can be executed by (or using) a hardware device, such as a microprocessor, a programmable computer, or electronic circuitry. In some embodiments, one or more of the most important method steps can be executed by the device.
[0386] According to a specific implementation, embodiments of the invention can be implemented in hardware or software. This implementation can be executed using a digital storage medium, such as a floppy disk, DVD, Blu-ray disc, CD, ROM, PROM, EPRO, EEPROM, or flash memory, having electrically readable control signals stored thereon that cooperate (or are capable of cooperating with) a programmable computer system to execute the corresponding method. Therefore, the digital storage medium can be computer-readable.
[0387] Some embodiments of the invention include a data carrier having electrically readable control signals, the data carrier being capable of cooperating with a programmable computer system to perform one of the methods described herein.
[0388] Typically, embodiments of the present invention can be implemented as a computer program product having program code, which, when run on a computer, can operate to perform one of the methods. For example, the program code can be stored on a machine-readable medium.
[0389] Other embodiments include a computer program stored on a machine-readable medium for performing one of the methods described herein.
[0390] Therefore, in other words, the implementation of the method of the present invention is a computer program, which, when run on a computer, has program code for performing one of the methods described herein.
[0391] Therefore, a further embodiment of the method of the present invention is a data carrier (or digital storage medium or computer-readable medium) including a computer program recorded thereon for performing one of the methods described herein, and the data carrier, digital storage medium or recording medium is generally volatile and / or non-volatile.
[0392] Therefore, a further embodiment of the method of the present invention is a data stream or signal sequence representing a computer program for performing one of the methods described herein. For example, the data stream or signal sequence may be configured to be transmitted via a data communication connection (e.g., via the Internet).
[0393] Further implementations include processing means (e.g., a computer) or programmable logic devices configured or adapted to perform one of the methods described herein.
[0394] A further embodiment includes a computer having a computer program installed thereon for performing one of the methods described herein.
[0395] Further embodiments of the invention include an apparatus or system configured to transmit a computer program (e.g., electrically or optically) for performing one of the methods described herein to a receiver. For example, the receiver may be a computer, a mobile device, a memory device, etc. For example, the apparatus or system may include a file server for transmitting the computer program to the receiver.
[0396] In some embodiments, some or all of the functions of the methods described herein may be performed using a programmable logic device (e.g., a field-programmable gate array). In some embodiments, the field-programmable gate array may cooperate with a microprocessor to perform one of the methods described herein. In summary, preferably, the methods are performed by any hardware device.
[0397] The apparatus described herein can be implemented using hardware devices, or using a computer, or a combination of hardware devices and a computer.
[0398] The methods described herein can be performed using hardware devices, or using a computer, or a combination of hardware devices and a computer.
[0399] The embodiments described above merely illustrate the principles of the invention. It should be understood that variations and modifications of the arrangements and details described herein will be apparent to those skilled in the art, and therefore are intended to be limited only by the scope of the pending patent claims, and not by the specific details presented in the description and illustration of the embodiments herein.
[0400] According to the present invention, it may include the following:
[0401] 1. A video decoder for decoding a multi-layer video data stream (40) into scenes encoded in the layers using inter-layer prediction from the first layer to the second layer, the video decoder supporting parallel decoding of the multi-layer video data stream in spatial segments (80) into which the images (12, 15) of the layers are subdivided, wherein the decoder is configured to:
[0402] Detect the long-term syntax element structure (606; e.g., tile_boundaries_aligned_flag) of the multi-layer video data stream to
[0403] Deciphering the long-term syntax element structure, which assumes that values outside the first set of possible values (e.g., tile_boundaries_aligned_flag = 1) are used as a guarantee for subdividing the second-layer image (15) within a predetermined time period (608) such that the boundaries between spatial segments of the second-layer image overlap with each boundary of a spatial segment of the first-layer image (12), and periodically determining the subdivision of the first-layer and second-layer images into spatial segments based on short-term syntax elements (602; e.g., column_width_minus1[i] and column_width_minus1[i]) of the multi-layer video data stream within a time interval (604) shorter than the predetermined time period; and
[0404] If the long-term syntax element structure assumes a value outside the second possible value set (e.g., tile_boundaries_aligned_flag = 0), then the short-term syntax elements from the multi-layer video data stream are periodically used to determine the subdivision of the layer's images into spatial segments at time intervals shorter than the predetermined time period, such that, at least for a first possible value of the short-term syntax element, there exists a boundary between the spatial segments of the second layer's images that does not overlap with any of the boundaries of the spatial segments of the first layer, and at least for a second possible value of the short-term syntax element, there exists a boundary between the spatial segments of the second layer's images that overlaps with each boundary of the spatial segments of the first layer.
[0405] 2. The video decoder according to claim 1, wherein the video decoder is configured to:
[0406] Intra-frame spatial prediction is used to decode the images of the layer, and the intra-frame spatial prediction for each spatial segment is interrupted at the boundary line of the corresponding spatial segment; or
[0407] The image of the layer is decoded using the intra-frame spatial prediction method, which decodes the spatial segments of the first layer image in parallel and supports the boundary lines of the spatial segments of the first layer image, and conforms to the decoding delay between the decoding of the spatial segments of the first layer image; and the image of the second layer is decoded using the intra-frame spatial prediction method, which decodes the spatial segments of the second layer image in parallel and supports the boundary lines of the spatial segments of the second layer image, and conforms to the decoding delay between the decoding of the spatial segments of the second layer image.
[0408] 3. The video decoder according to item 1 or 2, supporting slice-parallel decoding of the multi-layer video data stream in slices formed by subdividing the images of the layer, wherein the decoder is configured to:
[0409] Deciphering the long-term syntax element, which assumes values outside the first possible set as a guarantee for subdividing the second layer of the image within a predetermined time period, such that the boundaries between the slices of the second layer image overlap with each boundary of the slices of the first layer, and periodically determining the slice refinement of the second layer image relative to the first layer based on the short-term syntax element within time intervals shorter than the predetermined time period; and
[0410] If the long-term syntax element assumes a value outside the second set of possible values, then the short-term syntax element of the multi-layer video data stream periodically determines the subdivision of the layer's images into slices within a time interval shorter than the predetermined time period, such that, at least for a first possible value of the short-term syntax element, there exists a boundary between slices of the second layer's images that does not overlap with any of the boundaries of slices of the first layer, and at least for a second possible value of the short-term syntax element, there exists a boundary between slices of the second layer's images that overlaps with each boundary of a slice of the first layer.
[0411] 4. The video decoder according to item 3, wherein the video decoder is configured to:
[0412] The intra-frame spatial prediction is used to decode the images of the layer, and the intra-frame spatial prediction for each slice is interrupted at the boundary line of the corresponding slice.
[0413] 5. The video decoder according to item 1 or 2, wherein the decoder is configured to:
[0414] Deciphering the long-term syntax element structure, which assumes that values outside the first set of possible values guarantee the subdivision of the second-layer image within a predetermined time period, such that each spatial segment of the first-layer image is precisely composed of n spatial segments of the second-layer image, and n depends on the value of the long-term syntax element structure; and
[0415] If the long-term syntax element is set to a value outside the second possible value set, then the inter-layer offset within the predetermined time period is periodically determined based on the short-term syntax element of the multi-layer video data stream within a time interval shorter than the predetermined time period.
[0416] 6. The video decoder according to item 1 or 2, wherein the decoder is configured to determine whether to begin or not to begin attempting to decode the second layer of the multi-layer video data stream based on whether the long-term syntax element structure assumes values other than the first possible value.
[0417] 7. The video decoder according to any one of items 1 to 6, wherein the video decoder is a hybrid video decoder.
[0418] 8. A video encoder for encoding a scene into a multi-layer video data stream at the layer level using inter-layer prediction from a first layer to a second layer, such that the multi-layer video data stream can be decoded in parallel within spatial segments subdivided from the images of the layer, wherein the encoder is configured to:
[0419] Long-term syntax element structure (606) and short-term syntax element (602) are inserted into the multi-layer video data stream, wherein the short-term syntax element defines the subdivision of the first-layer and second-layer images into spatial segments within a time interval; and
[0420] Switch between setting the long-term syntax element structure to the following values:
[0421] Values outside the first possible set, and within a predetermined time period (608) longer than the time interval, the short-term syntax element is set to an appropriate subset outside the possible set, the appropriate subset being selected such that the second layer image is subdivided within the predetermined time period, such that the boundaries between spatial segments of the second layer image overlap with each boundary of a spatial segment of the first layer; or
[0422] Values other than the second set of possible values, and setting the short-term syntax element to any one of the set of possible settings within the predetermined time period, the set of possible settings including at least one setting and at least another setting, according to the at least one setting, there are boundaries between spatial segments of the second layer of the image that do not overlap with any of the boundaries of the spatial segments of the first layer, and according to the at least another setting, there are boundaries between spatial segments of the second layer of the image that overlap with each boundary of the spatial segments of the first layer.
[0423] 9. The video encoder according to claim 8, wherein the video encoder is configured to:
[0424] The images of the layer are encoded using intra-frame spatial prediction, and the intra-frame spatial prediction for each spatial segment is interrupted at the boundary line of the corresponding spatial segment; or
[0425] The images of the layer are encoded using intra-frame image spatial prediction that supports the boundary lines of spatial segments of the first layer's images and initializes the entropy context probability for entropy encoding a subset of spatial segments of the first layer's images individually, or based on the entropy context probability of the previous subset of spatial segments of the first layer's images in an intermediate adaptation state, according to the order between the subsets; and the images of the layer are encoded using intra-frame image spatial prediction and entropy encoding adapted to the entropy context probability, based on the entropy context probability of the previous subset of spatial segments of the second layer's images in an intermediate adaptation state, by supporting intra-frame image spatial prediction that supports the boundary lines of spatial segments of the second layer's images and initializing the entropy context probability for entropy encoding a subset of spatial segments of the second layer's images individually, or based on the entropy context probability of the previous subset of spatial segments of the second layer's images in an intermediate adaptation state, according to the order between the subsets.
[0426] 10. The video encoder according to item 8 or 9, wherein the spatial segment is a piece, and
[0427] The encoder is configured as follows:
[0428] If the structure of the long-term syntax element is set as follows:
[0429] For values outside the first set of possible values (tile_boundaries_aligned_flag = 1), the short-term syntax element is set to an appropriate subset outside the set of possible settings for a predetermined time period longer than the time interval. The appropriate subset is selected such that the image of the second layer subdivided into pieces is consistent with the image of the first layer subdivided into pieces for the predetermined time period, or the image of the first layer subdivided into pieces is refined.
[0430] or
[0431] If a value other than the second possible value set (tile_boundaries_aligned_flag = 0) is used, then the short-term syntax element is set to any one of the possible settings within the predetermined time period, such that for at least one time interval within the predetermined time period, the short-term syntax element is set to a first possible value of the possible settings, according to the first possible value, where there are boundaries between tiles of the second layer that do not overlap with any of the boundaries of tiles of the first layer; and for at least another time interval within the predetermined time period, the short-term syntax element is set to a second possible value of the possible settings, according to the second possible value, where there are boundaries between tiles of the second layer that overlap with each boundary of a tile of the first layer.
[0432] 11. The video encoder of claim 10, wherein the video encoder is configured to encode images of the layer using the intra-frame picture spatial prediction, and the intra-frame picture spatial prediction for each slice is interrupted at the boundary line of the respective slice.
[0433] 12. The video encoder according to item 8 or 9, wherein the encoder is configured to:
[0434] If the structure of the long-term syntax element is set as follows:
[0435] For values outside the first set of possible values, the short-term syntax element is set to an appropriate subset outside the set of possible settings using a predetermined time period longer than the time interval. The appropriate subset is selected such that, within the predetermined time period, each spatial segment of the first layer's image is precisely composed of n spatial segments of the second layer's image, where n depends on the value of the long-term syntax element structure.
[0436] 13. A decoder for decoding a spatially scalable bitstream (40) into an image, the image being encoded in different spatial layers, and for at least one of the spatial layers, the image being encoded in a first spatial segment, wherein the decoder is configured to:
[0437] The image (12) of the first spatial layer is upsampled to obtain an upsampled reference image, and the image (15) of the second spatial layer is predicted using the upsampled reference image, wherein,
[0438] The decoder responds to syntax elements (616; e.g., independent_tile_upsampling_idc) in the spatially scalable bitstream to insert (620) according to the syntax elements.
[0439] The image of the first spatial layer,
[0440] This makes any portion of a partition (622) of the upsampled reference image that depends on the first spatial segment independent of portions of the first spatial layer image covered by any of the other portions of the partition; or
[0441] This causes any portion of the partition (622) of the upsampled reference image to depend on a portion of the first spatial layer image covered by another portion of the partition that is spatially adjacent to the corresponding portion.
[0442] 14. The decoder according to item 13, wherein the decoder is configured to decode the different spatial layers in parallel.
[0443] 15. The decoder according to item 13 or 14, wherein the decoder responds to the syntax element (616) in the spatially scalable bitstream to insert (620) an image of the first spatial layer according to the syntax element.
[0444] Such that any portion of the upsampled reference image spatially covered by any of the first spatial segments is independent of a portion of the image of the first spatial layer covered by another of the first spatial segments; or
[0445] Such that any portion of the upsampled reference image spatially covered by any of the first spatial segments depends on a portion of the image of the first spatial segment covered by another spatial segment adjacent to the corresponding spatial segment.
[0446] 16. The decoder according to any one of items 13 to 15, wherein the spatially scalable bitstream has an image of the second spatial layer encoded in the first spatial segment into the spatially scalable bitstream.
[0447] 17. The decoder according to any one of claims 13 to 16, wherein the decoder is configured to perform decoding using intra-frame picture space prediction.
[0448] And the intra-frame spatial prediction used for each first spatial segment is interrupted at the boundary line of the corresponding first spatial segment; or
[0449] Furthermore, it supports intra-frame image spatial prediction across the boundary lines of the first spatial segment, uses the adaptability of entropy context probability to entropy decode each first spatial segment and initializes the entropy context probability of the first spatial segment, the entropy context probability of the first spatial segment being independent of any other first spatial segment, or adapting the entropy context probability of the previous first spatial segment according to the order within the first spatial segment to adapt to the middle position of the previous first spatial segment.
[0450] 18. The decoder according to item 13 or 14, wherein the spatially scalable bitstream has an image of the first spatial layer encoded in the first spatial segment of the spatially scalable bitstream, wherein the spatially scalable bitstream has an image of the second spatial layer encoded in the second spatial segment of the spatially scalable bitstream, wherein the decoder responds to the syntax element (606) in the spatially scalable bitstream to interpolate the image of the first spatial layer according to the syntax element.
[0451] (e.g., independent_tile_upsampling_idc = 2), such that any portion of the upsampled reference image spatially covered by any one of the second tiles is independent of a portion of the image of the first spatial layer spatially covered by another of the second tiles; or
[0452] (For example, independent_tile_upsampling_idc = 1), such that any portion of the upsampled reference image spatially defined by the spatial co-location boundary line of the first and second tiles is independent of a portion of the image of the first spatial layer covered by another portion of the upsampled reference image spatially defined by the spatial co-location boundary line of the first and second tiles; or
[0453] (e.g., independent_tile_upsampling_idc = 0), such that any portion of the upsampled reference image spatially covered by any of the second tiles depends on a portion of the image of the first spatial layer covered by another second tile adjacent to the corresponding tile.
[0454] 19. The decoder according to any one of claims 13 to 18, wherein the decoder is configured to fill fragments of a filter kernel used in an image interpolated from the first spatial layer using a backoff rule, the filter kernel extending into any portion of the image of the first spatial layer and covered by any other partition of the partition, the fragments being filled independently of the corresponding portion of the image of the first spatial layer into which the filter kernel extends, according to the backoff rule, to achieve independence.
[0455] 20. The decoder according to item 19, wherein the decoder is configured to use the backoff rule while still filling fragments of the filter kernel extending beyond the outer boundary of the image of the first spatial layer.
[0456] 21. The decoder according to any one of items 13 to 20, wherein the decoder is a video decoder and the decoder is configured to respond to the syntax element (606) in the spatially scalable bitstream on a per-picture or per-picture sequence basis.
[0457] 22. The decoder according to any one of claims 13 to 21, wherein the spatially scalable bitstream has an image of the first spatial layer encoded in the first spatial segment as the first spatial layer of the spatially scalable bitstream, wherein the spatially scalable bitstream has an image of the second spatial layer encoded in the second spatial segment as the second spatial layer of the spatially scalable bitstream, wherein,
[0458] The boundary of the partition corresponds to the logical AND of the spatial overlap between the boundary of the first spatial segment and the boundary of the second spatial segment or the boundary of the second spatial segment, wherein the decoder responds to the syntax element (606) in the spatially scalable bitstream according to the syntax element to fill the fragment of the filter kernel used in the interpolation of the image of the first spatial layer, the filter kernel extending from one partition to the adjacent partition of the partition using a backoff rule, filling the fragment independently of the corresponding portion of the image of the first spatial layer to which the filter kernel extends, or using the corresponding portion of the image of the first spatial layer to which the filter kernel extends.
[0459] 23. The decoder according to any one of items 13 to 22, wherein the decoder is configured to decode the first layer and the second layer in parallel using inter-layer offsets dependent on the syntax element (606).
[0460] 24. The decoder according to item 13 or 23, wherein the decoder is configured to place the boundaries of the partitions according to the syntax elements to correspond to a logical AND of the spatial overlap of the boundaries of the first spatial segment and the second spatial segment or the boundaries of the second spatial segment.
[0461] 25. An encoder for encoding an image into a spatially scalable bitstream in different spatial layers and encoding the image in a first spatial segment for at least one of the spatial layers, wherein the encoder is configured to:
[0462] An image of the first spatial layer is upsampled to obtain an upsampled reference image, and the upsampled reference image is used to predict an image of the second spatial layer. The encoder is configured to set and insert syntax elements (606) into the spatially scalable bitstream according to syntax elements and interpolate the image of the first spatial layer.
[0463] Such that any portion of a partition of the upsampled reference image dependent on the first spatial segment is independent of portions of the first spatial layer image covered by any of the other portions of the partition; or
[0464] This causes any portion of the partition of the upsampled reference image to depend on a portion of the first spatial layer image covered by another portion of the partition that is spatially adjacent to the corresponding portion.
[0465] 26. The decoder according to item 25, wherein the encoder is configured to set according to the syntax element and insert the syntax element into the spatially scalable bitstream and interpolate the image of the first spatial layer.
[0466] Such that any portion of the upsampled reference image covered by any space in the first spatial segment is independent of any portion of the image of the first spatial layer covered by any other portion of the first spatial segment; or
[0467] Such that any portion of the upsampled reference image covered by any one of the first spatial segments depends on a portion of the image of the first spatial layer covered by another spatial segment of the first spatial segment adjacent to the corresponding first spatial segment.
[0468] 27. The encoder according to item 25 or 26, wherein the encoder is configured to encode an image of the first spatial layer into the spatially scalable bitstream in the first spatial segment.
[0469] 28. The encoder according to claim 27, wherein the encoder is configured to encode images of the first spatial layer using intra-frame image spatial prediction, and the intra-frame image spatial prediction for each first spatial segment is interrupted at the boundary line of the corresponding first spatial segment.
[0470] 29. The encoder according to item 27 or 28, wherein the encoder is configured to encode an image of the second spatial layer into the spatially scalable bitstream in a second spatial segment, wherein the encoder is configured to set according to the syntax element and insert the syntax element into the spatially scalable bitstream and interpolate the image of the first spatial layer.
[0471] Such that any portion of the upsampled reference image spatially covered by any one of the second spatial segments is independent of a portion of the image of the first spatial layer spatially covered by another of the second spatial segments; or
[0472] Such that any portion of the upsampled reference image spatially defined by the spatially co-located boundary line of the first spatial segment and the second spatial segment is independent of a portion of the image of the first spatial layer spatially defined by the spatially co-located boundary line of the first spatial segment and the second spatial segment and covered by another partition of the upsampled reference image; or
[0473] This causes any portion of the upsampled reference image spatially covered by any of the second spatial segments to depend on a portion of the image of the first spatial layer covered by another second spatial segment adjacent to the corresponding spatial segment.
[0474] 30. The encoder according to claim 29, wherein the encoder is configured to encode images of the second spatial layer using intra-frame image spatial prediction, and the intra-frame image spatial prediction for each second spatial segment is interrupted at the boundary line of the respective second spatial segment.
[0475] 31. The encoder according to any one of claims 25 to 30, wherein the encoder is configured to use a fill fragment of a filter kernel to achieve independence from any portion of the image of the first spatial layer when interpolating an image of the first spatial layer, the filter kernel extending into any portion of the image of the first spatial layer using a back-back rule, the fill fragment being independent of the corresponding portion of the image of the first spatial layer into which the filter kernel extends, according to the back-back rule.
[0476] 32. The encoder according to item 31, wherein the encoder is configured to use the backoff rule while still filling fragments of the filter kernel extending beyond the outer boundary of the image of the first spatial layer.
[0477] 33. The encoder according to any one of items 25 to 32, wherein the encoder is a video encoder and the encoder is configured to set on a per-picture or per-picture sequence basis and insert the syntax elements into the spatially scalable bitstream.
[0478] 34. A video decoder for decoding a multi-layer video data stream into scenes encoded in layers using inter-layer prediction from a first layer to a second layer, the video decoder utilizing the inter-layer delay between the traversal of spatial segments of images of the first layer and the traversal of spatial segments of images of the second layer, sequentially traversing spatial segments in a temporally overlapping manner to support parallel decoding of the multi-layer video data stream in spatial segments partitioned from the images of the layers, the video decoder being configured to:
[0479] Detecting the long-term syntax element structure (642; e.g., min_spatial_segment_delay) of the multi-layer video data stream such that:
[0480] If the long-term syntax element structure (e.g., min_spatial_segment_delay≠0) is set to a value in the first set of possible values, then the inter-layer offset within a predetermined time period is predetermined using the value of the long-term syntax element, and the size and location of the spatial segments of the first layer image and the second layer image, as well as the spatial sampling resolution of the first layer image and the second layer image, are periodically determined based on the short-term syntax element (602) of the multi-layer video data stream in time intervals shorter than the predetermined time period.
[0481] If the long-term syntax element is set to a value in a second set of possible values that do not intersect with the first set of possible values (e.g., min_spatial_segment_delay = 0), then the short-term syntax element based on the multi-layer video data stream periodically determines the inter-layer offset within the predetermined time period in time intervals shorter than the predetermined time period.
[0482] 35. The video decoder according to claim 34, wherein the video decoder is configured to utilize the inter-frame picture substream delay between traversals of consecutive substreams of the same picture and the inter-layer offset of the traversal of the first layer's picture substream relative to the traversal of the second layer's picture substream, to decode the multi-layer video data stream in parallel through wavefront parallel processing of the substreams in a temporally overlapping manner into substreams composed of rows of blocks formed by partitioning the pictures of the layer and regularly subdividing the pictures of the layer.
[0483] 36. The video decoder according to claim 35, wherein the video decoder is configured to:
[0484] The substreams are decoded in parallel and intra-frame picture space prediction is supported across the boundaries of the substreams.
[0485] 37. The video decoder according to claim 34, wherein the video decoder is configured to decode the multi-layer video data stream into slices of images partitioned from the layers, and to traverse the slices in slice order between the slices within each of the images of the first layer and the second layer, and to decode the consecutive slices of images partitioned from the layers in parallel using an inter-layer offset of the traversal of slices of the first layer images relative to the traversal of slices of the second layer images.
[0486] 38. The video decoder according to claim 37, wherein the video decoder is configured to:
[0487] Intra-frame spatial prediction is used to decode the images of the first layer and the second layer, and the intra-frame spatial prediction for each slice is interrupted at the boundary line of the corresponding slice.
[0488] 39. The video decoder according to any one of claims 34 to 38, wherein the video decoder is configured to use the value of the long-term syntax element when determining the inter-layer offset by using the value of the long-term syntax element as a measurement of the inter-layer offset in a unit of a spatial segment of a picture in the first layer.
[0489] 40. The video decoder according to any one of claims 34 to 39, wherein the video decoder is configured to use the value of the long-term syntax element when determining the inter-layer offset by using the value of the long-term syntax element as the number of spatial segments of the first layer's picture, wherein the decoding of the first spatial segment of the second layer's picture by the value is delayed relative to the decoding and traversal of the first layer's picture.
[0490] 41. A video encoder for encoding a multi-layer video data stream using inter-layer predictions from a first layer to a second layer, wherein the multi-layer video data stream can be decoded in parallel within spatial segments partitioned by the layers by sequentially traversing spatial segments in a temporally overlapping manner, utilizing inter-layer offsets between the traversal of spatial segments of the first layer's images and the traversal of spatial segments of the second layer's images, the video encoder being configured to:
[0491] Set a long-term syntax element structure (min_spatial_segment_delay) and a short-term syntax element, and insert the long-term syntax element structure and the short-term syntax element into the multi-layer video data stream. The short-term syntax element of the multi-layer video data stream periodically defines the size and location of the spatial segment of the first layer image and the spatial segment of the second layer image within the time interval, as well as the spatial sampling resolution of the first layer image and the second layer image.
[0492] The encoder is configured to switch between the following settings:
[0493] The long-term syntax element structure (min_spatial_segment_delay≠0) is set to a value from a first set of possible values, and this value sends the signaling of the inter-layer offset within a predetermined time period longer than the time interval. Within the predetermined time period, the short-term syntax element is set to an appropriate subset outside the set of possible settings. This appropriate subset is selected such that within the predetermined time period, the actual inter-layer offset between the traversal of the spatial segments of the first-layer image and the traversal of the spatial segments of the second-layer image, is less than or equal to the inter-layer offset signaled by the long-term syntax element. The size and location of the spatial segments of the first-layer image and the second-layer image, as well as the spatial sampling resolution of the first-layer image and the second-layer image, can support the decoding of the multi-layer video data stream by sequentially traversing the spatial segments in a time-overlapping manner.
[0494] The long-term syntax element is set to a value in a second set of possible values that do not intersect with the first set of possible values (min_spatial_segment_delay = 0), and the short-term syntax element is set to any one of the possible settings within the predetermined time period. The set of possible settings includes at least one setting and at least another setting. Based on the at least one setting, the actual inter-layer offset, which is less than or equal to the inter-layer offset signaled by the long-term syntax element, is used between the traversal of the spatial segment of the first layer image and the traversal of the spatial segment of the second layer image. The sizes and values of the spatial segments of the first layer image and the spatial segments of the second layer image are determined by the long-term syntax element. The spatial sampling resolutions of the first-layer and second-layer images respectively disable decoding of the multi-layer video stream by sequentially traversing spatial segments in a time-overlapping manner, and according to the at least one other setting, using the actual inter-layer offset between the traversal of the spatial segments of the first-layer image and the traversal of the spatial segments of the second-layer image, which is less than or equal to the inter-layer offset signaled by the long-term syntax element, the size and positioning of the spatial segments of the first-layer and second-layer images, as well as the spatial sampling resolutions of the first-layer and second-layer images, respectively, can be used to decode the multi-layer video stream by sequentially traversing spatial segments in a time-overlapping manner.
[0495] 42. The video encoder according to claim 41, wherein the video encoder is configured to perform encoding such that the spatial segment is a substream formed by partitioning the images of the layer, and
[0496] The substream is composed of rows of blocks into which the images of the layer are regularly subdivided. In this way, by utilizing the intra-frame image substream delay between traversals of the same image in the middle consecutive substreams and the inter-layer offset between the traversal of the substream of the first layer of images and the traversal of the substream of the second layer of images, wavefront parallel processing is used to decode the multi-layer video data stream into the substream in parallel by traversing the substreams sequentially in a time-overlapping manner.
[0497] 43. The video encoder of claim 42, wherein the video encoder is configured to initialize an entropy context probability by supporting intra-frame picture space prediction across the boundaries of the substreams, the entropy context probability being used to entropy encode the substreams individually, or to encode the substreams using the intra-frame picture space prediction and entropy encoding with the adapted entropy context probability based on the entropy context probability of the preceding substream in an intermediate self-adaptive state according to the order between the substreams.
[0498] 44. The video encoder of claim 41, wherein the video encoder is configured to perform encoding such that the spatial segment is a piece of the images of the layer partitioned, thereby allowing the multi-layer video data stream to be decoded into the piece by traversing the piece in slice order within each of the images of the first layer and the second layer, and using the inter-layer offset of the slice traversal of the images of the first layer relative to the slice traversal of the images of the second layer to decode the immediately following consecutive pieces of the images of the first layer and the immediately following consecutive pieces of the images of the second layer in parallel.
[0499] 45. The video encoder according to claim 44, wherein the video encoder is configured to encode the images of the first layer and the images of the second layer using intra-frame picture spatial prediction.
[0500] Furthermore, the intra-frame image spatial prediction used for each slice is interrupted at the boundary line of the corresponding slice.
[0501] 46. The video encoder according to any one of claims 41 to 45, wherein the video encoder is configured such that the value of the long-term syntax element is limited to the measurement of the interlayer offset in the unit of the spatial segment of the picture of the first layer.
[0502] 47. The video decoder according to any one of items 41 to 46, wherein the video encoder is configured to set the value of the long-term syntax element to signal the number of spatial segments of the first layer of images, by means of which the decoding of the first spatial segment of the second layer of images is delayed relative to the decoding and traversal of the first layer of images.
[0503] 48. A network entity for processing multi-layer video data streams into scenarios, the scenarios being encoded in layers such that, in each layer, the scenarios are encoded at different operational points in a scalable space spanned by a scalable dimension, wherein the multi-layer video data stream comprises first NAL units and second NAL units, each of the first NAL units being associated with one of the layers, and the second NAL units being distributed across the first NAL units.
[0504] Within the unit, and presenting overall information about the multi-layer video data stream, the network entity is configured as follows:
[0505] Detect the type indicator field (609; e.g., dedicated_scalability_ids_flag) in the second NAL unit;
[0506] If the type indicator field has a first state (e.g., dedicated_scalability_ids_flag = 0), then the mapping information (e.g., layer_id_in_nuh[i], dimension_id[i][j]) that maps the possible values of the layer indicator field (e.g., layer_id) in the header of the first NAL unit to the operation point is read from the second NAL unit, and the first NAL unit is associated with the operation point in the first NAL unit through the layer indicator field and the mapping information.
[0507] If the type indicator field has a second state (dedicated_scalability_ids_flag = 1), the first NAL unit is associated with the operation point by dividing the layer indicator field in the first NAL unit into more than one part and using the value of that part as the coordinate of the vector within the scalable space.
[0508] 49. The network entity according to claim 48, wherein the network entity is configured as follows:
[0509] If the type indicator field has the second state (dedicated_scalability_ids_flag = 1), then according to the syntax element (dimension_id_len_minus1) in the second NAL unit, the layer indicator field in the first NAL unit is divided into more than one part, and the value of that part is used as the coordinate of the vector within the scalable space to locate the operation point of the first NAL unit. Furthermore, according to another syntax element (scalability_mask) in the second NAL unit, the scalable dimension is semantically determined so that the first NAL unit...
[0510] The unit is associated with the operation point.
[0511] 50. The network entity according to item 48 or 49, wherein the network entity is configured as follows:
[0512] If the type indicator field has the first state (dedicated_scalability_ids_flag = 0), then the number p of the scalable dimensions and the semantic meaning are determined from another syntax element (scalability_mask) in the second NAL unit and by reading a list of p-dimensional vectors from the second NAL unit (708).
[0513] Associate the possible values of the layer indicator field with the operation point.
[0514] 51. The network entity according to item 50, wherein the network entity is configured to skip reading the list from the second NAL unit if the type indicator field has the second state.
[0515] 52. The network entity according to any one of items 49 or 51, wherein the network entity is configured to read the additional syntax element from the second NAL unit without considering the type indicator field having the first state or the second state;
[0516] And ensures that the size of the layer indicator field is the same regardless of whether the type indicator field has the first state or the second state.
[0517] 53. The network entity according to any one of items 48 or 52, wherein the network entity includes a video decoder.
[0518] 54. A video encoder for encoding a scene into a multi-layer video data stream in layers, such that in each layer, the scene is encoded at different operational points in a scalable space spanned by a scalable dimension, wherein the multi-layer video data stream comprises first NAL units and second NAL units, each of the first NAL units being associated with one of the layers, and the second NAL units being distributed within the first NAL units and presenting overall information about the multi-layer video data stream, the video encoder being configured to:
[0519] Insert the type indicator field into the second NAL unit and in the following settings
[0520] Switch between settings:
[0521] The type indicator field is set to have a first state, and the mapping information of the possible values of the layer indicator field in the first NAL unit header to the operation point is inserted into the second NAL unit. The layer indicator field in the first NAL unit is also set such that the mapping information enables the operation point to...
[0522] The operation point of a NAL cell is associated with the corresponding layer indicator field;
[0523] The type indicator field is set to have a second state (dedicated_scalability_ids_flag = 1), and the first NAL is set by dividing the layer indicator field in the first NAL unit into one or more parts and setting the value of the one or more parts to correspond to the coordinates of the vector in the scalable space, thereby pointing to the operation point associated with the corresponding first NAL unit.
[0524] The layer indicator field in the unit.
[0525] 55. The video encoder according to claim 54, wherein the video encoder is configured to:
[0526] When the type indicator field is set such that the type indicator field has the second state, a syntax element is set and inserted into the second NAL unit, the syntax element defining the division of the type indicator field in the first NAL unit into the one or more parts, and another syntax element is set and inserted into the second NAL unit, the other syntax element semantically defining the scalable dimension.
[0527] 56. A method for converting a multi-layer video data stream into a scene encoded in layers, such that, in each layer, the scene is encoded at different operational points in a scalable space spanned by a scalable dimension, wherein the multi-layer video data stream comprises first NAL units and second NAL units, each of the first NAL units being associated with one of the layers, and
[0528] The second NAL unit is distributed within the first NAL unit and presents overall information about the multi-layer video data stream, wherein the second NAL unit is configured according to the following conditions.
[0529] The cell contains a type indicator field (696; for example, dedicated_scalability_ids_flag);
[0530] If the type indicator field has a first state (e.g., dedicated_scalability_ids_flag = 0), then the mapping information in the second NAL unit maps the possible values of the layer indicator field (e.g., layer_id) in the header of the first NAL unit to the operation point.
[0531] If the type indicator field has a second state (dedicated_scalability_ids_flag = 1), then the layer indicator field in the first NAL unit is divided into more than one part, and the operation point of the first NAL unit is limited by the value of the part to the coordinates of the vector within the scalable space.
[0532] 57. A transport layer decoder for transforming a multi-layer video data stream into a scene encoded in layers for decoding by a multi-standard multi-layer decoder, wherein the multi-layer video data stream is composed of NAL units, each of the NAL units being associated with one of the layers, wherein the layers are associated with different codecs, such that for each layer the NAL units associated with the corresponding layer are encoded using the codec associated with the corresponding layer, the transport layer decoder being configured to:
[0533] For each NAL unit, the same codec associated with the NAL unit is identified, and the NAL unit of the multi-layer video data stream is handed over to the multi-standard multi-layer decoder, which decodes the multi-layer video data stream using inter-layer prediction between layers associated with different codecs.
[0534] 58. The video decoder according to item 57 is further configured as follows:
[0535] Using NAL with a state set to indicate the corresponding layer of the codec.
[0536] The NAL unit header encapsulation of the predefined codec of the unit type indicator is identified as the NAL unit of a layer associated with any codec other than the predefined codec.
[0537] 59. The video decoder according to item 57 or 58 is further configured to:
[0538] Identification is performed based on the channels to which the NAL units arrive.
[0539] 60. The video decoder according to any one of items 57 or 59 is further configured to perform a handover, such that the NAL associated with different codecs is transferred on different channels.
[0540] The unit is transferred to the multi-standard multilayer decoder.
[0541] 61. The video decoder according to item 57 or 60 is further configured as follows:
[0542] Provide each NAL unit with metadata indicating the codec associated with the layer associated with the corresponding NAL unit.
[0543] 62. A video decoder for decoding a multi-layer video data stream into a scene, encoding the scene in a layer hierarchy using inter-layer prediction of a co-located portion from a first layer to a second layer, wherein images of the first layer are subdivided into an array of first blocks and images of the second layer are subdivided into an array of second blocks, wherein a raster scan decoding order is defined in the first block and the second block, the video decoder being configured to:
[0544] Using the inter-layer offset measured in units of the first block relative to the spatial traversal of the first block of the first layer image and the second block of the second layer image, the inter-layer offset for parallel decoding of the first layer image and the second layer image is determined by sequentially traversing the first block and the second block in a time-overlapping manner according to the syntax element structure of the multi-layer video data stream (e.g., ctb_delay_enabled_flag, min_spatial_segment_delay).
[0545] 63. The video decoder according to claim 62, wherein the syntax element structure is a long-term syntax element structure, and the video decoder is configured to:
[0546] The determination is executed in advance within a predetermined time period; and
[0547] Within a time interval shorter than the predetermined time period, the size and location of the first block of the first layer image and the second block of the second layer image, as well as the spatial sampling resolution of the first layer image and the second layer image, are periodically determined based on the short-term syntax elements of the multi-layer video data stream.
[0548] 64. The video decoder according to claim 63, wherein the video decoder utilizes the inter-layer offset of the traversal of spatial segments of the first layer's images relative to the traversal of spatial segments of the second layer's images, to support parallel decoding of the multi-layer video data stream into partitioned images of the layers by sequentially traversing spatial segments in a time-overlapping manner, in spatial segments arranged sequentially along the raster scan decoding order, the video decoder being configured to:
[0549] Detecting the long-term syntax element structure of the multi-layer video data stream enables:
[0550] If the long-term syntax element structure (e.g., ctb_delay_enabled_flag = 0, min_spatial_segment_delay ≠ 0) is set to a value from the first set of possible values, then the inter-layer offset within a predetermined time period is predetermined using the value of the long-term syntax element structure, thereby measuring the inter-layer offset in units of spatial segments of the first layer's images, and periodically determining the size and location of the spatial segments of the first layer's images and the spatial segments of the second layer's images, as well as the spatial sampling resolution of the first layer's images and the second layer's images, based on short-term syntax elements of the multi-layer video data stream, within time intervals shorter than the predetermined time period;
[0551] If the long-term syntax element structure is set to the value of a second set of possible values that do not intersect with the first set of possible values (e.g., min_spatial_segment_delay = 0), then the inter-segment offset within the predetermined time period is periodically determined based on the short-term syntax elements of the multi-layer video data stream within time intervals shorter than the predetermined time period; and
[0552] If the long-term syntax element is set to the value of a third set of possible values that do not intersect with the first and second sets of possible values (e.g., ctb_delay_enabled_flag = 1, min_spatial_segment_delay ≠ 0), then the inter-layer offset is determined in units of the first block, and the size and positioning of the first block of the first layer image and the second block of the second layer image, as well as the spatial sampling resolution of the first layer image and the second layer image, are determined periodically, respectively.
[0553] 65. The video decoder according to item 63 or 64, wherein the video decoder utilizes the intra-frame substream delay between traversals of intermediate consecutive substreams of the same image and the inter-layer offset between the traversal of the substream of the first layer of images and the traversal of the substream of the second layer of images to sequentially traverse the substreams in a time-overlapping manner to support parallel decoding of the multi-layer video data stream in substreams formed by partitioning the images of the layer and consisting of rows of the first block and the second block.
[0554] 66. The video decoder according to any one of items 64 to 65, wherein the long-term syntax element structure includes unit tags (e.g., exemplarily, ctb_delay_enabled_flag).
[0555] and a delay indicator (e.g., exemplarily, min_spatial_segment_delay), where,
[0556] When detecting the structure of the long-term syntax elements, the video decoder is configured to:
[0557] The delay indicator is detected to determine whether to set it to zero; if the delay indicator is set to zero, it is determined that the value of the long-term syntax element structure is set to a value in the second set; and
[0558] If the delay indicator is set to a non-zero value, the value of the long-term syntax element structure is determined using the non-zero value, and if the cell flag is zero, the value of the long-term syntax element structure is determined to be set to the value of the first set, and if the cell flag is one, the value of the long-term syntax element structure is determined to be set to the value in the third set.
[0559] 67. The video decoder according to any one of items 65 to 66, wherein the video decoder is configured to depend on the inter-layer offset when starting to decode the first layer and the second layer in parallel.
[0560] 68. The video decoder according to any one of items 62 to 67, wherein the video decoder is configured to check whether a number of s spatial segments or coded blocks of the first layer have been fully decoded, the number s being uniquely dependent on the syntax element structure, and except where the check shows that at least s spatial segments or coded blocks of the first layer have been fully decoded, to postpone the start of decoding of the second layer during the decoding of the first layer.
[0561] 69. The video decoder according to any one of items 62 to 68, wherein the video decoder is configured to depend on the inter-layer offset when initiating and fully processing parallel decoding of the first layer and the second layer.
[0562] 70. A video decoder according to any one of claims 62 to 69, wherein the video decoder is configured to check whether a number of s spatial segments or coded blocks of the first layer have been fully decoded, the number s being uniquely dependent on the syntax element structure and a number of t-1 decoded spatial segments or coded blocks of the second layer, and, except where the check shows that at least s spatial segments or coded blocks of the first layer have been fully decoded, postpone the start of the t-th segment of the second layer during the decoding of the first layer.
[0563] Decoding of a spatial segment or coded block.
[0564] 71. A method for decoding a multi-layer video data stream (40) into a scene encoded in the layers using inter-layer prediction from the first layer to the second layer, the video decoder supporting parallel decoding of the multi-layer video data stream in spatial segments (80) into which the images (12, 15) of the layers are subdivided, wherein the method comprises:
[0565] Detect the long-term syntax element structure (606; e.g., tile_boundaries_aligned_flag) of the multi-layer video data stream to:
[0566] Deciphering the long-term syntax element structure, which assumes that values outside the first set of possible values (e.g., tile_boundaries_aligned_flag = 1) are used as a guarantee for subdividing the second-layer image (15) within a predetermined time period (608) such that the boundaries between spatial segments of the second-layer image overlap with each boundary of a spatial segment of the first-layer image (12), and periodically determining the subdivision of the first-layer and second-layer images into spatial segments based on short-term syntax elements (602; e.g., column_width_minus1[i] and column_width_minus1[i]) of the multi-layer video data stream within a time interval (604) shorter than the predetermined time period; and
[0567] If the long-term syntax element structure assumes a value outside the second possible value set (e.g., tile_boundaries_aligned_flag = 0), then the short-term syntax elements from the multi-layer video data stream are periodically used to determine the subdivision of the layer's images into spatial segments at time intervals shorter than the predetermined time period, such that, at least for a first possible value of the short-term syntax element, there exists a boundary between the spatial segments of the second layer's images that does not overlap with any of the boundaries of the spatial segments of the first layer, and at least for a second possible value of the short-term syntax element, there exists a boundary between the spatial segments of the second layer's images that overlaps with each boundary of the spatial segments of the first layer.
[0568] 72. A method for encoding a scene into a multi-layer video data stream at the layer hierarchy using inter-layer prediction from a first layer to a second layer, such that the multi-layer video data stream can be decoded in parallel within spatial segments subdivided from the images of the layer, wherein the method comprises:
[0569] Long-term syntax element structure (606) and short-term syntax element (602) are inserted into the multi-layer video data stream, wherein the short-term syntax element defines the subdivision of the first-layer and second-layer images into spatial segments within a time interval; and
[0570] Switch between setting the long-term syntax element structure to the following values:
[0571] Values outside the first possible set, and within a predetermined time period (608) longer than the time interval, the short-term syntax element is set to an appropriate subset outside the possible set, the appropriate subset being selected such that the second layer image is subdivided within the predetermined time period, such that the boundaries between spatial segments of the second layer image overlap with each boundary of a spatial segment of the first layer; or
[0572] Values other than the second set of possible values, and setting the short-term syntax element to any one of the set of possible settings within the predetermined time period, the set of possible settings including at least one setting and at least another setting, according to the at least one setting, there are boundaries between spatial segments of the second layer of the image that do not overlap with any of the boundaries of the spatial segments of the first layer, and according to the at least another setting, there are boundaries between spatial segments of the second layer of the image that overlap with each boundary of the spatial segments of the first layer.
[0573] 73. A method for decoding a spatially scalable bitstream (40) into an image, the image being encoded in different spatial layers, and for at least one of the spatial layers, the image being encoded in a first spatial segment, wherein the method comprises:
[0574] The image (12) of the first spatial layer is upsampled to obtain an upsampled reference image, and the image (15) of the second spatial layer is predicted using the upsampled reference image, wherein,
[0575] The method for decoding responds to a syntax element (616; e.g., independent_tile_upsampling_idc) in the spatially scalable bitstream to interpolate (620) an image of the first spatial layer according to the syntax element.
[0576] This is such that any portion of a partition (622) of the upsampled reference image that depends on the first spatial segment is independent of a portion of the first spatial layer image covered by any of the other portions of the partition; or that any portion of a partition (622) of the upsampled reference image depends on a portion of the first spatial layer image covered by another portion of the partition that is spatially adjacent to the corresponding portion.
[0577] 74. A method for encoding images in different spatial layers into a spatially scalable bitstream, wherein for at least one of the spatial layers, the image is encoded in a first spatial segment, wherein the method comprises:
[0578] The method involves upsampling an image of a first spatial layer to obtain an upsampled reference image and using the upsampled reference image to predict an image of a second spatial layer. The method includes setting a syntax element (606) and inserting the syntax element (606) into the spatially scalable bitstream, and interpolating the image of the first spatial layer based on the syntax element.
[0579] Such that any portion of a partition of the upsampled reference image dependent on the first spatial segment is independent of portions of the first spatial layer image covered by any of the other portions of the partition; or
[0580] This causes any portion of the partition of the upsampled reference image to depend on a portion of the first spatial layer image covered by another partition in the partition that is spatially adjacent to the corresponding portion.
[0581] 75. A method for decoding a multi-layer video data stream into a scene encoded in the layers using inter-layer prediction from a first layer to a second layer, the video decoder supporting the parallel decoding of the multi-layer video data stream into spatial segments subdivided from the images of the layers by sequentially traversing spatial segments in a temporally overlapping manner, utilizing the inter-layer delay of traversing spatial segments of images of the first layer relative to traversing spatial segments of images of the second layer, the method comprising:
[0582] Detecting the long-term syntax element structure (642; e.g., min_spatial_segment_delay) of the multi-layer video data stream such that:
[0583] If the long-term syntax element structure (e.g., min_spatial_segment_delay≠0) is set to a value in the first set of possible values, then the inter-layer offset within a predetermined time period is predetermined using the value of the long-term syntax element, and the size and location of the spatial segments of the first layer image and the second layer image, as well as the spatial sampling resolution of the first layer image and the second layer image, are periodically determined based on the short-term syntax element (602) of the multi-layer video data stream in a time interval shorter than the predetermined time period.
[0584] If the long-term syntax element is set to a value of a second set of possible values that do not intersect with the first set of possible values (e.g., min_spatial_segment_delay = 0), then the inter-segment offset within the predetermined time period is periodically determined based on the short-term syntax element of the multi-layer video data stream within a time interval shorter than the predetermined time period.
[0585] 76. A method for encoding a scene into a multi-layer video data stream using inter-layer prediction from a first layer to a second layer, such that the multi-layer video data stream can be decoded in spatial segments partitioned by the layers of the images by sequentially traversing spatial segments in a temporally overlapping manner, utilizing an inter-layer offset of the traversal of spatial segments of the images of the first layer relative to the traversal of spatial segments of the images of the second layer, the method comprising:
[0586] The long-term syntax element structure (min_spatial_segment_delay) and the short-term syntax element are inserted and set into the multi-layer video data stream. The short-term syntax element of the multi-layer video data stream defines the size and location of the spatial segment of the first layer image and the spatial segment of the second layer image, as well as the spatial sampling resolution of the first layer image and the second layer image, respectively, based on the period within the time interval.
[0587] The method includes switching between the following settings:
[0588] The long-term syntax element structure (min_spatial_segment_delay≠0) is set to a value in a first set of possible values, and this value emits the signal of the inter-layer offset within a predetermined time period larger than the time interval. Within the predetermined time period, the short-term syntax element is set to an appropriate subset outside the possible set. This appropriate subset is selected such that within the predetermined time period, the actual inter-layer offset of the inter-layer offset signaled by the long-term syntax element is less than or equal to the traversal of the spatial segments of the first-layer image relative to the traversal of the spatial segments of the second-layer image. The size and location of the spatial segments of the first-layer image and the spatial segments of the second-layer image, as well as the spatial sampling resolution of the first-layer image and the second-layer image, are used to sequentially traverse the spatial segments through the temporal overlap method to enable decoding of the multi-layer video data stream.
[0589] The long-term syntax element is set to a value in a second set of possible values (min_spatial_segment_delay = 0) that does not intersect with the first set of possible values. During the predetermined time period, the short-term syntax element is set to any one of the possible settings in the set of possible settings, which includes at least one setting and at least another setting. Based on the at least one setting, the actual inter-layer offset, which is less than or equal to the inter-layer offset signaled by the long-term syntax element, is used to determine the size and location of the spatial segments of the first-layer and second-layer images. The spatial sampling resolutions of the first-layer and second-layer images are disabled by sequentially traversing spatial segments in a time-overlapping manner, and according to the at least one other setting, the actual inter-layer offset of the spatial segment traversal of the first-layer image relative to the spatial segment traversal of the second-layer image, which is less than or equal to the inter-layer offset signaled by the long-term syntax element, is used to enable the decoding of the multi-layer video data stream by sequentially traversing spatial segments in a time-overlapping manner, using the spatial segment traversal of the first-layer image and the spatial segment traversal of the second-layer image.
[0590] 77. A method for processing a multi-layer video data stream into a scene encoded in layers, such that the scene is encoded at different operational points in a scalable space spanned by a scalable dimension in each layer, wherein the multi-layer video data stream comprises first NAL units and second NAL units, each of the first NAL units being associated with one of the layers, and the second NAL units being distributed within the first NAL units and presenting overall information about the multi-layer video data stream, the method comprising:
[0591] Detect the type indicator field (696; for example, dedicated_scalability_ids_flag) in the second NAL unit;
[0592] If the type indicator field has a first state (e.g., dedicated_scalability_ids_flag = 0), then the mapping information (e.g., layer_id_in_nuh[i], dimension_id[i][j]) that maps the possible values of the layer indicator field (e.g., layer_id) in the header of the first NAL unit to the operation point is read from the second NAL unit, and the first NAL unit is associated with the operation point within the first NAL unit through the layer indicator field and the mapping information.
[0593] If the type indicator field has a second state (dedicated_scalability_ids_flag = 1), the first NAL unit is associated with the operation point by dividing the layer indicator field in the first NAL unit into more than one part, and the operation point of the first NAL unit is located by using the value of that part as the coordinate of the vector in the scalable space.
[0594] 78. A method for encoding a scene into a multi-layer video data stream in layers, such that, in each layer, the scene is encoded at different operational points in a scalable space spanned by a scalable dimension, wherein the multi-layer video data stream is composed of a first NAL unit and a second NAL unit.
[0595] The unit configuration is such that each of the first NAL units is associated with one of the layers, and the second NAL units are distributed within the first NAL units and present a relationship with respect to...
[0596] The method includes the following: (The overall information of the multi-layer video data stream)
[0597] Insert the type indicator field into the second NAL unit and set it as follows:
[0598] Switch between settings:
[0599] The type indicator field is set to have a first state, and the mapping information of the possible values of the layer indicator field in the first NAL unit header to the operation point is inserted into the second NAL unit, and the layer indicator field in the first NAL unit is set so that the operation point of the first NAL unit is associated with the indicator field of the corresponding layer through the mapping information.
[0600] The type indicator field is set to have a second state (dedicated_scalability_ids_flag = 1), and the layer indicator field in the first NAL unit is set by dividing the layer indicator field in the first NAL unit into more than one part, and the more than one part is set such that the value of the part corresponds to the coordinate of the vector in the scalable space, thereby pointing to the operation point associated with the corresponding first NAL unit.
[0601] 79. A method for converting a multi-layer video data stream into a scene encoded in layers for decoding by a multi-standard multi-layer decoder, wherein the multi-layer video data stream is generated by NAL.
[0602] The unit configuration, wherein each of the NAL units is associated with one of the layers, wherein,
[0603] The layers are associated with different codecs such that, for each layer, the NAL units associated with that layer are encoded using the codec associated with that layer.
[0604] include:
[0605] For each NAL unit, identify the same codec associated with the NAL unit;
[0606] and
[0607] The NAL units of the multi-layer video data stream are handed over to the multi-standard multi-layer decoder, which uses inter-layer prediction between layers associated with different codecs to decode the multi-layer video data stream.
[0608] 80. A method for decoding a multi-layer video data stream into a scene, encoding the scene in layers using inter-layer prediction of a co-located portion from a first layer to a second layer, wherein images of the first layer are subdivided into an array of first blocks and images of the second layer are subdivided into an array of second blocks, wherein the first block and the second block respectively define a raster scan decoding order, the method comprising:
[0609] Using the inter-layer offset measured within the cell of the first block, representing the spatial traversal of the first block of the first layer image relative to the spatial traversal of the second block of the second layer image, the inter-layer offset for parallel decoding of the first layer image and the second layer image is determined by sequentially traversing the first block and the second block in a time-overlapping manner, based on the syntax element structure of the multi-layer video data stream (e.g., ctb_delay_enabled_flag, min_spatial_segment_delay).
[0610] 81. A computer program having program code, which, when run on a computer, performs the method according to any one of claims 71 to 80.
Claims
1. A decoder for decoding a spatially scalable bitstream (40) into an image, said image being encoded in different spatial layers at different spatial resolutions, wherein, The decoder is configured as follows: Use upsampling to predict the image (15) of the second spatial layer from the image (12) of the first spatial layer. The decoder responds to a syntax element (616) in the spatially scalable bitstream to perform upsampling by interpolating (620) an image of the first spatial layer based on the syntax element. If the syntax element has a first value, then a fragment of the filter kernel used in the image interpolated from the first spatial layer is filled using a fallback rule. This fragment of the filter kernel extends beyond the boundary of a partition in the image of the first spatial layer, and the filter kernel is located within the first spatial layer. According to the fallback rule, the fragment is filled independently of the outside of the partition. This ensures that if the syntax element has a second value, the segment of the filter kernel extending beyond the boundary of the partition is filled with portions of the first spatial layer image outside the partition where the segment of the filter kernel overlaps. Detecting the long syntax element structure (606) in the bitstream, such that: The interpretation assumes that the long-term syntax element structure contains values outside the first possible set of values, as follows: within a predetermined time period (608), the second-layer picture (15) is subdivided such that the boundaries between the spatial segments of the second-layer picture cover the boundaries of each spatial segment of the first-layer picture (12); and within a time interval (604) shorter than the predetermined time period, the short-term syntax element (602) based on the bitstream period periodically determines the guarantee that the first-layer and second-layer pictures are subdivided into spatial segments, and If the long-term syntax element structure assumes values outside the second possible set of the long-term syntax element structure, then within a time interval shorter than the predetermined time period, the spatial segment subdivisions of the layer's picture are periodically determined from the short-term syntax elements of the bitstream, such that, at least for the first possible value of the short-term syntax element, there are boundaries between the spatial segments of the second layer's picture that do not cover any boundaries of the first layer's spatial segments, and, at least for the second possible value of the short-term syntax element, the boundaries between the spatial segments of the second layer's picture cover each boundary of the first layer's spatial segments.
2. The decoder according to claim 1, wherein, The decoder is configured as follows: Decode the image of the first spatial layer using intra-frame image spatial prediction, and The intra-frame image spatial prediction is interrupted, so that the boundary lines of the first spatial segments into which the image does not cross the first spatial layer are divided. The decoder responds to a syntax element (616) in the spatially scalable bitstream to perform upsampling by interpolating (620) an image of the first spatial layer based on the syntax element. If the syntax element has the first value, then a segment of the filter kernel used in the image interpolated from the first spatial layer is filled using a backoff rule. This segment of the filter kernel extends beyond the boundary of a predetermined first spatial segment of the image in the first spatial layer. The filter kernel is located within the first spatial layer and, according to the backoff rule, is filled independently of the outside of the predetermined first spatial segment. This ensures that if the syntax element has the second value, a portion of the adjacent first spatial segment covered by the segment of the filter kernel fills the segment of the filter kernel that extends beyond the boundary line of the predetermined first spatial segment.
3. The decoder according to claim 2, wherein, The decoder is configured as follows: Decode the image of the first spatial layer using intra-frame image spatial prediction, and The intra-frame image spatial prediction is interrupted, such that the boundary lines of the second spatial segments into which the image does not cross the second spatial layer are segmented, wherein the first spatial segment and the second spatial segment are spatially aligned.
4. The decoder according to claim 1, wherein, The spatially scalable bitstream has images encoded into a first spatial layer of the spatially scalable bitstream in a manner that divides them into independently coded first spatial segments, wherein the decoder responds to the syntax element (616) in the spatially scalable bitstream to perform the upsampling by interpolating (620) the images of the first spatial layer according to the syntax element. If the syntax element has the first value, then a fragment of the filter kernel used in the image interpolated from the first spatial layer is filled using a fallback rule. This fragment of the filter kernel extends beyond the boundary of a predetermined first spatial segment of the image in the first spatial layer. The filter kernel is located within the first spatial layer. According to the fallback rule, the fragment is filled independently of the outside of the predetermined first spatial segment. This ensures that if the syntax element has the second value, the portion of the adjacent first spatial segment covered by the segment of the filter kernel fills the segment of the filter kernel that extends beyond the boundary line of the predetermined first spatial segment.
5. The decoder according to claim 4, wherein, The spatially scalable bitstream has an image encoded in the second spatial layer of the spatially scalable bitstream in a manner that divides it into independently encoded second spatial segments, wherein the first spatial segment and the second spatial segment are spatially aligned.
6. The decoder according to claim 2 or 4, wherein, The decoder responds to the syntax element (616) in the spatially scalable bitstream to perform the upsampling based on the syntax element by interpolating (620) the image of the first spatial layer. If the syntax element has the first value, then a fragment of the filter kernel used in the image interpolated from the first spatial layer is filled using a fallback rule. This fragment of the filter kernel extends beyond the boundary of a predetermined first spatial segment of the image in the first spatial layer. The filter kernel is located within the first spatial layer. According to the fallback rule, the fragment is filled independently of the outside of the predetermined first spatial segment. If the syntax element has a third value, then a segment of the filter kernel used in the interpolated first spatial layer image is filled using a fallback rule. This segment of the filter kernel extends beyond the boundary of a predetermined first spatial segment of the first spatial layer image, and the filter kernel is located within the first spatial layer. According to the fallback rule, the segment is filled independently of the outside of the predetermined first spatial segment. Furthermore, another segment of the filter kernel is filled using the fallback rule. This other segment of the filter kernel extends beyond the boundary of a predetermined second spatial segment of the second spatial layer image, overlapping the boundary of the first spatial layer image. This filter kernel is located within the first spatial layer, and according to the fallback rule, the segment is filled independently of the outside of the predetermined second spatial segment. If the syntax element has the second value, then the segment of the filter kernel extending beyond the boundary line of the predetermined first spatial segment is filled with the portion of the adjacent first spatial segment whose segments overlap, and another segment of the filter kernel extending beyond the boundary line of the predetermined second spatial segment on the image of the first spatial layer is filled with the portion of the overlapping adjacent second spatial segment whose segments overlap.
7. The decoder according to claim 1, wherein, The decoder is configured to use the backoff rule while still filling fragments of the filter kernel that extend beyond the outer boundary of the image in the first spatial layer.
8. The decoder according to claim 1, wherein, The decoder is a video decoder and is configured to respond to the syntax element (616) in the spatially scalable bitstream on a per-picture or per-picture sequence basis.
9. The decoder according to claim 1, wherein, The decoder is configured to decode the first spatial layer and the second spatial layer in parallel using inter-layer offsets that depend on the syntax element (616).
10. An encoder for encoding images into a spatially scalable bitstream at different spatial layers with different spatial resolutions, wherein, The encoder is configured as follows: Use upsampling to predict images of the second spatial layer from images of the first spatial layer (15). The encoder responds to the insertion of a syntax element (616) into the spatially scalable bitstream and performs upsampling based on the syntax element by interpolating (620) an image of the first spatial layer. If the syntax element has a first value, then a fragment of the filter kernel used in the image interpolated from the first spatial layer is filled using a fallback rule. This fragment of the filter kernel extends beyond the boundary of a partition in the image of the first spatial layer, and the filter kernel is located within the first spatial layer. According to the fallback rule, the fragment is filled independently of the outside of the partition. This ensures that if the syntax element has a second value, the fragment of the filter kernel extending beyond the boundary of the partition is filled with a portion of the first spatial layer image outside the partition where the fragment of the filter kernel overlaps. Long-term syntax element structure (606) and short-term syntax element (602) are inserted into the bitstream, wherein the short-term syntax element defines the subdivision of the first-layer image and the second-layer image into spatial segments within a time interval, and Switch between the following settings: The long-term syntax element structure is set to a value outside the first set of possible values. Within a predetermined time period (608) longer than the time interval, the short-term syntax element is set to an appropriate subset of the set of possible settings. This appropriate subset is selected such that the second-layer image is subdivided within the predetermined time period, such that the boundaries between spatial segments of the second-layer image cover each boundary of the spatial segment of the first-layer image; or The long-term syntax element structure is set to a value outside the second set of possible values, and the short-term syntax element is set to any one of the set of possible settings within a predetermined time period. The set of possible settings includes at least one setting and at least another setting. According to the at least one setting, there are boundaries between spatial segments of the second layer of the image that do not cover any of the boundaries of the spatial segments of the first layer, and according to the at least another setting, the boundaries between spatial segments of the second layer of the image cover each boundary of the spatial segment of the first layer.
11. The encoder of claim 10, configured as follows: The images of the first spatial layer are encoded using intra-frame image spatial prediction, and The intra-frame image spatial prediction is interrupted, so that the boundary lines of the first spatial segments into which the image does not cross the first spatial layer are divided. in, The encoder is configured to perform the upsampling based on the syntax element by interpolating (620) the image of the first spatial layer. If the syntax element has the first value, then a segment of the filter kernel used in the image interpolated from the first spatial layer is filled using a backoff rule. This segment of the filter kernel extends beyond the boundary of a predetermined first spatial segment of the image in the first spatial layer. The filter kernel is located within the first spatial layer and, according to the backoff rule, is filled independently of the outside of the predetermined first spatial segment. This ensures that if the syntax element has the second value, the portion of the adjacent first spatial segment covered by the segment of the filter kernel fills the segment of the filter kernel that extends beyond the boundary line of the predetermined first spatial segment.
12. The encoder of claim 10, configured to encode an image of the first spatial layer into the spatially scalable bitstream by segmenting it into independently coded first spatial segments, wherein, The encoder is configured to perform the upsampling based on the syntax element by interpolating (620) the image of the first spatial layer. If the syntax element has the first value, then a fragment of the filter kernel used in the image interpolated from the first spatial layer is filled using a fallback rule. This fragment of the filter kernel extends beyond the boundary of a predetermined first spatial segment of the image in the first spatial layer. The filter kernel is located within the first spatial layer. According to the fallback rule, the fragment is filled independently of the outside of the predetermined first spatial segment. This ensures that if the syntax element has the second value, the portion of the adjacent first spatial segment covered by the segment of the filter kernel fills the segment of the filter kernel that extends beyond the boundary line of the predetermined first spatial segment.
13. The encoder according to claim 10, wherein, The encoder is configured to use the backoff rule while still filling fragments of the filter kernel that extend beyond the outer boundary of the image in the first spatial layer.
14. A method for decoding a spatially scalable bitstream (40) into an image, said image being encoded in different spatial layers at different spatial resolutions, wherein, The method includes: The image (12) of the first spatial layer is upsampled to obtain an upsampled reference image, and the upsampled reference image is used to predict the image (15) of the second spatial layer, which is encoded in a way that spatially segments into independently coded tiles. The method for decoding is in response to a syntax element (616) in the spatially scalable bitstream to interpolate (620) the image of the first spatial layer according to the syntax element. This causes a segment of the filter kernel used in the image interpolated from the first spatial layer to be filled using a backoff rule. The segment of the filter kernel extends from the tile containing the filter kernel to adjacent tiles. According to the backoff rule, the segment is filled independently of the adjacent tiles. This causes the filter kernel to cross the boundary line of the tile, thereby making the upsampled reference image at the location of the filter kernel dependent on the tiles of the first spatial layer that overlap the location of the filter kernel and the adjacent tiles of the first spatial layer; Detecting the long syntax element structure (606) in the bitstream, such that: The interpretation assumes that the long-term syntax element structure contains values outside the first possible set of values, as follows: within a predetermined time period (608), the second-layer picture (15) is subdivided such that the boundaries between the spatial segments of the second-layer picture cover the boundaries of each spatial segment of the first-layer picture (12); and within a time interval (604) shorter than the predetermined time period, the short-term syntax element (602) based on the bitstream period periodically determines the guarantee that the first-layer and second-layer pictures are subdivided into spatial segments, and If the long-term syntax element structure assumes values outside the second possible set of the long-term syntax element structure, then within a time interval shorter than the predetermined time period, the spatial segment subdivisions of the layer's picture are periodically determined from the short-term syntax elements of the bitstream, such that, at least for the first possible value of the short-term syntax element, there are boundaries between the spatial segments of the second layer's picture that do not cover any boundaries of the first layer's spatial segments, and, at least for the second possible value of the short-term syntax element, the boundaries between the spatial segments of the second layer's picture cover each boundary of the first layer's spatial segments.
15. A method for encoding an image into a spatially scalable bitstream in different spatial layers with different spatial resolutions, wherein, The method includes: The image (12) of the first spatial layer is upsampled to obtain an upsampled reference image, and the upsampled reference image is used to predict the image (15) of the second spatial layer, which is encoded in a way that spatially segments into independently coded tiles. The method includes setting a syntax element (616) and inserting the syntax element into the spatially scalable bitstream, and interpolating (620) the image of the first spatial layer according to the syntax element. This causes a segment of the filter kernel used in the image interpolated from the first spatial layer to be filled using a backoff rule. The segment of the filter kernel extends from the tile containing the filter kernel to adjacent tiles. According to the backoff rule, the segment is filled independently of the adjacent tiles. This causes the filter kernel to cross the boundary line of the tile, thereby making the upsampled reference image at the location of the filter kernel dependent on the tiles of the first spatial layer that overlap the location of the filter kernel and the adjacent tiles of the first spatial layer; Long-term syntax element structure (606) and short-term syntax element (602) are inserted into the bitstream, wherein the short-term syntax element defines the subdivision of the first-layer image and the second-layer image into spatial segments within a time interval, and Switch between the following settings: The long-term syntax element structure is set to a value outside the first set of possible values. Within a predetermined time period (608) longer than the time interval, the short-term syntax element is set to an appropriate subset of the set of possible settings. This appropriate subset is selected such that the second-layer image is subdivided within the predetermined time period, such that the boundaries between spatial segments of the second-layer image cover each boundary of the spatial segment of the first-layer image; or The long-term syntax element structure is set to a value outside the second set of possible values, and the short-term syntax element is set to any one of the set of possible settings within a predetermined time period. The set of possible settings includes at least one setting and at least another setting. According to the at least one setting, there are boundaries between spatial segments of the second layer of the image that do not cover any of the boundaries of the spatial segments of the first layer, and according to the at least another setting, the boundaries between spatial segments of the second layer of the image cover each boundary of the spatial segment of the first layer.
16. A method for storing images, the method comprising: A spatially scalable bitstream is stored on a digital storage medium, the spatially scalable bitstream having the image encoded into the spatially scalable bitstream by performing the method of claim 15.
17. A digital storage medium storing a computer program and a spatially scalable bitstream, the spatially scalable bitstream having images encoded into the spatially scalable bitstream, wherein, When the computer program is executed by the processor, it implements the method of claim 15 to encode the image into the spatially scalable bitstream.
18. A decoder for decoding a spatially scalable bitstream (40) into an image, said image being encoded in different spatial layers at different spatial resolutions, wherein, The decoder is configured as follows: Upsample the image (12) of the first spatial layer to obtain an upsampled reference image, and The upsampled reference image is used to predict the image of the second spatial layer (15) which is encoded in a way that spatially segments it into independently coded tiles. The decoder responds to a syntax element (616) in the spatially scalable bitstream to interpolate (620) an image of the first spatial layer based on the syntax element. This causes a segment of the filter kernel used in the image interpolated from the first spatial layer to be filled using a backoff rule. The segment of the filter kernel extends from the tile containing the filter kernel to adjacent tiles. According to the backoff rule, the segment is filled independently of the adjacent tiles. This causes the filter kernel to cross the boundary line of the tile, thereby making the upsampled reference image at the location of the filter kernel dependent on the tiles of the first spatial layer that overlap the location of the filter kernel and the adjacent tiles of the first spatial layer; Detecting the long syntax element structure (606) in the bitstream, such that: The interpretation assumes that the long-term syntax element structure contains values outside the first possible set of values, as follows: within a predetermined time period (608), the second-layer picture (15) is subdivided such that the boundaries between the spatial segments of the second-layer picture cover the boundaries of each spatial segment of the first-layer picture (12); and within a time interval (604) shorter than the predetermined time period, the short-term syntax element (602) based on the bitstream period periodically determines the guarantee that the first-layer and second-layer pictures are subdivided into spatial segments, and If the long-term syntax element structure assumes values outside the second possible set of the long-term syntax element structure, then within a time interval shorter than the predetermined time period, the spatial segment subdivisions of the layer's picture are periodically determined from the short-term syntax elements of the bitstream, such that, at least for the first possible value of the short-term syntax element, there are boundaries between the spatial segments of the second layer's picture that do not cover any boundaries of the first layer's spatial segments, and, at least for the second possible value of the short-term syntax element, the boundaries between the spatial segments of the second layer's picture cover each boundary of the first layer's spatial segments.
19. The decoder according to claim 18, wherein, The decoder is configured to decode the different spatial layers in parallel.
20. The decoder according to claim 18, wherein, The decoder is configured to decode the image of the second spatial layer using intra-frame picture spatial prediction, interrupting the intra-frame picture spatial prediction for each tile so as not to cross the boundary line of the corresponding tile.
21. The decoder according to claim 18, wherein, The spatially scalable bitstream has an image of a first spatial layer encoded as other independently coded tiles, wherein the decoder responds to the syntax element (616) in the spatially scalable bitstream to interpolate the image of the first spatial layer according to the syntax element. This causes a segment of the filter kernel used in the image interpolated from the first spatial layer to be filled using a backoff rule. The segment of the filter kernel extends from the tile containing the filter kernel to adjacent tiles. According to the backoff rule, the segment is filled independently of the adjacent tiles. This causes a segment of the filter kernel to be filled using a backoff rule, the segment of the filter kernel extending from a portion of a partition including the location of the filter kernel to an adjacent portion, the segment being filled independently of adjacent tiles according to the backoff rule, wherein the portion and the adjacent portion are defined by a partition, the partition being defined by a logical AND of the spatial overlap of the boundaries of the first tile and the second tile, or This causes the filter kernel to span the boundary line of the tile, thereby making the upsampled reference image at the location of the filter kernel dependent on the tiles of the first spatial layer that overlap the location of the filter kernel and the adjacent tiles of the first spatial layer.
22. The decoder according to claim 18, wherein, The decoder is configured to use the backoff rule while still filling fragments of the filter kernel that extend beyond the outer boundary of the image in the first spatial layer.
23. The decoder according to claim 18, wherein, The decoder is a video decoder and is configured to respond to the syntax element (616) in the spatially scalable bitstream on a per-picture or per-picture sequence basis.
24. The decoder according to claim 18, wherein, The spatially scalable bitstream has an image of the first spatial layer encoded in a first spatial segment, wherein the spatially scalable bitstream has an image of the second spatial layer encoded in a second spatial segment, wherein the boundary of the partition corresponds to a logical AND of the spatial overlap of the boundary of the first spatial segment and the boundary of the second spatial segment or the boundary of the second spatial segment, wherein the decoder, in response to the syntax element (616) in the spatially scalable bitstream, fills a segment of the filter kernel used in the interpolation of the image of the first spatial layer using a backoff rule, the segment of the filter kernel extending from one partition to an adjacent partition of the partition, and, according to the backoff rule, fills the segment independently of the corresponding portion of the image of the first spatial layer to which the filter kernel extends, or using the corresponding portion of the image of the first spatial layer to which the filter kernel extends.
25. The decoder according to claim 18, wherein, The decoder is configured to decode the first spatial layer and the second spatial layer in parallel using inter-layer offsets that depend on the syntax element (616).
26. An encoder for encoding images into spatially scalable bitstreams in different spatial layers with different spatial resolutions, wherein, The encoder is configured as follows: The image (12) of the first spatial layer is upsampled to obtain an upsampled reference image, and the upsampled reference image is used to predict the image (15) of the second spatial layer, which is encoded in a way that spatially segments into independently coded tiles. The encoder sets syntax elements and inserts these syntax elements into the spatially scalable bitstream, and interpolates (620) the image of the first spatial layer according to the syntax elements (616). This causes a segment of the filter kernel used in the image interpolated from the first spatial layer to be filled using a backoff rule. The segment of the filter kernel extends from the tile containing the filter kernel to adjacent tiles. According to the backoff rule, the segment is filled independently of the adjacent tiles. This causes the filter kernel to cross the boundary line of the tile, thereby making the upsampled reference image at the location of the filter kernel dependent on the tiles of the first spatial layer that overlap the location of the filter kernel and the adjacent tiles of the first spatial layer; Long-term syntax element structure (606) and short-term syntax element (602) are inserted into the bitstream, wherein the short-term syntax element defines the subdivision of the first-layer image and the second-layer image into spatial segments within a time interval, and Switch between the following settings: The long-term syntax element structure is set to a value outside the first set of possible values. Within a predetermined time period (608) longer than the time interval, the short-term syntax element is set to an appropriate subset of the set of possible settings. This appropriate subset is selected such that the second-layer image is subdivided within the predetermined time period, such that the boundaries between spatial segments of the second-layer image cover each boundary of the spatial segment of the first-layer image; or The long-term syntax element structure is set to a value outside the second set of possible values, and the short-term syntax element is set to any one of the set of possible settings within a predetermined time period. The set of possible settings includes at least one setting and at least another setting. According to the at least one setting, there are boundaries between spatial segments of the second layer of the image that do not cover any of the boundaries of the spatial segments of the first layer, and according to the at least another setting, the boundaries between spatial segments of the second layer of the image cover each boundary of the spatial segment of the first layer.
27. The encoder according to claim 26, wherein, The encoder is configured to encode the image of the first spatial layer using intra-frame image spatial prediction, and the intra-frame image spatial prediction for each tile is interrupted so as not to cross the boundary line of the corresponding tile.
28. The encoder according to claim 26, wherein, The encoder is configured to encode the image of the first spatial layer into the spatially scalable bitstream by segmenting it into other independently coded tiles. The encoder is configured to set syntax elements and insert these syntax elements into the spatially scalable bitstream, and to interpolate the image of the first spatial layer based on these syntax elements. This causes a segment of the filter kernel used in the image interpolated from the first spatial layer to be filled using a backoff rule. The segment of the filter kernel extends from the tile containing the filter kernel to adjacent tiles. According to the backoff rule, the segment is filled independently of the adjacent tiles. This causes a segment of the filter kernel to be filled using a backoff rule, the segment of the filter kernel extending from a portion of a partition including the location of the filter kernel to an adjacent portion, the segment being filled independently of adjacent tiles according to the backoff rule, wherein the portion and the adjacent portion are defined by a partition, the partition being defined by a logical AND of the spatial overlap of the boundaries of the first tile and the second tile, or This causes the filter kernel to span the boundary line of the tile, thereby making the upsampled reference image at the location of the filter kernel dependent on the tiles of the first spatial layer that overlap the location of the filter kernel and the adjacent tiles of the first spatial layer.
29. The encoder according to claim 26, wherein, The encoder is configured to use the backoff rule while still filling fragments of the filter kernel that extend beyond the outer boundary of the image in the first spatial layer.
30. The encoder according to claim 26, wherein, The encoder is a video encoder and is configured to set up on a per-image or per-image sequence basis and insert the syntax elements into the spatially scalable bitstream.
Citation Information
Patent Citations
High-efficiency scalable coding concept
CN110062240B
Methods and systems for extended spatial scalability with picture-level adaptation
CN101379511A
Discardable lower layer adaptations in scalable video coding
CN101558651A