Output layer set processing for video coding
By introducing the output layer set indication, the efficiency problem of sub-bitstream extraction in multi-layer video bitstream is solved, efficient decoding and resource utilization are achieved, signaling overhead is reduced, and decoder level requirements are met.
Patent Information
- Application Number
- CN202180037217.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-05-22
- Filing Date
- 2021-05-20
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2041-05-20
AI Technical Summary
Existing video coding technologies have difficulty in efficiently extracting sub-bitstreams when processing multi-layer video bitstreams, resulting in wasted decoder resources and excessive signaling overhead, and ineffective use of decoder capabilities.
By introducing the output layer set (OLS) indication in the video bitstream, the decoder is allowed to selectively extract and decode the randomly accessible sub-bitstream, and the inter-layer reference relationship and temporal sub-layer indication are utilized to accurately extract the sub-bitstream, avoiding unnecessary decoding and signaling overhead.
This achieves efficient extraction of sub-bitstreams, reduces decoder resource waste and signaling overhead, improves decoding efficiency, and ensures that the decoder can meet level requirements.
Smart Images

Figure CN115668928B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments of the present invention relate to an apparatus for encoding a video into a video bitstream, an apparatus for decoding a video bitstream, and an apparatus for processing a video bitstream, e.g. extracting a bitstream such as a sub-bitstream from a video bitstream. Other embodiments relate to a method for encoding, a method for decoding, a method for processing a video bitstream, e.g. a method for extracting. Other embodiments relate to a video bitstream. BACKGROUND
[0002] It is envisaged that the emerging VVC codec supports from the start layered coding of temporal, fidelity and spatial scalability, i.e. a coded video bitstream structured into so-called layers and (temporal) sub-layers and coded picture data corresponding to a temporal instance, i.e. so-called access units (AU) can contain pictures within each layer, which can be predicted from each other and some of which are output after decoding. So-called output layer sets (OLS) conceive to indicate reference relationships to a decoder and which layers are to be output when decoding the bitstream. OLSs can also be used to identify corresponding HRD related timing / buffering information in the form of buffering period, picture timing and decoding unit information SEI messages, which are encapsulated in so-called scalable nesting SEI messages to be carried in the bitstream.
[0003] It is desirable to have conceptions for processing output layer sets to allow for extraction of sub-bitstreams from a video bitstream, which provide an improved trade-off between precise definition of extractable sub-bitstreams by the output layer sets (in terms of precisely describing which parts of the video bitstream are to be extracted), efficient utilization of decoder resources (e.g. in terms of avoiding extraction of parts which are not necessary for decoding a selected sub-bitstream, or in terms of providing precise information about decoder settings or requirements for decoding a selected sub-bitstream), and less signaling overhead. SUMMARY
[0004] According to a first aspect of the present invention conceptions for indicating, extracting and / or decoding a randomly accessible sub-bitstream from a multi-layer video bitstream are provided. According to the first aspect, the extracted randomly accessible sub-bitstream selectively comprises bitstream portions of access units of the multi-layer video bitstream which are associated with output layers of the randomly accessible sub-bitstream indicated by an output layer set of the randomly accessible sub-bitstream, or bitstream portions required for decoding the randomly accessible bitstream portions of the output layers.
[0005] The second aspect of the present invention provides a concept of a multi-layer video bitstream having a plurality of layers and a plurality of temporal layers. The multi-layer video bitstream comprises an indication of an output layer set comprising one or more layers of the multi-layer video bitstream and a reference layer indication indicating an inter-layer reference of a layer of the output layer set. The multi-layer video bitstream comprises an indication, e.g. a temporal layer indication or an intra temporal layer indication, which allows to identify, in combination with the way the multi-layer video bitstream is encoded, bitstream portions of the output layer set which belong to the output layer set among the layers of the output layer set. This concept allows to identify the bitstream portions of the OLS by the bitstream portion type of the bitstream portions and / or by the dependency between the layers of the OLS indicated by the reference layer indication. Embodiments of the second aspect thus allow an exact extraction of sub-bitstreams while avoiding unnecessary high signaling overhead.
[0006] The third aspect of the present invention provides a concept which allows a decoder to determine an output layer set to be decoded for decoding a video bitstream based on properties provided to the decoder in the video bitstream. This concept thus enables the decoder to select an OLS without an indication by the decoder which OLS to decode. The ability of the decoder to select an OLS without an indication can for example ensure by an indication in the video bitstream that the bitstream decoded by the decoder fulfills level requirements known to the decoder.
[0007] The fourth aspect of the present invention provides a concept for extracting a sub-bitstream from a multi-layer video data stream such that within the extracted sub-bitstream, access units exclusively comprising one (e.g. the same) type of picture or bitstream portion of a predetermined bitstream portion type or picture type (e.g. a randomly accessible or independently encoded bitstream portion type or picture type) set are indicated by a sequence start indicator, even in case the respective access unit is a non-sequence start access unit in the original multi-layer video bitstream from which the sub-bitstream is extracted. The frequency of sequence start access units in the sub-bitstream can thus be higher than in the multi-layer video data stream and, therefore, a decoder can benefit from having more sequence start access units available, avoiding unnecessary long waiting times until decoding of a video sequence can start.
[0008] The fifth aspect of the application provides a concept allowing extraction of a sub-bitstream from a multi-layer video bitstream such that the sub-bitstream exclusively comprises pictures belonging to one or more temporal sub-layers associated with an output layer set describing the sub-bitstream to be extracted. To this end, a syntax element in the multi-layer video bitstream is used which indicates a predetermined temporal sub-layer of an OLS in a manner distinguishing different states, including a state according to which the predetermined temporal sub-layer is below a maximum temporal sub-layer among the temporal sub-layers within an access unit to which pictures of at least one layer of the layer subset belong. Avoiding forwarding unnecessary sub-layers of the multi-layer video bitstream can reduce the size of the sub-bitstream and can reduce the requirements of a decoder to decode the sub-bitstream.
[0009] According to embodiments, a decoder capability related parameter of a sub-bitstream exclusively comprising pictures belonging to temporal sub-layers of an OLS describing the sub-bitstream is signaled in the sub-bitstream and / or in the multi-layer video data stream. Thus, decoder capabilities can be efficiently utilized since pictures not belonging to the OLS can be omitted when determining the decoder related capability parameter.
[0010] The sixth aspect of the application provides a concept of handling temporal sub-layers in the signaling of video parameters for an output layer set of a multi-layer video bitstream. According to embodiments, an OLS is associated with one of one or more bitstream conformance sets, one of one or more buffer requirement sets, and one of one or more decoder requirement sets signaled in the video bitstream, wherein each of the bitstream conformance set, the buffer requirement set, and the decoder requirement set is valid for one or more temporal sub-layers indicated by a constraint on a maximum temporal sub-layer (e.g., a hierarchically ordered temporal sub-layer). Embodiments provide a concept of a relationship between the bitstream conformance set, the buffer requirement set, and the decoder requirement set associated with the OLS with respect to the maximum temporal sub-layer they are associated with, thereby allowing a decoder to easily determine parameters of the OLS which are associated with the bitstream conformance set, the buffer requirement set, and the decoder requirement set. For example, embodiments can allow a decoder to conclude that the parameters given in the bitstream conformance set, the buffer requirement set, and the decoder requirement set of the OLS are fully valid for the OLS. Other embodiments allow a decoder to conclude that the parameters given in the bitstream conformance set, the buffer requirement set, and the decoder requirement set of the OLS are valid for the OLS to a certain extent.
[0011] According to embodiments, the maximum temporal sub-layer indicated by the decoder requirement set associated with the OLS is less than or equal to the maximum temporal sub-layer indicated by each of the buffer requirement set and the bitstream conformance set associated with the OLS, and parameters within the buffer requirement set and the bitstream conformance set are only valid for the OLS as long as the buffer requirement set and the bitstream conformance set relate to temporal sub-layers equal to or smaller than the maximum temporal sub-layer indicated by the decoder requirement set associated with the OLS. Thus, if the maximum temporal sub-layer indicated by the decoder requirement set associated with the OLS is less than or equal to the maximum temporal sub-layer indicated by each of the buffer requirement set and the bitstream conformance set associated with the OLS, the decoder can infer that parameters in the buffer requirement set and the bitstream conformance set associated with the OLS are only valid for the OLS as long as the buffer requirement set and the bitstream conformance set relate to temporal sub-layers equal to or smaller than the maximum temporal sub-layer indicated by the decoder requirement set associated with the OLS. Thus, embodiments can enable the decoder to determine video parameters for the OLS based on an indication of a constraint on the maximum temporal sub-layer signaled for the respective set of parameters, so that a complex analysis of the OLS and the video parameters can be avoided. Moreover, since this concept allows to associate an OLS with a buffer requirement set and a bitstream conformance set whose constraint on the maximum temporal sub-layer is greater than the constraint for the decoder requirement set associated with the OLS, dedicated buffer requirement sets and dedicated bitstream conformance sets related to the same maximum temporal sub-layer as the decoder requirement set can be omitted, so that the signaling overhead for the video parameter sets is reduced.
[0012] A seventh aspect of the application provides a concept for handling a missing picture of a multi-layer video bitstream, e.g. due to a bitstream error or a transmission loss, which picture is coded using inter-layer prediction. In case a picture which is part of a first layer is missing, this picture can be replaced by another picture of a second layer, the picture of the second layer being used for inter-layer prediction of the picture of the first layer. The concept comprises replacing the picture by the other picture depending on a coincidence of a scaling window defined for the picture with a picture boundary of the picture and a coincidence of a scaling window defined for the other picture with a picture boundary of the other picture. Replacing the picture by the other picture in case the scaling window defined for the picture coincides with a picture boundary of the picture and the scaling window defined for the other picture coincides with a picture boundary of the other picture can for example not result in a change of a display window of a rendered content, e.g. from a detailed view to an overview. BRIEF DESCRIPTION OF DRAWINGS
[0013] Further embodiments and advantageous implementations of the present disclosure are described in more detail below with reference to the accompanying drawings, in which:
[0014] Figure 1Examples of an encoder, extractor, decoder and multi-layer video bitstream according to an embodiment are shown,
[0015] Figure 2 Examples of an output layer set of a random accessible sub-bitstream are shown,
[0016] Figure 3 Examples of an extracted random accessible sub-bitstream with unused pictures are shown,
[0017] Figure 4 Examples of a three-layer bitstream with aligned independently coded pictures on three layers are shown,
[0018] Figure 5 Examples of a three-layer bitstream with unaligned independently coded pictures are shown,
[0019] Figure 6 Examples of a four-layer bitstream are shown, where pictures of access units including independently coded pictures include pictures referring to another temporal sub-layer,
[0020] Figure 7 Examples of a decoder according to an embodiment are shown,
[0021] Figure 8 Examples of a multi-layer video data stream including access units with random accessible pictures and non-random accessible pictures are shown,
[0022] Figure 9 Examples of a sub-bitstream of a multi-layer video data stream of Figure 8 are shown,
[0023] Figure 10 Examples of a multi-layer video bitstream with two layers are shown, each layer of the two layers having a different picture rate,
[0024] Figure 11 Examples of an encoder, extractor, multi-layer video bitstream and sub-bitstream according to an embodiment are shown,
[0025] Figure 12 Examples of a video parameter set and mapping to output layer sets are shown,
[0026] Figure 13 Examples for sharing video parameters between different output layer sets are shown,
[0027] Figure 14 Examples of sharing video parameters between different OLS according to an embodiment are shown,
[0028] Figure 15 Examples of sharing video parameters between different OLS according to another embodiment are shown. DETAILED DESCRIPTION
[0029] Hereinafter, embodiments are discussed in detail; however, it should be understood that the embodiments provide many applicable concepts that can be embodied in various video coding concepts. The specific embodiments discussed merely illustrate specific ways to implement and use the concepts of the present invention, and do not limit the scope of the embodiments. In the following description, a number of details are set forth to provide a more thorough explanation of the embodiments of the present disclosure. However, it will be apparent to those skilled in the art that other embodiments can be practiced without these specific details. In other instances, well-known structures and devices are shown in block diagram form rather than specifically to avoid confusion with the embodiments described herein. In addition, unless otherwise specifically indicated, the features of the different embodiments described herein may be combined with each other.
[0030] In the following description of the embodiments, the same or similar elements or elements having the same function are given the same reference numerals or identified by the same names, and repeated descriptions of the elements given the same reference numerals or identified by the same names are generally omitted. Therefore, the descriptions provided for the elements having the same or similar reference numerals or identified by the same names can be interchanged with each other or can be applied to each other in different embodiments.
[0031] 0. The encoder 10, extractor 30, decoder 50 and video bitstream 12, 14 of Figure 1 0. The encoder 10, extractor 30, decoder 50 and video bitstream 12, 14 of
[0032] The embodiments described in this section provide examples of frameworks in which embodiments of the present invention may be embedded. In the following, descriptions of embodiments of the concepts of the present invention and how these concepts may be built into Figure 1 Although, refer to the following Figure 2 The embodiments described in the following figures can also be used to form Figure 1 It should also be noted that although the encoder, extractor and decoder are Figure 1 The extractor and decoder are described together for illustrative purposes, but they can be implemented separately from each other. It should also be noted that the extractor and decoder can be combined in one device, or one of them can be implemented as part of the other.
[0033] Figure 1 An example of an encoder 10, an extractor 30, a decoder 50, a video bitstream 14 (also referred to as a video data stream or data stream), and a sub-bitstream 12 is shown. The encoder 10 is used to encode a video sequence 20 into the video bitstream 14. The encoder 10 encodes the video sequence 20 into the video bitstream 14 in units of pictures 26, each of which belongs to a time instance, such as a frame of the video sequence. The encoded video data belonging to a common time instance may be referred to as an access unit (AU) 22. In Figure 1, an exemplary application of three access units 221, 222, 223 of a video sequence 20 is shown. Note that a description referring to an access unit 22 may refer to any of the exemplary access units 221, 222, 223. Each of the access units 22 includes or has one or more pictures 26 encoded therein, each of the pictures 26 being associated with one of the multiple layers of the video bitstream 14. Figure 1 In FIG, an example of picture 26 is represented by picture 261 and picture 262. Picture 261 is associated with the first layer 241 of the video bitstream 14, and picture 262 is associated with the second layer 242 of the video bitstream 14. Figure 1 , although each of the access units 22 includes pictures of both the first layer 241 and the second layer 242, the video bitstream 14 may include access units 22 that do not necessarily include pictures 26 of each layer in the layers 24 of the video bitstream 14. Figure 1 In addition to the first and second layers shown, the video bitstream 14 may also include other layers. The encoder 10 is configured to encode each of the pictures 26 into one or more bitstream portions (e.g., NAL units) of the video sequence 14. For example, each of the bitstream portions 16 into which the picture 26 is encoded may encode a portion of the picture 26, such as a slice of the picture 26. The bitstream portion 16 into which the picture 26 is actually encoded may be referred to as a video coding layer (VCL) NAL unit. The bitstream 14 may also include descriptive data indicating information describing the encoded video data, such as non-VCL NAL units. For example, in addition to the bitstream portion that signals the decoded video data, each of the access units 22 may include a bitstream portion that signals descriptive data for the corresponding access unit. The video bitstream 14 may also include descriptive data relating to multiple access units, or portions of one or more access units. For example, the video bitstream 14 may have an output layer set (OLS) indication 18 encoded therein that indicates one or more output layer sets.
[0034] The OLS can be an indication for a sub-bitstream extractable from the video bitstream 14. The OLS can indicate one or more or all of the layers of the video bitstream 14 as output layers of the sub-bitstream described by the respective OLS. Note that the set of layers indicated by the OLS can not necessarily be a proper subset of the layers of the video bitstream 14. In other words, all layers of the video bitstream 14 can be included in the OLS. The OLS can also optionally include a description of the sub-bitstream described by the OLS and / or decoder requirements for decoding the sub-bitstream indicated by the OLS. Note that the sub-bitstream described by the OLS can be defined by further parameters than the layers, e.g., temporal sub-layers or sub-pictures. For example, the pictures 26 of the layer 24 can be associated with one or more of the temporal sub-layers of the layer 24. A temporal sub-layer can include pictures of the temporal instances associated with the respective temporal sub-layer. For example, the pictures of a first temporal sub-layer can be associated with temporal instances forming a sequence of a first frame rate, and the pictures of a second temporal sub-layer can be associated with temporal instances falling in between the temporal instances associated with the pictures of the first temporal sub-layer, such that the combination of the first temporal sub-layer and the second temporal sub-layer can provide a video sequence with a higher frame rate than a single one of the first temporal sub-layer and the second temporal sub-layer. The OLS can optionally indicate the temporal sub-layers for describing which bitstream portions or pictures 26 belong to the sub-bitstream described by the OLS. The temporal sub-layers of a bitstream or coded video sequence can be hierarchically ordered, e.g., by an index. For example, the hierarchical order can mean that decoding a picture of a bitstream including a certain temporal sub-layer requires all temporal sub-layers lower in the hierarchical order.
[0035] Note that the OLS can include one or more output layers and optionally also one or more non-output layers. In other words, the OLS can indicate one or more of the layers included in the OLS as output layers of the OLS and optionally one or more of the layers of the OLS as non-output layers. For example, a layer including reference pictures for the pictures of the output layers of the OLS can be included in the OLS as a non-output layer, as decoding the pictures of the output layers of the OLS can require the pictures of this non-output layer.
[0036] The OLS can also include level information about the bitstream described by the OLS, which indicates one or more bitstream constraints, e.g., a maximum value for one or more of the bit rate, picture size, frame rate, or is associated with one or more bitstream constraints.
[0037] Optionally, the bitstream 14 can also include an extractability indication 19 of the OLS. For example, the extractability indication can be part of the OLS indication. The extractability indication can indicate a (not necessarily true) subset of the bitstream portions 16 that form a decodable sub-bitstream associated with the OLS. That is, the extractability indication can indicate which bitstream portions 16 belong to the OLS.
[0038] The pictures 26 can be coded into the video bitstream 14 with reference to other pictures (e.g., for prediction of residuals, motion vectors, and / or syntax elements). For example, a picture can refer to another picture of the same access unit (referred to as a reference picture of the picture), which is associated with another layer, which can be referred to as an inter-layer reference picture. Additionally or alternatively, a picture can refer to a reference picture that is part of the same layer but of a different access unit than the picture.
[0039] The extractor 30 can receive the video bitstream 14 and can select an OLS from the one or more OLSs indicated by the video bitstream 14, e.g., based on the indication 32 provided to the extractor 30. The extractor 30 can provide the sub-bitstream 12 indicated by the selected OLS by forwarding at least the bitstream portions 16 belonging to the selected OLS in the sub-bitstream 12. Note that the extractor 30 can modify or adapt one or more of the bitstream portions 16, such that the forwarded bitstream portions do not necessarily correspond exactly to the bitstream portions 16 signaled in the video bitstream 14. In Figure 1 In the figures, an apostrophe (e.g., reference sign 16') is used for bitstream portions of the sub-bitstream 12 in order to indicate potential changes of the bitstream portions of the video bitstream 14 when forwarded into the sub-bitstream 12.
[0040] The sub-bitstream 12 can be decoded by the decoder 50 in order to obtain a decoded video sequence represented by the sub-bitstream 12. Note that in addition to the fact that the decoded video sequence can optionally represent only a part of the video sequence 20 such that the decoded video sequence differs from the video sequence 20 in terms of resolution, fidelity, frame rate, picture size, and video content (in terms of sub-picture extraction), the decoded video sequence can have distortions due to quantization losses.
[0041] Pictures 26 of video sequence 20 can comprise independently coded pictures that do not reference pictures of other access units. That is, for example, independently coded pictures are coded into video bitstream 14 without using inter-prediction (although, layer-wise prediction can optionally be used to code independently coded pictures, in terms of temporal prediction). Due to independent coding, a decoder can start decoding a video sequence at an access unit of an independently coded picture. Independently coded pictures can be referred to as Instantaneous Random Access Points (IRAP). Examples of IRAP pictures are IDR pictures and CRA pictures. In contrast, trailing (TRAIL) pictures can refer to pictures that reference pictures of another access unit that can precede the trailing picture in coding order (the order in which pictures 26 are coded into video stream 14). Bitstream portions 16 into which independently coded pictures are coded can be referred to as independently coded bitstream portions, e.g., IRAP NAL units, while bitstream portions 16 into which non-independently coded pictures 26 are coded can be referred to as non-independent bitstream portions, e.g., non-IRAP NAL units. It is also noted that not necessarily all bitstream portions of a picture are coded in the same way, independently coded and non-independently coded. For example, a first portion of pictures 26 of a first access unit (e.g., access unit 221) can be independently coded, while a second portion of pictures 26 of the first access unit can be non-independently coded. In such a case, in pictures 26 of a second access unit (e.g., access unit 222), a first portion of pictures 26 of the second access unit can be non-independently coded, and a second portion can be independently coded. In this way, the higher data rate of independent coding relative to non-independent coding can be distributed over multiple access units. Such coding can be referred to as General Decoder Refresh (GDR) because decoder 50 can have to decode several access units, i.e., a sequence of pictures on which the independently coded portions covering the entire pictures are distributed, before having decoded the entire pictures of the access units preceding the GDR period independently.
[0042] In the following, several concepts and embodiments will be described with reference to Figure 1 It is noted that features described with respect to an encoder, a video bitstream, an extractor or a decoder are to be understood as being descriptive of the other of these entities as well. For example, a feature described as being present in a video data stream is to be understood as being descriptive of an encoder configured to encode the feature into a video bitstream and a decoder or extractor configured to read the feature from a video bitstream. It is further noted that inferring information based on indications encoded into a video bitstream can equally be performed on the encoder side and the decoder side. It is further noted that the various aspects described in the following sections can be combined with each other.
[0043] 1. Randomly accessible sub-bitstream indication
[0044] This section refers to Figure 1 Embodiments according to the first aspect are described, wherein the details described in section 0 can optionally apply to embodiments according to the first aspect. Furthermore, details described with reference to other aspects can optionally be implemented in embodiments described in this section.
[0045] A randomly accessible bitstream portion can refer to a Figure 1 separately coded bitstream portion as described with reference to Figure 1 a separately coded picture (or correspondingly, bitstream portion) as described with reference to
[0046] Some embodiments according to the first aspect can relate to full IRAP level indication for unaligned IRAPs. Embodiments can relate to IRAP alignment implications in case of max_tid_il_ref_pics_plus1 == 0 (e.g. only reference IDR, or only reference one of IDR, CRA or GRD with ph_recovery_poc_cnt equal to 0, or any kind).
[0047] Figure 2 An example of an output layer set of a video sequence (e.g. video sequence 20) is shown. In other words, Figure 2 The video sequence of can represent a video sequence formed by the layers of the OLS, including the output layer LI and the non-output layer L0. In access unit 22*, the multi-layer OLS bitstream of Figure 2 includes an example of unaligned IRAPs, i.e. access unit 22* includes a non-randomly accessible bitstream portion in one of the layers of the OLS, i.e. Figure 2 L1 in.
[0048] A bitstream including an indication of the level information of a full IRAP sub-bitstream (i.e. the result of discarding all non-IRAP NAL units from the bitstream) can be useful for trick mode playback (e.g. fast forward based on IRAP pictures only). This level information refers to a level_idc indication pointing to a list of limits defined for parameters such as maximum picture size, maximum picture rate, maximum bit rate, maximum buffer size, maximum slices / tile / subpictures per picture and minimum compression ratio. However, in the multi-layer case, it is not uncommon that IRAP pictures are not aligned between layers, e.g. a higher (non-independent) layer has a larger IRAP distance, such that IRAPs are not as frequent in the higher layer as they are in the lower (reference) layer. This is the case in Figure 2The example of a decoded multi-layer video bitstream is shown in Figure 2, wherein the higher layer LI of the depicted two-layer OLS does contain indeed trailing NAL units at the position of POC (Picture Order Count) 3 (i.e. in access unit 22*), while the lower layer L0 contains an IRAP NAL unit at the same position.
[0049] Figure 3 An example of an extracted full IRAP sub-bitstream extracted from a multi-layer video bitstream of Figure 2 Figure 2 is shown, as it can be performed conventionally. In view of the fact that not all layers of an OLS are to be output by the decoder, e.g. only the higher layer LI is marked as output layer in the example of Figure 2 Figure 2, it would be a waste of decoder resources to keep the lower (non-output) layer IDR in the bitstream, as long as there is no IDR at the corresponding position of the higher (output) layer, as is the case for picture 260*. The reason is that the output of the decoder is the same whether or not the L0 IDR at POC 3 is in the full IRAP sub-bitstream, the decoder does not output any picture when decoding the full IRAP sub-bitstream at POC 3. Moreover, when using the full IRAP level indication to approximate the maximum playback speed of a full IRAP presentation (e.g. in relation to the levels and playback speeds of a full OLS bitstream), decoding said L0 IRAP at POC 3 reduces the achievable maximum playback speed of the full IRAP sub-bitstream.
[0050] It is therefore part of embodiments of the present application to omit / discard such bitstreams, thereby excluding from consideration all IRAP NAL units in all access units that do not have an IRAP NAL unit in the corresponding output layer of the OLS of such full IRAP sub-bitstream, also by the respective level indication.
[0051] According to embodiments of the first aspect, e.g. with reference to Figure 1The described video bitstream 14 represents an encoded video sequence 20, the video bitstream 14 comprising a sequence of access units 22, each access unit comprising one or more bitstream portions 16, wherein each bitstream portion is associated with one of a plurality of layers 24 of the video bitstream 14. Each of the bitstream portions 16 is one of a bitstream portion type, including a random-accessible bitstream portion type, e.g. an independently encoded bitstream portion type, like an IRAP type. The video bitstream 14 comprises, e.g. encoded by the encoder 10 and to be detected by the extractor 30, an OLS indication 18 for an OLS of the video bitstream 14 and an extractability indication 19 for a random-accessible sub-bitstream described by the OLS, the OLS comprising one or more output layers and one or more non-output layers. The random-accessible sub-bitstream can be, e.g., a full IRAP sub-bitstream. The OLS indication can comprise, e.g., a level indication for the full IRAP sub-bitstream. Note that the term full IRAP sub-bitstream is not to be understood such that the full IRAP sub-bitstream must exclusively comprise random-accessible or IRAP bitstream portions. Rather, the random-accessible sub-bitstream or the full IRAP sub-bitstream can comprise, in some cases, non-IRAP or non-random-accessible bitstream portions, e.g. bitstream portions of reference pictures of the random-accessible bitstream portions. In other examples, the random-accessible sub-bitstream can exclusively comprise random-accessible bitstream portions.
[0052] According to an embodiment, the encoder 10 provides the video bitstream 14 such that, for each layer in the OLS, for each of the access units 22, outside of the bitstream portions of the respective access unit 22, if the respective access unit 22 comprises one of the random-accessible bitstream portions, the bitstream portions of all of the output layers in the output layer are random-accessible bitstream portions.
[0053] In other words, in one embodiment, the requirement of bitstream conformance is that the bitstream for which the full IRAP level indication is indicated does not contain access units without output pictures in all output layers.
[0054] There can be use cases where not all output layers have pictures in each access unit in the original bitstream, e.g. stereoscopic video with different frame rates for each eye. In this case, the bitstream requirement should be less strict.
[0055] In another embodiment, the requirement of bitstream conformance is that the bitstream for which the full IRAP level indication is indicated does not contain access units without output pictures in one output layer.
[0056] According to embodiments, the encoder 10 according to the first aspect can thus provide the video bitstream 14 such that, for each layer indicated by the OLS of the randomly accessible sub-bitstream, for each of the access units 22, outside the bitstream portion of the respective access unit 22, the bitstream portion 16 of at least one of the output layers which comprises one of the randomly accessible bitstream portions is a randomly accessible bitstream portion.
[0057] The result of both (the randomly accessible bitstream portion in at least one of the output layers or in all of the output layers) is that access units which do not satisfy the bitstream constraint are either not created at the encoder side or discarded during extraction. In other words, the bitstream containing only the indication of the IRAP level is constrained such that there is no AU with IRAP in a non-output layer and with a non-IRAP NAL unit in a temporally co-located picture in an output layer.
[0058] Alternatively, for each access unit, for each of the bitstream portions of the respective access unit, the OLS indication 18 and the extractability indication 19 are descriptive (the OLS indication and the extractability indication are encoded by the encoder 10 into the video bitstream 14) if one of the following two conditions is satisfied for the respective bitstream portion. The respective bitstream portion (e.g., the bitstream portion of the picture 26* in Figure 2 The first condition is satisfied if the respective bitstream portion is a randomly accessible bitstream portion and the respective bitstream portion is associated with one of the one or more output layers. The second condition is satisfied if the respective bitstream portion is associated with a reference layer of one of the output layers and the respective bitstream portion is associated with one of the one or more non-output layers and additionally, outside the bitstream portions of the respective access unit, the bitstream portion of at least one of the output layers is a randomly accessible bitstream portion (the bitstream portion of the picture 26** in Figure 2 The first condition is satisfied if the respective bitstream portion is a randomly accessible bitstream portion and the respective bitstream portion is associated with one of the one or more output layers. The second condition is satisfied if the respective bitstream portion is associated with a reference layer of one of the output layers and the respective bitstream portion is associated with one of the one or more non-output layers and additionally, outside the bitstream portions of the respective access unit, the bitstream portion of at least one of the output layers is a randomly accessible bitstream portion (the bitstream portion of the picture 26** in Figure 2 The same is true in the case of the picture 26** of
[0059] According to embodiments, the apparatus (e.g., the apparatus 30 of Figure 1 is configured to provide the sub-bitstream 12 indicated by the OLS indication of the randomly accessible sub-bitstream.
[0060] In other words, as an alternative to the option of bitstream constraint, i.e. the full IRAP level indication is indicated that the bitstream does not contain access units without output pictures in all output layers, according to embodiments, the indicated level does not include such AUs with this "mix" of IRAP units and non-IRAP NAL units, thus requiring discarding such AUs when expecting an IRAP-only bitstream with the indicated level.
[0061] A similar case will be considered when the output layer has an IRAP NAL unit but the reference layer does not have an IRAP NAL unit. In another embodiment, as long as the output layer has an IRAP NAL unit, the NAL units in the co- temporal reference layer are not discarded and considered for the indicated level. For this, there is a bitstream constraint, i.e. the temporal reference in the co-located reference layer does not have an IRAP NAL unit only refers to pictures that are also co-temporal with an IRAP NAL unit in the output layer. Alternatively, the indicated level only applies to AUs with all NAL units being IRAP NAL units and all other NAL units are discarded when considering such a bitstream for such a level (IRAP-only).
[0062] In another embodiment, instead of referring to the selection of layers for OLS, the requirement of considering AUs with all NAL unit types being IRAP types for IRAP level indication only applies to the whole bitstream. In this case, since such AUs contain an access unit delimiter with aud_irap_or_gdr_au_flag being 1, the presence of an access unit delimiter with aud_irap_or_gdr_au_flag being 1 is used to determine whether the AU is considered for IRAP-only level indication when there is an IRAP NAL unit (i.e. GDR case is not considered).
[0063] According to examples of the first aspect, the video bitstream 14 comprises a level indication for a randomly accessible sub-bitstream (e.g. a randomly accessible sub-bitstream described as extractable according to extractability information 19). The level indication (also referred to as level information) can for example indicate a level associated with a bitstream constraint by means of a list of levels as described in the 0th part. For example, the level indication is associated with one or more of a CPD size, a DPB size, a picture size, a picture rate, a minimum compression ratio, a picture partitioning restriction (e.g. tiles / slices / sub-pictures), a HRD timing (e.g. access unit / DU removal time, DPB output time).
[0064] In other words, in addition to the level indication, other parameters are relevant when considering an extracted bitstream with IRAP access units only. These parameters are DPB parameters and HRD parameters.
[0065] According to an embodiment of the first aspect, the decoder (e.g. decoder 50) is configured for checking whether the picture buffer complies with the random accessible sub-bitstream according to the extractability information 19. For example, the decoder 50 can further check the level indication in the extractability information, the HRD parameters in the above-mentioned parameters, the DPB parameters. The picture buffer can refer to the coded picture buffer and / or the decoded picture buffer of the decoder. Optionally, the decoder 50 can be configured for deriving timing information for the random accessible sub-bitstream (e.g. for the picture buffer) indicated by the OLS indication 18 from the video bitstream 12. The decoder 50 can decode the random accessible sub-bitstream based on the timing information.
[0066] In other words, in fact, omitting such full IRAP variant of IRAPs in non-output layers that are not accompanied by an IRAP in the output layer within the same access unit would also allow to reduce the DPB requirements (i.e. DPB size in terms of picture slots) as the decoding of non-output pictures that are not used for reference would be omitted. Note that this is a separate part of the level restrictions of the bitstream and not directly related to the restrictions defined by level_idc of the bitstream. This level sets a limit in combination with the picture size of the bitstream on the maximum number of pictures that can be held in the DPB. However, the DPB parameters include more information, e.g. how much maximum reordering of the output pictures is there, i.e. how many pictures can precede another picture in output order but follow it in output order. This information can be different when the extracted bitstream contains only IRAP pictures. Therefore, part of the invention is to signal additional DPB parameters for this representation to allow the decoder to better utilize its resources. One embodiment of the invention is given in table 1 below.
[0067] Table 1
[0068]
[0069]
[0070] In table 1, level_indication_for_all_irap_present indicates that there is level information for a full IRAP representation excluding non-output IRAP pictures that are not accompanied by an output layer IRAP in their respective access unit.
[0071] For example, vps_ols_dpb_params_all_irap_idx[ i ] specifies the index of the dpb_parameters( ) syntax structure in the list of dpb_parameters( ) syntax structures in the VPS that applies to the i-th multi-layer OLS when considering only IRAP sub-bitstreams. When present, the value of vps_ols_dpb_params_idx[ i ] shall be in the range of 0 to VpsNumDpbParams - 1, inclusive. When vps_ols_dpb_params_all_irap_idx[ i ] is not present, it is inferred to be equal to vps_ols_dpb_params_idx[ i ]. For single-layer OLS, the applicable dpb_parameters( ) syntax structure is present in the SPS referred to by the layer in the OLS. Each dpb_parameters( ) syntax structure in the VPS shall be referred to by at least one of the values of vps_ols_dpb_params_idx[ i ] or vps_ols_dpb_params_all_irap_idx[ i ], where i is in the range of 0 to NumMultiLayerOlss - 1, inclusive.
[0072] As noted, another additional information that an extracted bitstream containing only IRAP NAL units can need is HRD parameters. HRD parameters can include, for example, one or more or all of: required CPB size, time at which access units are removed from the CPB, bit rate at which the CPB is fed, or whether the resulting bitstream after extraction corresponds to a constant bit rate representation.
[0073] 2. Reference picture alignment
[0074] Part 2 References Figure 1 Embodiments according to the second aspect are described, wherein the details described in part 0 can optionally apply to embodiments according to the second aspect. Furthermore, details described with reference to other aspects can optionally be implemented in embodiments described in this part.
[0075] In VVC, output layer set defines the prediction dependency relationship between layers of a bitstream. The syntax element vps_max_tid_il_ref_pics_plus1[ i ][ j ] signaled for all direct reference layers of a given layer allows to further limit the number of pictures in the reference layers used for prediction, as follows:
[0076] vps_max_tid_il_ref_pics_plus1[ i ][ j ] equal to 0 specifies that no picture in the j-th layer that is neither an IRAP picture nor a GDR picture with ph recovery poc cnt equal to 0 is used as an ILRP for decoding pictures of the i-th layer. vps_max_tid_il_ref_pics_plus1[ i ][ j ] greater than 0 specifies that for decoding pictures of the i-th layer, no picture from the j-th layer with Temporalld greater than vps_max_tid_il_ref_pics_plus1[ i ][ j ] - 1 is used as an ILRP. When not present, the value of vps_max_tid_il_ref_pics_plus1[ i ][ j ] is inferred to be equal to vps_max_sublayers_minus1 + 1.
[0077] Note that when not present, the value is inferred to be vps_max_sublayer_minus1 + 1, where vps_max_sublayer_minus1 is the maximum number of sub-layers in any layer present in the bitstream. However, the value of the maximum number of sub-layers can be smaller for a particular layer.
[0078] This syntax element indicates a special mode (vps_max_tid_il_ref_pics_plus1[ i ][ j ] equal to 0) where inter-layer referencing is not used for some sub-layers or some sub-layers of the reference layer do not need to be decoded and only IRAP NAL units or GDR NAL units with ph recovery poc cnt equal to 0 in the reference layer need to be decoded. Furthermore, the output layer set of the bitstream delivered to the decoder does not include the unnecessary NAL units as indicated by this syntax element vps_max_tid_il_ref_pics_plus1[ i ][ j ] or such NAL units are discarded in a particular decoder implementation of the extraction process defined in the implementation specification.
[0079] The syntax element vps_max_tid_il_ref_pics_plus1[ i ][ j ] is only present for direct reference layers. For example, consider an OLS with 3 layers as shown in Figure 4 Figure 4 Three-layer example is shown where vps_max_tid_il_ref_pics_plus1 is equal to 0 for all reference layers L0, L1, L2, where L2 is a direct reference layer of L1 and an indirect reference layer of L0, as L1 uses L0 as a direct reference layer. In this case, if vps_max_tid_il_ref_pics_plus1[2][1] is equal to 0, only IRAP NAL units or GDR NAL units with ph_recovery_poc_cnt equal to 0 are kept in L1 and thus in L0, as indicated by the specification. More specifically, a variable indicating the number of sub-layers kept is derived for each layer in the OLS, i.e. NumSubLayersInLayerInOLS[i][j], where i is the OLS index and j is the layer index. The value of this variable is set to the maximum temporalld expected by the bitstream for the layer when it is an output layer. When it is not an output layer but it is a reference layer of each layer k in the OLS using layer j as reference, the value of NumSubLayersInLayerInOLS[i][j] is set to the maximum of min(NumSubLayersInLayerInOLS[i][k], vps_max_tid_il_ref_pics_plus1[k][j]). That is, for each layer k, the minimum between how many sub-layers are needed in layer k (NumSubLayersInLayerInOLS[i][k]) and how many sub-layers are needed in layer j if all sub-layers in layer k are considered (vps_max_tid_il_ref_pics_plus1[k][j]) is checked. The minimum of both is selected because if less sub-layers are needed in layer k than indicated in vps_max_tid_il_ref_pics_plus1[k][j], then only the same number of sub-layers is needed in layer j. Then, a check is also made for other layers k that also use j as reference, and if the other layers indicate more sub-layers are needed, the higher value is taken, i.e. the maximum value once all layers using layer j as reference are checked.
[0080] A problem arises when an IRAP NAL unit is not aligned between L0 and L1. For example, as shown in Figure 5 , it is assumed that L1 has an IRAP AU at some point but L0 does not have an IRAP AU, and the IRAP AU in L1 uses a non-IRAP AU in L0 as reference, Figure 5 A three-layer example is shown where the non-aligned IRAP is in the lower layer. In this case, the IRAP-based extraction process discards the non-IRAP Figure 5260* in the picture), so the IRAP AU in L1 ( Figure 5 The picture 261*) in cannot be decoded.
[0081] In an embodiment, when vps_max_tid_il_ref_pics_plus1[i][j] of layer i is 0, any direct or indirect layer of such layer i is required to have aligned IRAP or GDR NAL units with ph_recovery_poc_cnt equal to 0. In other words, in any indirect reference layer, NAL units of the indirect reference layer that depend on a synchronous NAL unit in the direct reference layer of the indirect reference layer being either an IRAP NAL unit or a GDR NAL unit with ph_recovery_poc_cnt equal to 0 are also required to be either an IRAP NAL unit or a GDR NAL unit with ph_recovery_poc_cnt equal to 0.
[0082] According to an embodiment of the second aspect, a video bitstream 14 comprises a sequence of access units 22, each access unit comprising one or more bitstream parts 16. Each of the bitstream parts 16 is associated with one of a plurality of layers 24 of the video bitstream 14 and a plurality of temporal layers of the video bitstream (e.g., reference 1). Figure 1 The bitstream portions 16 within the same access unit 22 are associated with the same temporal sub-layer. Furthermore, each of the bitstream portions is one of a set of bitstream portion types. For example, the set of predetermined bitstream portion types may include independently coded bitstream portion types (e.g., IDR) and, optionally, other access unit types that are dependent only on the set of predetermined bitstream portion types.
[0083] According to an embodiment of the second aspect, the encoder 10 is configured to provide, in the video bitstream 14, an OLS indication of an OLS of the video bitstream 14, the OLS comprising one or more layers of the video bitstream. Further, for each layer in the OLS, the encoder 10 provides, in the video bitstream, a reference layer indication indicating a set of reference layers on which the respective layer depends. Further, the encoder 10 provides, in the video bitstream 14, a temporal layer indication (e.g., vps_max_tid_il_ref_pics_plus1[i][j]) for each layer (e.g., i) in the OLS, for each reference layer (e.g., j) of the respective layer, indicating whether all bitstream parts of the respective reference layer on which the respective layer depends are of a predetermined bitstream part type of a set of predetermined bitstream part types, or, if not, indicating to which temporal layer the bitstream part belongs (i.e., a subset of the multiple temporal layers to which all bitstream parts of the respective reference layer belong (e.g., indicated by the maximum index indexing the temporal layers)) on which the respective layer depends.
[0084] The encoder 10 according to this embodiment is configured to provide the video bitstream such that, for each layer (e.g., i) in the OLS, the temporal layer indication indicates that all bitstream parts of a predetermined reference layer (of the reference layers of the respective layer) on which the respective layer depends are of a predetermined bitstream part (e.g., the same) type of a set of predetermined bitstream part types, for each other reference layer directly or indirectly depending on the predetermined reference layer, there are no bitstream parts other than the set of predetermined bitstream part types (e.g., direct dependency or direct reference is a dependency between a (non-independent) layer and its reference layers as indicated e.g. in the reference layer indication, and indirect dependency or reference is a dependency between a (non-independent) layer and a direct or indirect reference layer of a (non-independent) layer not indicated in the reference layer indication) including access units of the predetermined reference layer that are bitstream parts of the predetermined bitstream part type of the set of predetermined bitstream part types.
[0085] In fact, this is a bit more restrictive than necessary. As Figure 6 indicated, such indirect reference layers (L0) can also be reference layers of another layer (now assume the case of a 4th L3 layer indicating that sublayer 0 and sublayer 1 of L0 are required by vps_max_tid_il_ref_pics_plus1[3][0] being equal to 2), Figure 6 shows a four-layer example with direct reference to sublayers having Tid 1.
[0086] In this case, the IRAP alignment constraints discussed in the previous embodiments would not be necessary, as the non-IRAP NAL units in layer 0 that are required for the IRAP NAL units in L1 would remain in the OLS bitstream corresponding to L0+L1+L2+L3. Thus, the variable NumSubLayersInLayerInOLS[i][j] can instead be used to express the constraint, where this variable indicates the number of sub-layers (with the corresponding temporal ID) that are retained in the i-th OLS of the j-th layer (0 means only IRAP or GDR with ph recovery poc cnt equal to 0 are retained).
[0087] In an embodiment, when within the i-th OLS, two layers k and j (with k>j) have NumSubLayersInLayerInOLS[i][j] and NumSubLayersInLayerInOLS[i][k] equal to 0, if j is a (direct or indirect) reference layer of k, the IRAP NAL units or GDRs with ph recovery poc cnt equal to 0 are aligned.
[0088] Thus, as a reference Figure 4 and Figure 5 As an alternative to the embodiments described, the encoder 10 can provide the bitstream 14 such that, for each layer (e.g., i) 20B in an OLS, the temporal layer indication indicates that all bitstream portions of the predetermined reference layers (in the reference layers of the respective layer) on which the respective layer 20B depends are of a predetermined bitstream portion type (e.g., the same type) in a predetermined set of bitstream portion types, and for another reference layer 20D on which the predetermined reference layers 20C directly or indirectly depend, the following two criteria are met: as a first criterion, the access units 40A, 40B that include bitstream portions of the predetermined reference layers that are of a predetermined bitstream portion type in the predetermined set of bitstream portion types have no bitstream portions (40A, or not 40B) other than the predetermined set of bitstream portion types. As a second criterion, according to the reference layer indication, the respective other reference layer 20D is a reference layer of a direct reference layer 20A that depends on the respective layer 20B according to the reference layer indication.
[0089] According to an alternative embodiment, the encoder 10 is configured for providing, in the video bitstream 14, in addition to the OLS indication and the reference layer indication described with respect to the previous embodiments, an intra-layer temporal layer indication [e.g., NumSubLayersInLayerInOLS[i][j]] for each layer [e.g., j] in an OLS [e.g., i] indicating whether [e.g., by NumSubLayersInLayerInOLS[i][j] = 0] only bitstream parts of the respective layer are required that are of a predetermined bitstream part type of a set of predetermined bitstream part types as a predetermined bitstream part type part or, if not, indicating a temporal layer [e.g., a maximum temporal layer index] subset of bitstream parts required for the OLS including the respective layer.
[0090] According to this embodiment, the encoder 10 is configured for providing the bitstream 14 such that for each layer 24 in an OLS, the intra-layer temporal layer indication indicates that only bitstream parts of the respective layer are required for the OLS that are of a predetermined bitstream part type of a set of predetermined bitstream part types, for each access unit including a bitstream part of a predetermined bitstream part type of the set of predetermined bitstream part types, for each of the bitstream parts of the respective access unit, the following condition is fulfilled: if the respective bitstream part belongs to a layer in the OLS whose intra-layer temporal indication indicates that only bitstream parts of the respective layer are required for the OLS that are of a predetermined bitstream part type of the set of predetermined bitstream part types, the respective bitstream part is of the predetermined bitstream part type of the set of predetermined bitstream part types, or according to the reference layer indication, the respective layer does not depend on the layer of the respective bitstream part.
[0091] According to another embodiment, if IRAP is not aligned, inter-layer prediction is not used for such IRAP NAL units of a layer for which a non-RAP NAL unit is present at the same AU.
[0092] According to another embodiment of the second aspect, therefore, the encoder 10 is configured for providing in the video bitstream 14 the OLS indication, the reference layer indication and the temporal layer indication described in the previous embodiments of reference part 2. According to this embodiment, furthermore, the encoder 10 is configured to encode, for each layer (e.g. i) in the OLS, whose temporal layer indication indicates that all bitstream portion layers of a predetermined reference layer (in the reference layer of the respective layer) on which the respective layer depends are of a predetermined set of bitstream portion types, for each other reference layer directly or indirectly depended on the predetermined reference layer, the access units comprising bitstream portions of the predetermined reference layer that are of the predetermined set of bitstream portion types as the predetermined bitstream portion types, bitstream portions of the access units comprising bitstream portions of the predetermined reference layer that are of the predetermined set of bitstream portion types as the predetermined bitstream portion types, except for the predetermined set of bitstream portion types, without using an intra prediction method for bitstream portions belonging to layers that directly or indirectly refer to a reference layer that does not have bitstream portions other than the predetermined set of bitstream portion types, for the respective layer.
[0093] For example, the predetermined set of bitstream portion types can comprise one or more or all of the following: IRAP type and GDR type with ph_recovery_poc_cnt equal to 0.
[0094] Embodiments of the encoder 10 according to the second aspect can be configured for providing in the video bitstream 14 level indications of the bitstream 12 that can be extracted from the video bitstream according to the OLS. For example, the level indications comprise one or more of the following: coded picture buffer size, decoded picture buffer size, picture size, picture rate, minimum compression ratio, picture partitioning restrictions (e.g. tiles / slices / subpictures), bit rate, buffer scheduling (e.g. HRD timing (AU / DU removal times, DPB output times)).
[0095] 3. Bitstream based OLS determination
[0096] Reference part 3 Figure 1 Embodiments according to the third aspect of the application are described, wherein details described in part 0 can optionally apply to embodiments according to the third aspect. Furthermore, details described with reference to other aspects can optionally be implemented in embodiments described in this part.
[0097] Embodiments of the third aspect can provide an identification of the OLS corresponding to the bitstream. In other words, according to embodiments of the third aspect, it is allowed to infer from the video bitstream the OLS for which the OLS of the video bitstream is to be decoded or extracted. A decoder receiving a bitstream to decode can get additional information via its API about for which operation point it should decode. For example, in the current VVC draft specification, two variables are set as follows via external means.
[0098] The variable TargetOlsIdx identifying the OLS index of the target OLS to be decoded and the variable Htid identifying the highest temporal sub-layer to be decoded are set by some external means not specified by the present specification. The bitstream BitstreamToDecode does not contain any layer other than the layers included in the target OLS and does not contain any NAL unit with Temporalld greater than Htid.
[0099] The specification does not say what to do when these variables are not set, as it is expected that in this case the decoder simply decodes the whole bitstream it gets, and not a subset of the bitstream, e.g. in terms of temporal sub-layers.
[0100] However, there is a problem with output layer sets as follows. When a decoder gets a bitstream containing more than one layer and the parameter sets define more than one OLS containing all layers in the bitstream (e.g. variants with different output layers), it cannot simply be determined from the bitstream itself which output layer set the decoder should decode. Depending on the OLS properties, the OLS to be selected can e.g. impose different level requirements due to different DPS parameters, etc. Thus, it is essential to allow the decoder to select the OLS even in the absence of external signaling via its API. In other words, a fallback method is needed as in other cases without external means, e.g. selecting the highest temporal sub-layer in the bitstream to decode, etc.
[0101] In one embodiment, there is a constraint on bitstream conformance, i.e. the bitstream can only correspond to a single OLS, so that the decoder can clearly determine from the bitstream it gets which OLS to decode. This property can e.g. be instantiated by a syntax element indicating that all OLSs can be unambiguously determined from the layers present in the bitstream, i.e. there is a unique mapping from the number of layers to the OLS.
[0102] According to embodiments of the third aspect, an encoder 10 for providing a multi-layer video bitstream 14 is configured to indicate within the multi-layer video bitstream 14, e.g. Figure 1The OLSs in the OLS indication 18. Each of the OLSs indicates a subset of layers of the multi-layer video bitstream 14. Note that the subset of layers can not necessarily be a proper subset of the layers, i.e. the subset of layers can comprise all layers of the multi-layer video bitstream 14. The encoder 10 according to this embodiment provides the multi-layer video bitstream 14 such that for each of the OLSs, the sub-bitstream of the multi-layer video bitstream 14 defined by the respective OLS (e.g. the sub-bitstream 12) is distinguishable from the sub-bitstreams of the multi-layer video bitstream defined by any of the other OLSs. For example, the encoder 10 can provide the multi-layer video bitstream such that the OLSs indicate mutually different subsets of layers of the multi-layer video bitstream 14, such that the OLSs are distinguishable by their subset of layers.
[0103] For example, each of the OLSs can be defined by indicating the subset of layers in the OLS indication 18 for the OLS by means of layer indices. Optionally, the OLS indication can comprise further parameters defining the subset of bitstream portions of the layers of the OLS that belong to the OLS. For example, the OLS indication can indicate which temporal sub-layers belong to the OLS.
[0104] In an example, the encoder 10 can indicate within the multi-layer video bitstream 14 that the multi-layer video bitstream 14 can unambiguously belong to one of the OLSs. For example, the encoder 10 can indicate the plurality of OLSs such that for each of the OLSs, the subset of layers of the respective OLS is different from any of the subsets of layers of the other OLSs. Thus, in an example, the encoder 10 can indicate within the multi-layer video bitstream 14 that a set of layers of the multi-layer video bitstream 14, e.g. which can be indicated by a set of indices, can unambiguously belong to one of the OLSs.
[0105] According to an embodiment, the encoder 10 is configured for checking the consistency of the multi-layer video bitstream 14 by checking, for each of the plurality of OLSs, whether the sub-bitstream of the multi-layer video bitstream 14 defined by the respective OLS is distinguishable or different from the sub-bitstreams of the multi-layer video bitstream defined by any of the other OLSs. For example, the encoder 10 can reject the bitstream consistency if this is not the case.
[0106] Thus, an embodiment of a decoder for decoding a video bitstream, e.g. the decoder 50 of Figure 1 Thus, an embodiment of a decoder for decoding a video bitstream, e.g. the decoder 50 of Figure 1OLSs, each OLS indicating a subset of layers of the video bitstream, when the video bitstream is decoded. The decoder can detect an indication within the video bitstream that indicates that the video bitstream can be unambiguously attributed to one of the OLSs and can decode the one of the OLSs to which the video bitstream can be attributed.
[0107] For example, the decoder 50 can identify the OLSs to which the video bitstream can be attributed by identifying the layers included in the video bitstream and decode the OLSs, which exactly identifies the layers included in the video bitstream. Thus, the indication that indicates that the video bitstream can be unambiguously attributed to one of the OLSs can indicate that the set of layers within the video bitstream can be unambiguously attributed to one of the OLSs.
[0108] For example, the decoder 50 can determine the one of the OLSs to which the video bitstream can be attributed by examining the first of the access units of the coded video sequence. The first access unit can refer to the first received, the first in time order, the first in decoding order, or the first in output order. Alternatively, the decoder 50 can determine the one of the OLSs by examining the first of the access units that is a sequence-start access unit type (e.g., a CVS S access unit), the first access unit being defined by, for example, the reception order, the time order, the decoding order. For example, the decoder 50 can examine the first of the access units of the coded video sequence or the first of the access units that is a sequence-start access unit type with respect to the layers included in the respective access unit.
[0109] For example, the decoder 50 can determine the one of the OLSs such that, for the first access unit of the coded video sequence or the first of the access units that is a sequence-start access unit type, the respective access unit only includes pictures of the layers of the one OLS.
[0110] For example, in a multi-view two-layer scenario with an OLS outputting two views and one OLS outputting only one independently coded view, some of the above embodiments can have the drawback of prohibiting certain combinations of OLSs. To mitigate this limitation, another embodiment of the present application, for example, by one or more of the following combinations, in the OLSs corresponding to the bitstream or the first access unit of the bitstream or the CVS S AU of the bitstream:
[0111] • selecting the OLS with the highest or lowest index,
[0112] • selecting the OLS with the most number of output layers.
[0113] Figure 7 A decoder 50 according to an embodiment is shown, which can optionally be according to the latter embodiment with the selection algorithm. According to Figure 7The decoder 50 can optionally correspond to a decoder 50 according to Figure 1 The decoder 50 according to Figure 7 The decoder 50 according to Figure 1 The decoder 50 according to Figure 7 The video bitstream 12 can correspond to the video bitstream 14 according to Figure 1 The video bitstream 12 can correspond to the video bitstream 14 according to Figure 7 The video bitstream 12 can be, but is not necessarily, a multi-layer video bitstream, i.e. in examples it can be a single-layer video bitstream. The video bitstream 12 comprises access units 22 (e.g. access units 221, 222) of an encoded video sequence 20, each access unit comprising one or more pictures 26 (e.g. pictures 261, 262) of the encoded video sequence. Each of the pictures 26 belongs to one of one or more layers 24 of the video bitstream 12, e.g. as described in the 0th part. The decoder 50 according to Figure 7 The decoder 50 according to Figure 1 The decoder 50 according to Figure 7 The decoder 50 according to
[0114] As described with respect to the previous embodiment, the decoder 50 can determine the OLS subset and the one OLS by inspecting the first of the access units of the encoded video sequence or the first of the access units that is a sequence-start access unit type. For example, the access unit 221 as shown in Figure 7 may be the first (e.g. the first received or the first in encoding order or the first in time order or the first in output order) of the access units of the encoded video sequence of the video bitstream 12. In other examples, the encoded video sequence 20 can comprise further access units before the access unit 221, the preceding access units not being sequence-start access units, so that the access unit 221 is the first sequence-start access unit of the sequence 20. The decoder 50 can inspect the access unit 221 to detect the picture 261 of the first layer 241 and the picture 262 of the second layer 242. Based on this finding, the decoder 50 can conclude that the video bitstream 12 comprises the first layer 241 and the second layer 242.
[0115] according to Figure 7 In an embodiment of the present invention, the decoder 50 determines an OLS to be decoded based on one or more attributes of the OLS. The one or more attributes may include one or more of the following: an index of the corresponding OLS (i.e., OLS index) and / or the number of layers of the OLS and / or the number of output layers of the corresponding OLS. For example, the decoder determines the OLS with the highest or lowest OLS index and / or the largest number of layers and / or the largest number of output layers among the OLSs as the one OLS. In other words, the decoder 50 may evaluate which of the OLSs has the highest or lowest OLS index and / or which of the OLSs has the largest number of layers. Additionally or alternatively, the decoder 50 may evaluate which of the OLSs includes the output layer with the highest or lowest index. By selecting the one OLS according to the largest number of output layers and / or the largest number of layers, a bitstream that provides the highest quality output for the video sequence may be selected for decoding.
[0116] For example, the decoder 50 may select an OLS having the largest number of output layers, an OLS having the largest number of layers outside the OLS having the largest number of output layers, or an OLS having the lowest OLS index outside the OLS having the largest number of layers outside the OLS having the largest number of output layers as the one OLS.
[0117] According to an embodiment, the decoder 50 determines one OLS among the OLSs by evaluating which OLS among the OLSs has the largest number of layers, and in the case where there are multiple OLSs having the largest number of layers, the decoder 50 can evaluate which OLS among the OLSs with the largest number of layers has the largest number of output layers, and can select the OLS with the largest number of output layers as the one OLS out of the OLS with the largest number of layers.
[0118] In other words, in the absence of an indication of an OLS to encode, the decoder 50 may decode the OLS indicated in the OLS indication 18 where all required layers are present in the bitstream and use the OLS of the majority of the layers present, thereby providing a high fidelity video output.
[0119] Below, reference Figure 7 Another embodiment of the decoder 50 is described.
[0120] Furthermore, this further embodiment of the decoder 50 may optionally be based on the selection algorithm previously described and may optionally correspond to a method based on Figure 1 According to this other embodiment, the decoder 50 is configured to decode the video bitstream 12 (eg, a video bitstream 12 extracted from a multi-layer video bitstream 14, such as Figure 1 Alternatively, Figure 7The video bitstream 12 can correspond to Figure 1 The video bitstream 14. Figure 7 The video bitstream 12 can be, but is not necessarily, a multi-layer video bitstream, i.e. it can be, in an example, a single-layer video bitstream. The video bitstream 12 comprises access units 22 (e.g. access units 221, 222) of an encoded video sequence 20, each access unit comprising one or more pictures 26 (e.g. pictures 261, 262) of the encoded video sequence. Each of the pictures 26 belongs to one of one or more layers 24 of the video bitstream 12, e.g. as described in the 0th part. The decoder 50 according to this further embodiment is configured for deriving one or more OLSs from the video bitstream 12. For example, the video bitstream 12 comprises, e.g. in reference to Figure 1 The OLSs indicated 18 as described in the 0th part indicate one or more OLSs, such as OLS 181 and OLS 182, as shown in Figure 7 Each of the one or more OLSs indicates a (not necessarily true) set of one or more layers 24 of the video bitstream 12. In other words, each of the OLSs indicates one or more of the layers as part of the respective OLS. The decoder 50 according to this further embodiment determines a subset of OLSs from (or based on) the OLSs such that each of the OLSs of the subset of OLSs is attributable to the video bitstream 12. The decoder 50 according to this further embodiment determines one of the subset of OLSs based on one or more properties of each of the OLSs of the subset of OLSs and decodes the one OLS determined from the subset of OLSs.
[0121] For example, an OLS attributable to the video bitstream can represent that one or more sets of layers present in the video bitstream 12 correspond to the set of layers indicated in the respective OLS. In other words, the decoder 50 can determine the subset of OLSs attributable to the video bitstream 12 based on the set of layers present in the video bitstream 12. That is, the decoder 50 can determine the subset of OLSs such that for each of the OLSs of the subset of OLSs, the subset of layers indicated by the respective OLS corresponds to the set of layers of the video bitstream 12. For example, in Figure 7 In the example, the video bitstream 12 exemplarily comprises a picture 261 of a first layer 241 and a picture 262 of a second layer 242. The OLS 181 indicates that the first layer 241 and the second layer 242 are part of the OLS 181. Further, the OLS 182 indicates that both the first layer and the second layer are part of the OLS 182. Thus, according to the example, the decoder 50 can attribute both the OLS 181, 182 to the part of the index of the OLSs attributable to the video bitstream 12. Figure 7
[0122] Alternatively, the decoder 50 may consider for decoding only those OLSs that can be decoded by the decoder 50 according to the level information of the OLSs.
[0123] As described with respect to the previous embodiments, the decoder 50 may determine the subset of attributable OLSs and the one OLS by examining the first of the access units of the coded video sequence or the first of the access units that is a sequence start access unit type. Figure 7 The illustrated access unit 221 may be the first (e.g., the first received, the first in coding order, the first in temporal order, or the first in output order) access unit of the coded video sequence of the video bitstream 12. In other examples, the coded video sequence 20 may include other access units before the access unit 221, the preceding access unit not being a sequence-start access unit, and thus the access unit 221 is the first sequence-start access unit of the sequence 20. The decoder 50 may examine the access unit 221 to detect a picture 261 of the first layer 241 and a picture 262 of the second layer 242. Based on this finding, the decoder 50 may conclude that the video bitstream 12 includes the first layer 241 and the second layer 242.
[0124] According to an embodiment, the decoder 50 can determine the OLS subset so that for the first access unit of the encoded video sequence or the first access unit of the sequence start access unit type, for example, access unit 221, the corresponding access unit only includes pictures of the layers of each OLS in the OLS subset.
[0125] according to Figure 7 In an embodiment of the present invention, the decoder 50 determines one OLS to decode based on one or more properties of the subset of OLSs. The one or more properties may include one or more of the following: an index and / or number of output layers of the corresponding OLS, a highest or lowest index (i.e., highest or lowest layer index), and a maximum number of layers. In other words, the decoder 50 may evaluate which OLS of the subset of OLSs attributable to the video bitstream 12 includes a layer indexed by the highest or lowest layer index and / or which OLS of the OLSs has the maximum number of layers. Additionally or alternatively, the decoder 50 may evaluate which OLS of the subset of OLSs has the maximum or minimum number of output layers and / or which OLS of the subset of OLSs includes an output layer with the highest or lowest index.
[0126] According to an embodiment, the decoder 50 determines one OLS in the OLS subset by evaluating which OLS among the OLSs has the largest number of layers, and in the case where there are multiple OLSs with the largest number of layers, the decoder 50 can evaluate which OLS among the OLSs with the largest number of layers has the largest number of output layers, and can select the OLS with the largest number of output layers as the one OLS out of the OLS with the largest number of layers.
[0127] In other words, in the absence of an indication of an OLS to be decoded, the decoder 50 can decode all its required layers in the OLS indicated in the OLS indication 18 are present in the bitstream and use the OLS of the most present layers, thereby providing a high-fidelity video output.
[0128] In other words, embodiments of the third aspect comprise a decoder 50 for decoding a video bitstream 12, 14, wherein the video bitstream 14 comprises access units 22 of encoded video sequences 20 and wherein each access unit 22 comprises one or more pictures 26 of an encoded video sequence, wherein each of the pictures belongs to one of one or more layers 24 of the video bitstream 14. The decoder is configured to derive one or more output layer sets (OLSs) 181, 182 from the video bitstream 14, each OLS indicating a set of one or more layers of the video bitstream 14, determine a subset of the OLSs 181, 182 from the OLSs 181, 182, each OLS of the subset of OLSs being attributable to the video bitstream 14, determine one OLS of the subset of OLSs based on one or more properties of each OLS of the subset of OLSs, and decode the one OLS.
[0129] According to embodiments, the decoder 50 is configured to determine the subset of OLSs such that for each OLS of the subset of OLSs, the subset of layers indicated by the respective OLS corresponds to the set of layers of the video bitstream 14.
[0130] According to embodiments, the decoder 50 is configured to determine the subset of OLSs and the one OLS by inspecting a first of the access units of the encoded video sequence or a first of the access units being a sequence-start access unit type.
[0131] According to embodiments, the decoder 50 is configured to determine the subset of OLSs such that for a first of the access units 22 of the encoded video sequence or a first of the access units 22 being a sequence-start access unit type, the respective access unit only comprises pictures of the layers of each OLS of the subset of OLSs.
[0132] According to embodiments, the decoder 50 is configured to determine the one OLS of the OLSs by evaluating each criterion against the one or more properties of the OLSs.
[0133] 4. Sequence start access unit in sub-bitstreams
[0134] Part 4 refers to Figure 1 Embodiments of the fourth aspect of the application are described. The description provided in part 0 can optionally apply to embodiments of the fourth aspect. Furthermore, details described with reference to other aspects can optionally be implemented in embodiments described in this part.
[0135] Some embodiments of the fourth aspect can relate to an access unit delimiter (AUD) in supplemental enhancement information (SEI) to allow starting access units (CVSS AUs) that are not initially (i.e., from which a video bitstream such as video bitstream 12 is extracted) CVSS AUs in a video bitstream (e.g., video bitstream 14). For example, a CVSS AU can be an AU that is randomly accessible (e.g., having a randomly accessible or independently coded picture in each layer of the video bitstream), or an AU that can be decoded independently of previous AUs of the video bitstream.
[0136] Current specifications require a coded video sequence start (CVSS) AU to have an IRAP NAL unit type or a GDR NAL unit type at each layer, and the IRAP NAL unit type within the CVSS AU is the same. Furthermore, current specifications require the presence of an AUD (access unit delimiter) that indicates that the CVSS AU is an IRAP AU or a GDR AU.
[0137] Figure 8 An example of a multi-layer bitstream is shown, which does not have aligned IRAPs at different layers (i.e., not all IRAPs are aligned between layers) and the identification of CVSS AUs.
[0138] It can be seen that AUs 2, 4, and 6 have NAL unit types of IRAP type in the two lowest layers, but since not all layers have the same IRAP type in these AUs, these AUs are not CVSS AUs. To easily identify CVSS AUs without the need to parse all NAL units of an AU, an AUD nal unit is used in order to easily identify the CVSS AUs. This means that AUs 0 and 8 would contain an AUD with a flag indicating that these AUs are CVSS AUs (IRAP AUs).
[0139] Figure 9 An example of a video bitstream (e.g., video bitstream 12) after extraction is shown. However, when extracting a bitstream with only L0 and LI, new CVSS AUs are present in the extracted bitstream, as indicated by reference numeral 22*. In other words, after extraction, some AUs (e.g., AU 22*) become CVSS AUs.
[0140] Since AUs 2 and 6 become CVSS AUs or IRAP AUs, an AUD needs to be present in the bitstream at such AUs that indicate IRAP AU properties.
[0141] Embodiments according to the fourth aspect comprise a device for extracting a sub-bitstream from a multi-layer video bitstream, e.g., with reference to Figure 1An extractor 30 for extracting a sub-bitstream 12 from a multi-layer video bitstream 14 is described. According to a fourth aspect, the multi-layer video bitstream 14 represents an encoded video sequence, e.g. the encoded video sequence 20, and the multi-layer video bitstream comprises access units 22 of the encoded video sequence. Each of the access units 22 comprises one or more bitstream portions 16 of the multi-layer video bitstream 14, wherein each bitstream portion belongs to one of the layers 24 of the multi-layer video bitstream. According to the fourth aspect, the extractor 30 is configured for deriving one or more OLSs from the multi-layer video bitstream 14, each OLS indicating a (not necessarily true) subset of layers of the multi-layer video bitstream 14. For example, the multi-layer video bitstream comprises an OLS indication 18, the OLS indication 18 comprising information about or a description of one or more OLSs, e.g. the OLS 181, the OLS 182, as shown. Figure 7
[0142] The extractor 30 according to the fourth aspect is configured for providing within the sub-bitstream 12 the layers 24 of the multi-layer video bitstream 14 that are indicated by a predetermined one of the OLSs, i.e. the layers 24 that are indicated as part of the predetermined OLS. In other words, the extractor 30 can provide within the sub-bitstream 12 the bitstream portions 16 that belong to the respective layers of the predetermined OLS. For example, the predetermined OLS can be provided to the extractor 30 by external means, e.g. by the OLS indication 32, as shown. Figure 1 According to embodiments of the fourth aspect, e.g. for a sub-bitstream comprising the layers L0 and L1, the extractor 30 provides within the sub-bitstream 12 the layers L0 and L1 of the multi-layer video bitstream 14 that are indicated by a predetermined one of the OLSs, i.e. the layers L0 and L1 that are indicated as part of the predetermined OLS. Figure 8 and Figure 9 The access units 22* of the sub-bitstream described for L0 and L1, if all bitstream portions of the respective access unit are bitstream portions of the same predetermined bitstream portion type of a predetermined set of bitstream portion types, the extractor 30 provides within the sub-bitstream 12, for each access unit of the sub-bitstream 12, a sequence start indication indicating that the respective access unit is a start access unit of a sub-sequence of the encoded video sequence.
[0143] In other words, the extractor 30 can provide access units that the extractor 30 includes or provides within the sub-bitstream 12, which access units exclusively comprise bitstream portions of the same predetermined bitstream portion type of the predetermined set of bitstream portion types with a sequence start indication.
[0144] For example, the predetermined set of bitstream portion types can comprise one or more IRAP NAL unit types and / or GDR NAL unit types. For example, the predetermined set of bitstream portion types can comprise NAL unit types of IDR NUT, CRA NUT or GDR NUT.
[0145] For example, the extractor 30 can determine, for each access unit within the multi-layer video bitstream that does not have a sequence start indication, or alternatively, for each access unit of the sub-bitstream 12, whether all bitstream portions of the respective access unit are bitstream portions of the same predetermined bitstream portion type of a set of predetermined bitstream portion types. For example, the extractor 30 can parse to the associated information within the respective access unit or the multi-layer video bitstream to determine whether all bitstream portions of the respective access unit are bitstream portions of the same predetermined bitstream portion type of a set of predetermined bitstream portion types.
[0146] In other words, in one embodiment, the bitstream extraction process removes unnecessary layers and adds the AUD NUT to such use when these unnecessary layers are not present in the AU that needs the AUD after extraction.
[0147] According to an embodiment, the extractor 30 is configured to infer, for a predetermined OLS, from an indication within the multi-layer video bitstream 14 that one of the access units is a start access unit of a sub-sequence of the encoded video sequence represented by the predetermined OLS and to provide a sequence start indication within the sub-bitstream indicating that one access unit is the start access unit.
[0148] In other words, in another embodiment, such an indication is present in the bitstream at such an AU, i.e. an AU that becomes a CVSS AU, when extracting layers, such an AU becomes a CVSS AU or an IRAP AU, so that the insertion (addition) of an AUD becomes simpler and less parsing is needed.
[0149] According to another embodiment, the extractor 30 is configured to extract, for a predetermined OLS, nesting information, e.g. a nested SEI, indicating one or more access units, e.g. access units that are not the start access unit of the multi-layer video bitstream 14, are start access units of the OLS. According to this embodiment, the extractor 30 provides a sequence start indication within the sub-bitstream indicating that the one or more access units indicated within the nesting information are start access units. For example, the extractor 30 can provide a sequence start indication as described before for each of the indicated access units. Alternatively, the apparatus can provide a common indication for the indicated access units within the sub-bitstream.
[0150] In other words, in another embodiment, there is a nested SEI that can encapsulate AUDs, such that when a particular OLS is extracted that turns an AU into an IRAP / GDR AU, the encapsulated AUDs are unencapsulated and added to the bitstream. Currently the specification only includes nesting of other SEI messages. Therefore, there is a need for non-VCL non-SEI payloads to be allowed into the nested SEI. One way is to extend the existing nested SEI and indicate that non-SEI is included. Another way is to add a new SEI within the SEI that contains other non-VCL payloads.
[0151] The first option of implementation is shown in Table 2:
[0152] Table 2
[0153]
[0154]
[0155] The nested SEI encapsulates other SEI (sei_message()) and a given number (nesting_num_nonVclNuts_minus1) of non-VCL NAL units that are not SEI. This encapsulated non-VCL NAL units will be written into the nested SEI (nonVclNut) preceded by their length (length_minus1) such that their boundaries within the nested SEI can be found.
[0156] Table 3 shows another option, option 2:
[0157] Table 3
[0158]
[0159]
[0160] In option 2, a type can be added that also allows other non-VCL nal units, such that if in the future other non-VCL NAL units need to be included into the nested SEI, the nonVclNutPayload SEI message can be used.
[0161] In this case, a new SEI message is defined that directly includes a single non-VCL NAL unit, in this case, an access unit delimiter (AUD_rbsp()), and therefore this encapsulating SEI can be directly added to the nested SEI without any changes to the nested SEI (the nested SEI itself already includes other SEI).
[0162] According to further embodiments of the fourth aspect, the extractor 30 is configured for providing, for each access unit of the sub-bitstream, a sequence start indication indicating that the respective access unit is a start access unit of a sub-sequence of coded video sequences, if all bitstream parts of the respective access unit are bitstream parts of the same predetermined bitstream part type of a set of predetermined bitstream part types and the respective access unit comprises bitstream parts of two or more layers.
[0163] In another embodiment, AUDs are also required for access units that can become IRAP AUs or GDR AUs in case of OLS extraction (layer dropping), so that the extractor can ensure that they are present when needed and easily overwrite aud_irap_or_gdr_au_flag to 1 (i.e. the AU becomes an IRAP or GDR AU after extraction) when appropriate. One way to express this constraint in the specification is to extend the current text that AUDs must be present for IRAP or GDR AUs of the current bitstream:
[0164] There can be at most one EOB NAL unit in an AU, and there shall be one and only one AUD NAL unit in each IRAP or GDR AU when vps_max_layers_minus1 is greater than 0.
[0165] This would change to the following:
[0166] There can be at most one EOB NAL unit in an AU, and there shall be one and only one AUD NAL unit in each AU containing at least two layers with only IRAP or GDR NAL units when vps_max_layers_minus1 is greater than 0.
[0167] In other words, instead of using all NAL units the wording IRAP or GDR picture is used:
[0168] There can be at most one EOB NAL unit in an AU, and there shall be one and only one AUD NAL unit in each AU containing at least two IRAP or GDR pictures.
[0169] Accordingly, further embodiments according to the fourth aspect comprise an encoder 10 for providing a multi-layer video bitstream, e.g. with reference to Figure 1The encoder 10 described. Embodiments of the encoder 10 according to the fourth aspect are configured for providing a multi-layer video bitstream representing an encoded video sequence, e.g. the video sequence 20. The multi-layer video bitstream 14 provided by the encoder 10 comprises a sequence of access units 22, each access unit comprising one or more pictures, wherein each picture is associated with one of a plurality of layers 24 of the video bitstream. The encoder 10 is configured for providing, for each access unit of the multi-layer video bitstream 14, within a sub-bitstream 12, a sequence start indicator indicating whether all pictures of the respective access unit are pictures of a predetermined bitstream portion type of a predetermined set of picture types, if the respective access unit comprises at least two pictures of the predetermined bitstream portion type of the predetermined set of picture types. Optionally, the encoder 10 is configured for providing an output layer set (OLS) indication 18 in the multi-layer video bitstream 14, the OLS comprising one or more layers of the multi-layer video bitstream 14, and the encoder 10 provides the sequence start indicator for the respective access unit, if the respective access unit comprises at least two pictures of the predetermined bitstream portion type of the predetermined set of access unit types.
[0170] For example, the picture types comprise one or more IRAP picture types and / or GDR picture types. A picture belonging to one of the picture types can indicate that each of the one or more bitstream portions the picture is encoded into is of a predetermined bitstream portion type of a predetermined set of bitstream portion types, e.g. one or more IRAP NAL unit types or GDR NAL units.
[0171] For example, the sequence start indicator referred to throughout the embodiments of section 4 can be provided in the form of or as part of an access unit delimiter (AUD), the access unit delimiter can be a bitstream portion provided within the respective access unit. For example, an indication that an access unit is a sequence start access unit can be indicated by setting a flag of the AUD, e.g. aud_irap_or_gdr_au_flag, e.g. to a value of 1, to indicate that the respective access unit is a sequence start access unit.
[0172] Alternatively, the encoder 10 can not necessarily provide the sequence start indicator for each access unit comprising at least two pictures of the predetermined bitstream portion type of the predetermined set of picture types, but the encoder 10 can provide the sequence start indicator for each of the access units of the multi-layer video bitstream 14 comprising at least two pictures of the predetermined bitstream portion type of the predetermined set of picture types, each of the pictures belonging to one of the layers of an OLS indicated by the encoder in the OLS indication provided by the encoder 10 in the multi-layer video bitstream 14.
[0173] As an alternative implementation to the above embodiment, according to which AUDs are also required for access units that can become IRAP AUs or GDR AUs in case of OLS extraction (layer dropping), so that the extractor can be sure that they are present when needed and easily overwrite aud_irap_or_gdr_au_flag to 1 (i.e. the AU becomes an IRAP or GDR AU after extraction) if such access units would correspond to CVSS AUs in at least one OLS with more than one layer, the encoder would only need to write AUD NAL units. An example specification is shown below:
[0174] There can be at most one EOB NAL unit in an AU, and when vps_max_layers_minus1 is greater than 0, there shall be one and only one AUD NAL unit in each AU containing only IRAP or GDR NAL units in all layers of at least one multi-layer OLS.
[0175] In other words, the term IRAP or GDR picture is used:
[0176] There can be at most one EOB NAL unit in an AU, and when vps_max_layers_minus1 is greater than 0, there shall be one and only one AUD NAL unit in each AU containing only IRAP or GDR pictures in all layers of at least one multi-layer OLS.
[0177] In this embodiment, the related OLS extraction process would be extended by the following steps:
[0178] […]
[0179] The output sub-bitstream OutBitstream is derived as follows:
[0180] - Set the bitstream outBitstream to be the same as the bitstream inBitstream.
[0181] - Remove from outBitstream all NAL units with Temporalld greater than tIdTarget.
[0182] - Remove from outBitstream all NAL units with nal_unit_type not equal to any of VPS_NUT, DCI_NUT and EOB_NUT and nuh_layer_id not included in the list LayerIdInOls[targetOlsIdx].
[0183] When an AU contains only NAL units of a single type, nal_unit_type equal to IDR_NUT, CRA_NUT, or GDR_NUT, in two or more layers, the flag aud_irap_or_gdr_au_flag is rewritten to be equal to 1 for the AUD of the AU.
[0184] Therefore, the extractor 30 according to the fourth aspect can provide a sequence start indication by setting the value of the sequence start indicator for the corresponding access unit 22*. For example, the sequence start indicator can be a syntax element signaled in the access unit of the multi-layer video bitstream 14, and the extractor 30 can modify or maintain the value of the sequence start indicator when forwarding the access unit in the sub-bitstream 12.
[0185] In other words, according to an embodiment, the extractor 30 according to the fourth aspect may, for each of the access units 22 of the sub-bitstream, if all bitstream parts of the corresponding access unit (e.g., access unit 22*) are bitstream parts of the same type in a predetermined set of bitstream part types, provide a sequence start indication within the sub-bitstream 12 by setting the value of a sequence start indicator (e.g., aud_irap_or_gdr_flag) present in the corresponding access unit of the multi-layer video bitstream 14 (e.g., present in the AUD NAL unit of the corresponding access unit 22*) to a predetermined value (e.g., 1), indicating that the corresponding access unit is a starting access unit of a subsequence of the coded video sequence, wherein the predetermined value indicates that the corresponding access unit is a starting access unit of the subsequence of the coded video sequence. For example, if the sequence start indicator does not have the predetermined value in the multi-layer video bitstream 14, the extractor 30 may change the value of the sequence start indicator to the predetermined value.
[0186] Therefore, an embodiment of the encoder 10 for providing a multi-layer video bitstream 14 according to the fourth aspect provides an OLS indication 18 in the multi-layer video bitstream 14, the OLS indication 18 indicating the OLS of the layers (i.e., at least two layers) comprising the video bitstream 14. For each access unit of pictures of one of the predetermined picture types (e.g., the same type or not necessarily the same type) of the layer comprising the OLS, the encoder 10 may provide a sequence start indicator in the multi-layer video bitstream 14, the sequence start indicator indicating, for example by means of the value of the sequence start indicator, whether all pictures of the access unit (i.e., including pictures that are not part of the OLS) are of one of the predetermined picture types.
[0187] In other words, the encoder 10 can signal a sequence start indicator for an access unit (e.g., access unit 22) whose pictures (pictures belong to one of the layers of the OLS) are of one of the predetermined types.
[0188] 5. Handling of temporal sub-layers in the extraction process of an output layer set
[0189] Part 5 describes embodiments according to a fifth aspect of the application. Embodiments according to the fifth aspect can optionally be implemented according to the encoder 10 and extractor 30 embodiments described with reference to the other aspects. Figure 1 Furthermore, details described with reference to the other aspects can optionally be implemented in the embodiments described in this part.
[0190] Some embodiments according to the fifth aspect relate to the extraction process of OLS and vps_ptl_max_temporal_id[i][j]. Some embodiments according to the fifth aspect can relate to the derivation of NumSublayerInLayer[i][j].
[0191] In order to extract an output layer set (OLS), it is necessary to discard or remove from the bitstream the layers that do not belong to the OLS. However, note that the layers that belong to the OLS can have different number of sub-layers (time layers TLx) in Figure 10
[0192] Figure 10 An example of a two-layer bitstream is shown, where each of the two layers has a different frame rate. Figure 10 The bitstream of Fig. 2 includes an access unit 221 associated with a first temporal layer TL0 and an access unit 222 associated with a second temporal layer TL1. The first layer 241 includes pictures of both temporal sub-layers TL0, TL1, while the second layer 242 includes only pictures of TL0. Thus, the first layer 241 has twice the frame rate or picture rate of the second layer 242.
[0193] Figure 10 The bitstream of Fig. 2 can have two OLS: one composed of L0 only, which will have two operation points (e.g., 30fps TL0 and 60fps TL0+TL1). The other OLS can be composed of L0 and L1, but only have TL0.
[0194] The current specification allows to signal the profile and level of an OLS with TL0 and TL1 for L0, but the extraction process fails to generate a bitstream with only TL0.
[0195] Currently, when vps_max_tid_il_ref_pics_plus1[m][k] does not exist or layer j is the output layer in the i-th OLS, NumSublayerInLayer[i][j] indicating the maximum sublayer included in the i-th OLS for layer j is set to vps_max_sublayers_minus1+1.
[0196] Figure 11 An encoder 10 and an extractor 30 according to an embodiment of the fifth aspect are shown. Figure 11 The encoder 10 and extractor 30 may optionally correspond to a reference Figure 1 The encoder 10 and extractor 30 described above are described. Figure 1 The description can also optionally apply to Figure 11 The components shown. Figure 11 The encoder 10 is configured to encode a coded video sequence 20 into a multi-layer video bitstream 14. The multi-layer video bitstream 14 comprises access units 22, e.g. Figure 11 (or also Figure 1 ), each of the access units 22 includes one or more pictures 26 of the coded video sequence. For example, each of the access units 22 includes the same reference picture as the reference picture. Figure 1 One or more pictures 26 are related to a common time instance or frame of the coded video sequence described. Each of the pictures 26 belongs to one of the layers 24 of the multi-layer video bitstream. For example, in Figure 11 In the illustrative example of , the multi-layer video bitstream 14 includes a first layer 241 and a second layer 242, the first layer includes a picture 261, and the second layer includes a picture 262. According to the fifth aspect, each of the access units 22 belongs to a temporal sub-layer of the set of temporal sub-layers TL0, TL1 of the coded video sequence 20. For example, Figure 11 The access units 221 and 223 may belong to the first temporal sub-layer TL0, and the access unit 222 may belong to the second temporal sub-layer TL1, for example, as shown in FIG. Figure 10 The temporal sublayers may also be referred to as temporal subsets or temporal layers. For example, each temporal sublayer is indicated or indexed by a temporal identifier and may be characterized, for example, by a frame rate and / or a temporal relationship relative to other temporal subsets. For example, each of the access units may include or may have associated therewith a temporal identifier that associates the corresponding access unit with one of the temporal sublayers.
[0197] According to a fifth aspect, the encoder 10 is configured for providing, to the multi-layer video bitstream 14, a syntax element, e.g. the above-mentioned max_tid_within_ols, indicating a predetermined temporal sub-layer of an OLS, the OLS comprising or indicating a (not necessarily true) subset of layers of the multi-layer video bitstream 14. The syntax element indicates the predetermined temporal sub-layer of the OLS in a manner that distinguishes between different states, the different states including a state according to which the predetermined temporal sub-layer is below a maximum temporal sub-layer among the temporal sub-layers within an access unit to which pictures of at least one layer of the subset of layers belong. For example, the predetermined temporal sub-layer is the maximum temporal sub-layer comprised in the OLS.
[0198] For example, the encoder 10 can provide, within the multi-layer video data stream 14, an OLS indication 18, e.g. as described with reference to Figure 1 The OLS indication 18 can comprise a description or indication of one or more OLSs, e.g. the OLS 181 as shown in Figure 11 Each of the OLSs can be associated with a set of layers of the multi-layer video data stream 14. In Figure 11 the example shown, the OLS 181 is associated with the first layer 241 and the second layer 242. Note that the multi-layer video bitstream 14 can optionally comprise further layers.
[0199] For example, for the illustrative example of Figure 11 the OLS 181 can comprise the first temporal sub-layer (to which the access units 221 and 223 can belong), but the second temporal sub-layer to which the access unit 222 can belong in the example can not be comprised in the OLS, such that, according to the example, TL0 can be the maximum temporal sub-layer comprised in the OLS. Note that the temporal sub-layers can have a hierarchical order. In Figure 11 the example, the maximum value of the temporal sub-layers within the access unit 22 is TL1, and a picture of at least one layer of the subset of layers (the layer 241, which is part of the OLS), if it is the access unit 221, is the picture 261, is this maximum value. Thus, the maximum temporal sub-layer comprised in the OLS is below the maximum temporal sub-layer TL1. Thus, the predetermined temporal sub-layer can be indicated, for example, by indicating that the predetermined temporal sub-layer is below the maximum temporal sub-layer present in the access units comprised at least partially in the OLS. In another example, the predetermined temporal sub-layer can be indicated by indicating an index identifying the predetermined temporal sub-layer.
[0200] Thus, in the example of the OLS 181, the predetermined temporal sub-layer may be the first temporal sub-layer. A syntax element (e.g., max_tid_within_ols or vps_ptl_max_temporal_id) provided in the multi-layer video bitstream 14 indicates the predetermined temporal sub-layer of the OLS. The syntax element distinguishes different states. According to one of these states, the predetermined temporal sub-layer is below the maximum value among the temporal sub-layers within the access unit, and the picture of at least one layer in the layer subset is at the maximum value. For example, in Figure 11 In the example of OLS 181 described above, layers 241 and 242 are included. The largest temporal sub-layer within the access unit of the layer subset of OLS 181 is the second temporal sub-layer to which access unit 222 belongs. In the example where the second temporal sub-layer does not belong to OLS 181, a syntax element may indicate this state.
[0201] The extractor 30 according to the fifth aspect can derive syntax elements from the multi-layer video bitstream 14, and if a picture of the multi-layer video bitstream 14 belongs to one of the layers of the OLS 181, and if the picture belongs to an access unit 221, 223 (the access unit 221, 223 belongs to a temporal sublayer equal to or smaller than a predetermined temporal sublayer), a sub-bitstream 12 can be provided by forwarding the picture of the multi-layer video bitstream 14 in the sub-bitstream 12.
[0202] That is, if the picture belongs to a temporal sub-layer equal to or smaller than a predetermined temporal sub-layer, the extractor 30 may provide the bitstream portion of the corresponding picture 26 in the sub-bitstream 12, otherwise the picture may be discarded (ie, not forwarded).
[0203] In other words, extractor 30 may use syntax elements in the construction of sub-bitstream 12 to exclude pictures belonging to temporal sub-layers that are not part of the OLS to be decoded but are part of one of the layers indicated by the OLS from being forwarded in sub-bitstream 12 .
[0204] According to an embodiment, for each layer in the OLS, the multi-layer video bitstream 14 indicates a syntax element indicating a predetermined temporal sub-layer, such as a maximum temporal sub-layer, included in the corresponding OLS. The extractor 30 can distinguish between a bitstream portion belonging to a temporal sub-layer of the OLS and a bitstream portion belonging to a temporal sub-layer that is not part of the OLS based on the syntax element of the layer of the OLS, and consider forwarding these bitstream portions to the sub-bitstream 12 belonging to the temporal sub-layer of the OLS.
[0205] For example, the syntax element may be part of an OLS indication 18, for example, the syntax element may be part of the OLS 181 to which it relates. For example, the syntax element may be part of a video parameter set of the corresponding OLS.
[0206] In one embodiment, signaling (e.g. of the maximum temporal sublayer) is provided into the bitstream to indicate that the OLS has a maximum sublayer different from vps_max_sublayers_minus1 + 1 or different from the maximum layer among all layers present in the OLS. For this purpose, the existing syntax element vps_ptl_max_temporal_id[i][j] can be re-used to indicate the maximum sublayer present in the OLS.
[0207] Furthermore, according to some embodiments, NumSublayerInLayer[i][j] representing the maximum sublayer included in the i-th OLS of layer j is changed to vps_ptl_max_temporal_id[i][j] when vps_max_tid_il_ref_pics_plus[m][k] is not present or layer j is an output layer in the i-th OLS.
[0208] Alternatively, a new syntax element indicating the maximum sublayer within the OLS can be added, e.g. max_tid_within_ols.
[0209] According to embodiments, the encoder 10 and / or the extractor 30 are configured for deriving decoder capability related parameters for a substream (e.g. substream 12), the substream being obtained by selectively taking over respective pictures if the pictures belong to an OLS 181 and if the pictures belong to an access unit whose temporal sublayer is equal to or smaller than a predetermined temporal sublayer. In other words, the encoder 10 and / or the extractor 30 can derive decoder capability related parameters for a sub-bitstream, the sub-bitstream exclusively comprising pictures belonging to temporal sublayers of an OLS describing the sub-bitstream which are equal to or smaller than a predetermined temporal sublayer. The encoder 30 or the extractor 30 can signal the capability related parameters in the sub-bitstream 12. Thus, the encoder 10 can signal the capability related parameters in the multi-layer video bitstream 14. For example, the decoder capability related parameters can comprise the parameters as described in section 6.
[0210] 6. Handling of temporal sub-layers in video parameter signaling
[0211] Section 6 refers to Figure 11 and Figure 1 Embodiments according to the sixth aspect of the present application are described. Thus, Figure 1 and 11 The description of the above-mentioned aspects and embodiments can optionally apply to embodiments according to the sixth aspect. Furthermore, details described with respect to the other aspects can optionally be implemented in the embodiments described in this section.
[0212] Some examples according to the sixth aspect relate to constraints on vps_ptl_max_temporal_id[i], vps_dpb_max_temporal_id[i], vps_hrd_max_tid[i] to be consistent with a given OLS.
[0213] With regard to Figure 1 and / or Figure 11 The described multi-layer video bitstream 14 and / or sub-bitstream 12 can optionally comprise a video parameter set 81. The video parameter set 81 can comprise one or more decoder requirement sets (e.g., profile-tier-level sets (PTL sets)) and / or one or more buffer requirement sets (e.g., DPB parameter sets) and / or one or more bitstream conformance sets (e.g., hypothetical reference decoder (HRD) parameter sets). For example, each of the OLSs indicated in the OLS indication 18 can be associated with each of the decoder requirement sets, buffer requirement sets, and bitstream conformance sets that are applicable to the bitstreams described by the respective OLS. For each of the decoder requirement sets, buffer requirement sets, and bitstream conformance sets, the video parameter set can indicate the maximum temporal sub-layer that the respective set is concerned with, i.e., the maximum temporal sub-layer of the video bitstream or video sequence that the respective set is concerned with.
[0214] Figure 12 An example of a video parameter set 81 comprising a first decoder requirement set 821, a first buffer requirement set 841, and a first bitstream conformance set 861 associated with a first OLS 1 of the OLS indication 18 is shown. Furthermore, according to Figure 12 , the video parameter set 81 comprises a second decoder requirement set 822, a second buffer requirement set 842, and a second bitstream conformance set 862 associated with a second output layer set OLS 2.
[0215] For example, each of the OLSs described by the OLS indication 18 can be associated with one of the decoder requirement set 82, the buffer requirement set 84 and the bitstream conformance set 86 by associating respective indices pointing to the decoder requirement set, the buffer requirement set and the bitstream conformance set with the respective OLS. According to embodiments of the sixth aspect, the multi-layer video bitstream comprises access units, each access unit belonging to one of the temporal sub-layers of a temporal sub-layer set of a coded video sequence coded into the multi-layer video bitstream 14. The multi-layer video bitstream 14 according to the sixth aspect further comprises a video parameter set 81 and an OLS indication 18. For each of the bitstream conformance set 86, the buffer requirement set 84 and the decoder requirement set 82, a temporal subset indication indicates a constraint on a maximum temporal sub-layer (e.g. a maximum temporal sub-layer the respective bitstream conformance set / buffer requirement set / decoder requirement set is related to). For example, each of the bitstream conformance set 86, the buffer requirement set 84 and the decoder requirement set 82 signals a syntax element indicating the respective temporal subset indication (e.g. vps_ptl_max_temporal_id for the PTL set, vps_dpb_max_temporal_id for the DPB parameter set and vps_hrd_max_tid for the bitstream conformance set).
[0216] As Figure 12 illustrated, each of the bitstream conformance set 86, the buffer requirement set 84 and the decoder requirement set 82 can comprise a set of one or more parameters for each temporal sub-layer present in a layer set comprising layers of the bitstream portion of the video bitstream the respective bitstream conformance set 86, buffer requirement set 84 or decoder requirement set 82 is related to. For example, in Figure 12 , the OLS 1 comprises the layer L0 comprising the bitstream portion of the temporal layer TL0 and the OLS 2 comprises the layers L0 and L1 comprising the bitstream portions of the temporal layers TL0 and TL1. The bitstream conformance set 862 and the decoder requirement set 822 associated with the OLS 2 comprise parameters for L0, thus, according to this example, comprise parameters for the temporal layer TL0 and further comprise parameters for L1, thus also comprising parameters for the temporal layer TL1. The buffer requirement set 842 comprises a set of parameters for the DPB0 of the temporal sub-layer TL0 and for the DPB1 of the temporal sub-layer TL1.
[0217] Conventionally, there are three syntax structures present in the VPS, which are typically defined and subsequently mapped to a specific OLS:
[0218] • a profile-tier-level set (PTL), e.g. one or more decoder requirement sets
[0219] • DPB parameters, e.g. one or more buffer requirement sets
[0220] • HRD parameters, such as one or more bitstream conformance sets
[0221] The mapping of PTL to OLS is done in the VPS for all OLSs (with single layer or with multiple layers). However, the mapping of DPB parameters and HRD parameters to OLSs is done in the VPS only for OLSs with more than one layer. As shown in Figure 12 , the parameters of PTL, DPB and HRD are first described in the VPS and then the OLSs are mapped to indicate which of them use which parameters.
[0222] In the example shown in Figure 12 , there are 2 OLSs and 2 representations of each of these parameters. However, the definitions and mappings have been specified to allow more than one OLS to share the same parameters, so there is no need to repeat the same information multiple times, for example Figure 13 as shown.
[0223] Figure 13 An example is shown in which OLS2 and OLS3 have the same PTL and DPB parameters but different HRD parameters.
[0224] In the example of Figure 12 and Figure 13 , the values of vps_ptl_max_temporal_id[i], vps_dpb_max_temporal_id[i], vps_hrd_max_tid[i] for a given OLS are aligned (TL0 for OLS 1, TL1 for OLS2, TL1 for OLS2), but there is currently no necessity to do so. These three values associated with the same OLS are not currently limited to have the same value. Currently, none of these values are constrained in any way to be consistent or match the number of sub-layers in the bitstream. For example, in the example above, the bitstream can have a single sub-layer for OLS2 and OLS3, although these values are defined for two sub-layers. Therefore, when the matching becomes more complex, the decoder will not easily find out what the characteristics of the bitstream are.
[0225] In a first embodiment, the bitstream signals the maximum number of sub-layers present in the OLS (not necessarily in the bitstream, as some can have been discarded), but at least it can be understood as an upper limit, i.e. for an OLS, there cannot be more sub-layers in the bitstream than the signaled value (e.g. vps_ptl_max_temporal_id[i]). Therefore, the DPB parameters and HRD parameters are also used by the decoder.
[0226] If the values of vps_dpb_max_temporal_id[i], vps_hrd_max_tid[i] are different from vps_ptl_max_temporal_id[i], the decoder will need to perform a more complex mapping. Therefore, in one embodiment, there is a bitstream constraint that if the OLS index the PTL structure, the DPB structure and the HRD parameter structure with vps_ptl_max_temporal_id[i], vps_dpb_max_temporal_id[i], vps_hrd_max_tid[i] respectively, vps_dpb_max_temporal_id[i], vps_hrd_max_tid[i] should be equal to vps_ptl_max_temporal_id[i].
[0227] According to embodiments, the encoder 10 (e.g., Figure 1 or Figure 11 ) is configured for forming the OLS indication 18 such that the maximum temporal sublayer indicated by the bitstream conformance set 86, the buffer requirement set 84 and the decoder requirement set 82 associated with the OLS are equal to each other and the parameters within the bitstream conformance set 86, the buffer requirement set 84 and the decoder requirement set 82 are fully valid for the OLS.
[0228] However, looking at the example in Figure 12 , the parameters for OLS1 with a single sublayer (TL0), i.e. level 0 (L0 in PTL0 (TL0)) DPB parameter 0 (DPB0 in DPB 0) and HRD parameter 0 (HRD0 in HRD 0) are also described in PTL 1 822, DPB1 842 and HRD 1 862. In order not to repeat so many parameters, one option is not to include DPB 0 841 and HRD 0 861 and to take the values for OLS1 from the DPB parameters and HRD parameters including more sublayers. Figure 14 An example is shown in
[0229] Figure 14 An example of PTL, DPB and HRD definitions and sharing of sublayer information independent of some OLS between different OLS is shown. Since PTL0 indicates that there is only one sublayer, only the parameters for TL0 in DPB1 and HRD1 are used.
[0230] Accordingly, in another embodiment, there is a bitstream constraint, i.e. if the OLS is indexed with vps_ptl_max_temporal_id[i], vps_dpb_max_temporal_id[i], vps_hrd_max_tid[i] for the PTL structure, the DPB structure and the HRD parameter structure, respectively, vps_dpb_max_temporal_id[i], vps_hrd_max_tid[i] shall be greater than or equal to vps_ptl_max_temporal_id[i], and the greater values corresponding to higher sub-layers of the DPB parameters and the HRD parameters are ignored for the OLS.
[0231] Accordingly, according to another embodiment, the encoder 10 is configured for forming the OLS indication 18 and / or the video parameter set 81 (or, generally, the multi-layer video bitstream 14) such that the maximum temporal sub-layer indicated by the decoder requirement set 82 associated with the OLS is less than or equal to the maximum temporal sub-layer indicated by each of the buffer requirement set 84 and the bitstream conformance set 86 associated with the OLS, and the parameters within the buffer requirement set 84 and the bitstream conformance set 86 are only valid for the OLS as long as the buffer requirement set 84 and the bitstream conformance set 86 relate to temporal sub-layers equal to or smaller than the maximum temporal sub-layer indicated by the decoder requirement set 82 associated with the OLS.
[0232] In other words, the encoder 10 can provide the OLS indication 18 and / or the video parameter set 81 such that the maximum temporal sub-layer indicated by the buffer requirement set 84 associated with the OLS is greater than or equal to the maximum temporal sub-layer indicated by the decoder requirement set 82 associated with the OLS, and such that the maximum temporal sub-layer indicated by the bitstream conformance set 86 associated with the OLS is greater than or equal to the maximum temporal sub-layer indicated by the decoder requirement set 82 associated with the OLS.
[0233] For example, Figure 14An example of a video parameter set 81 is shown, which includes a first decoder requirement set 821 for a first layer set (e.g., layer L0, which includes access units of a first temporal sub-layer TL0). The video parameter set 81 also includes a second decoder requirement set 822 that relates to a second layer set, which includes layer L0 and layer L1, and the second set includes access units of a first temporal sub-layer and a second temporal sub-layer (i.e., TL0 and TL1). Thus, the maximum temporal sub-layer of the second layer set is the second temporal sub-layer TL1. The video parameter set 81 also includes a DPB parameter set 842 that relates to the second layer set, a bitstream conformance set 862 that relates to the second layer set, and a bitstream conformance set 863 that relates to a third layer set, which includes access units of the first temporal sub-layer and access units of a third layer L2 having the second temporal sub-layer. A first OLS (OLS1) is associated with the first layer set, and the decoder requirements of the first layer set are described by the first decoder requirement set 821 that indicates that the maximum temporal sub-layer of the first layer set is the first temporal sub-layer. Since the first temporal sub-layer is less than or equal to (note that temporal sub-layers are hierarchically ordered) the maximum temporal sub-layer indicated by the DPB parameter set 842 associated with OLS1 and the bitstream conformance set 862 associated with OLS1, the DPB parameter set 842 and the bitstream conformance set 862 include information about the first layer set. Thus, as long as the DPB parameter set 842 and the bitstream conformance set 862 relate to the first layer set, they are valid for OLS1, and the access units of the first layer set belong to the first temporal sub-layer. For example, as described with reference to Figure 12 Figure 13 Figure 14 and shown, the decoder requirement set 822, the buffer requirement set 842, and the bitstream conformance set 862 include the parameter sets for each temporal sub-layer included in the bitstream that they relate to. According to this embodiment, the parameters related to the first temporal sub-layer TL0 are valid for OLS1 because the first temporal sub-layer is equal to or less than the maximum temporal sub-layer indicated by the decoder requirement set 822.
[0234] In other words, for an OLS to be decoded, the decoder 50 can use those (and in examples only those) parameters of the decoder requirement set 82, the buffer requirement set 84, and the bitstream conformance set 86 associated with that OLS that relate to temporal sub-layers that are equal to or less than the maximum temporal sub-layer associated with the decoder requirement set 82 of that OLS.
[0235] Figure 15 Another example of a video parameter set 81 and OLS indication 18 is shown. Figure 15 An alternative that can occur when the parameters of TL1 and TL0 are the same or when the value of TL1 is the maximum value allowed for the level is shown. In this case, instead of not including DPB0 and HRD0 in the VPS as shown previously, both DPB0 and HRD0 can be included without including DPB1 and HRD1. Then, the value of the higher sub-layer of OLS1 can be derived as equal to the value signaled for TL0 or the maximum value allowed for the level. Thus, Figure 15 An example of PTL, DPB and HRD definitions and sharing between different OLSs is shown, where sub-layer information needs to be inferred when it is not present for some OLSs.
[0236] Thus, in another embodiment, there is no bitstream constraint on the values vps_ptl_max_temporal_id[i], vps_dpb_max_temporal_id[i], vps_hrd_max_tid[i] but for values greater than vps_dpb_max_temporal_id[i], vps_hrd_max_tid[i], vps_ptl_max_temporal_id[i], the DPB and HRD parameters for i > vps_dpb_max_temporal_id[i] up to vps_hrd_max_tid[i] of vps_ptl_max_temporal_id[i] should be inferred to be the maximum value specified by the profile level or equal to the highest signaled DPB parameter and HRD parameter.
[0237] Thus, according to another embodiment, the encoder 10 is configured for forming the OLS indication and / or the video parameter set 81 such that the maximum temporal sub-layer indicated by the decoder requirement set 82 associated with the OLS is greater than or equal to the maximum temporal sub-layer indicated by each of the buffer requirement set 84 and the bitstream conformance set 86 associated with the OLS. According to these embodiments, the parameters related to temporal sub-layers above the maximum temporal sub-layer indicated by each of the buffer requirement set 84 and the bitstream conformance set 86 associated with the OLS (e.g., OLS2) are missing within the buffer requirement set 84 and the bitstream conformance set 86 associated with the OLS and are set equal to the fourth parameter or equal to the parameters within the buffer requirement set 84 and the bitstream conformance set 86 associated with the OLS related to the maximum temporal sub-layer indicated by each of the buffer requirement set 84 and the bitstream conformance set 86. Figure 15
[0238] Thus, a decoder for decoding a multi-layer video bitstream (e.g., the decoder 20) is configured for: Figure 1 Embodiments of the decoder 50) can be configured to make an inference that parameters of the buffer requirement set 84 and the bitstream conformance set 86 associated with the OLS related to temporal sub-layers above the maximum temporal sub-layer indicated by each of the buffer requirement set 84 and the bitstream conformance set 86 will be set, for each parameter of the buffer requirement set 84 and the bitstream conformance set 86, to equal a default value, e.g., a maximum value of the respective parameter indicated in the decoder requirement set 82, or a value of the respective parameter within the buffer requirement set 84 or the bitstream conformance set 86 associated with the OLS that is related to the maximum temporal sub-layer indicated by each of the buffer requirement set 84 and the bitstream conformance set 86, if the maximum temporal sub-layer indicated by the decoder requirement set 82 associated with the OLS (i.e., the OLS to be decoded) is greater than or equal to the maximum temporal sub-layer indicated by each of the buffer requirement set 84 and the bitstream conformance set 86 associated with the OLS. For example, the selection of whether to use a default value or a value of the respective parameter within the buffer requirement set or the bitstream conformance set related to the maximum temporal sub-layer indicated by each of the buffer requirement set and the bitstream conformance set can be different for each parameter of the buffer requirement set 84 and the bitstream conformance set 86.
[0239] 7. Output layer selection in region of interest application
[0240] Section 7 refers to Figure 1 Embodiments according to the seventh aspect are described. Thus, Figure 1 The description of the decoder 50) can optionally apply to embodiments according to the seventh aspect. Furthermore, details described with respect to other aspects can optionally be implemented in embodiments described in this section.
[0241] Some embodiments according to the seventh aspect relate to PicOutputFlag derivation in RoI applications.
[0242] When using multi-layer bitstreams (e.g., video bitstream 14) and pictures of the specified output layer are not available at the decoder side (e.g., bitstream error or transmission loss), a suboptimal user experience can result without adhering to certain considerations. Typically, when an access unit does not contain pictures in the output layer, it is up to the implementation to select pictures from non-output layers for output to compensate for the error / loss, as evident from the derivation of the PicOutputFlag variable from the following note:
[0243] The variable PictureOutputFlag for the current picture is derived as follows:
[0244] - If sps_video_parameter_set_id is greater than 0 and the current layer is not an output layer (i.e., nuh layer id is not equal to OutputLayerldlnOls[TargetOlsldx][i] for any value of i in the range of 0 to NumOutputLayerslnOls[TargetOlsldx] - 1, inclusive), or one of the following conditions is true, PictureOutputFlag is set equal to 0:
[0245] - The current picture is a RASL picture and the associated IRAP picture has NoOutputBeforeRecoveryFlag equal to 1.
[0246] - The current picture is a GDR picture with NoOutputBeforeRecoveryFlag equal to 1 or a recovery picture of a GDR picture with NoOutputBeforeRecoveryFlag equal to 1.
[0247] - Otherwise, PictureOutputFlag is set equal to ph_pic_output_flag.
[0248] Note - In implementations, a decoder can output pictures that do not belong to an output layer. For example, when there is only one output layer available and the pictures of the output layer are not available in an AU, e.g., due to loss or layer down-switching, the decoder can set PictureOutputFlag equal to 1 for the picture with the highest value of nuh layer id among all pictures of the AU available to the decoder and with ph_pic_output_flag equal to 1, and set PictureOutputFlag equal to 0 for all other pictures of the AU available to the decoder.
[0249] However, when the bitstream is produced for application to a region of interest (RoI), it is not desirable to change between layers in the decoder output on a short time frame, i.e., a higher layer (via the use of a scaling window) depicts only a subset of the lower layer picture, as this would result in fast switching between overview and detail. Therefore, as part of the present invention, in one embodiment, when a scaling window is used that does not cover the entire picture plane, the decoder is not allowed to freely choose the output layer, as follows:
[0250] In implementations, the decoder can output pictures that do not belong to the output layer as long as the scaling window covers the complete picture plane. For example, when only one output layer is present, e.g. due to loss or layer down-switching, and pictures of the output layer are not available in an AU, the decoder can set PictureOutputFlag equal to 1 for the picture with the highest value of nuh layer id among all pictures of the AU available to the decoder and having ph_pic_output_flag equal to 1, and set PictureOutputFlag equal to 0 for all other pictures of the AU available to the decoder.
[0251] According to embodiments of the seventh aspect, a decoder 50 for decoding a multi-layer video bitstream (e.g. the multi-layer video bitstream 14 or a sub-bitstream 12) is configured for using, according to the relative size and relative position of the scaling windows of the prediction pictures and the reference pictures defined in the multi-layer video bitstream 14, a scaled and offset prediction vector for a vector-based inter-layer prediction of a prediction picture 262 of a second layer 241 from a reference picture 261 of a first layer 242. For example, the picture 262 of the layer 242 can be encoded into the multi-layer video data stream 14 using inter-layer prediction from a picture 261 of the layer 241 (e.g. a picture 261 of the same access unit 221). According to the seventh aspect, the multi-layer video bitstream 14 can comprise an OLS indicating a subset of layers of the multi-layer video bitstream 14, the OLS comprising one or more output layers, including the first layer 241, and one or more non-output layers, including the second layer. Figure 1
[0252] In case of a missing predetermined picture (e.g. the picture 262) of the first layer 242 of the OLS, the decoder 50 according to the seventh aspect is configured for replacing the predetermined picture 262 by another predetermined picture of the second layer 241 of the OLS in case the scaling window defined for the predetermined picture 262 coincides with the picture boundary of the predetermined picture and the scaling window defined for the other predetermined picture in the same access unit 22 as the predetermined picture 262 coincides with the picture boundary of the other predetermined picture. In case at least one of the scaling pictures defined for the predetermined picture does not coincide with the picture boundary of the predetermined picture and the scaling window defined for the predetermined picture does not coincide with the picture boundary of the other predetermined picture, the decoder 50 is configured for replacing the predetermined picture by other means or not at all.
[0253] 8. Other embodiments
[0254] In the foregoing sections, although some aspects have been described as features in the context of an apparatus, it should be clear that such description can also be seen as a description of corresponding features of a method. Although some aspects have been described as features in the context of a method, it should be clear that such description can also be seen as a description of corresponding features of a function on an apparatus.
[0255] Some or all of the method steps can be executed by (or using) a hardware apparatus, like for example, a microprocessor, a programmable computer or electronic circuit.
[0256] The inventive encoded image signal can be stored on a digital storage medium or can be transmitted on a transmission medium such as a wireless transmission medium or a wired transmission medium (e.g. the Internet).
[0257] Depending on certain implementation requirements, embodiments of the application can be implemented in hardware or in software, or in a combination of hardware and software. The implementation can be performed using a digital storage medium, for example a floppy disk, a DVD, a Blu-Ray, a CD, a ROM, a PROM, an EPROM, an EEPROM or a FLASH memory, having electronically readable control signals stored thereon, which cooperate (or are capable of cooperating) with a programmable computer system such that the respective method is performed. Therefore, the digital storage medium can be computer readable.
[0258] Some embodiments according to the application comprise a data carrier having electronically readable control signals, which are capable of cooperating with a programmable computer system, such that one of the methods described herein is performed.
[0259] Generally, embodiments of the present application can be implemented as a computer program product, having a program code, the program code being operative for performing one of the methods when the computer program product runs on a computer. The program code may, for example, be stored on a machine readable carrier.
[0260] A further embodiment comprises a computer having installed thereon the computer program for performing one of the methods described herein.
[0261] In other words, an embodiment of the inventive method is, therefore, a computer program having a program code for performing one of the methods, when the computer program runs on a computer.
[0262] A further embodiment of the inventive methods is, therefore, a data carrier (or a digital storage medium, or a computer-readable medium) comprising, recorded thereon, the computer program for performing one of the methods described herein. The data carrier, the digital storage medium or the recorded medium are typically tangible and / or non-transitionary. Thus, an embodiment of the inventive method is, therefore, a data stream or a sequence of signals representing the computer program for performing one of the methods described herein. The data stream or the sequence of signals can, for example, be for communication about the inventive method over a computer network.
[0263] A further embodiment of the inventive method is, therefore, a data stream or a sequence of signals representing the computer program for performing any of the methods described herein. The data stream or the sequence of signals can for example be configured to cause a corresponding apparatus, for example a computer, to execute the method. The data stream or the sequence of signals can for example be configured to cause a corresponding apparatus, for example a computer, to execute the method.
[0264] A further embodiment comprises an apparatus configured to or adapted for performing any of the methods described herein.
[0265] A further embodiment comprises a computer having installed thereon the computer program for performing any of the methods described herein.
[0266] A further embodiment comprises an apparatus or a system configured to transfer (for example, electronically or optically) a computer program for performing any of the methods described herein to a receiver. The receiver may, for example, be a computer, a mobile device, a memory device or the like. The apparatus or system may, for example, comprise a file server for transferring the computer program to the receiver.
[0267] In some embodiments, a programmable logic device (for example a field programmable gate array) can be used to perform some or all of the functionalities of the methods described herein. In some embodiments, a field programmable gate array can cooperate with a microprocessor in order to perform one of the methods described herein. Generally, the methods are preferably performed by any hardware apparatus.
[0268] The apparatuses described herein can be implemented using a hardware apparatus, or using a computer, or using a combination of hardware and computer.
[0269] The methods described herein can be performed using a hardware apparatus, or using a computer, or using a combination of hardware and computer.
[0270] In the detailed description above, it can be seen that various features are grouped together in examples for the purpose of streamlining the disclosure. This method of disclosure should not be interpreted as reflecting an intention that the claimed examples require more features than are explicitly recited in each claim. Rather, inventive subject matter can lie in less than all features of a particular disclosed example. Thus, the following claims are hereby incorporated into the detailed description, where each claim can stand as a separate example. While each claim can stand as a separate example, the examples can be combined in any way deemed to be desirable. Thus, the claims can be combined in any way to form a new example, even if the combination is not expressly disclosed in the detailed description. Further, any one feature of an example can be used in combination with any other feature of another example, even if that feature is not expressly disclosed in the other example. Thus, the claims are to be affixed to the examples in this manner so that any element of any claim can be used in combination with any element of any other claim, even if that element is not expressly disclosed in the other claim.
[0271] The above examples are merely illustrative of the principles of the disclosure. Numerous modifications and adaptations will be apparent to those skilled in the art in view of the examples described herein. Therefore, the above description should not be construed as limiting but merely as illustrative of the present disclosure.
Claims
1. A decoder (50) for decoding a video bitstream (12, 14), the decoder (50) comprising: processor; as well as A memory stores instructions, which, when executed by the processor, cause the processor to perform the following operations: obtaining a video bitstream (12, 14), the video bitstream (12, 14) comprising access units (22), each of the access units (22) comprising one or more pictures (26), and each of the one or more pictures corresponding to one layer of a plurality of layers (24) of the video bitstream (14), deriving one or more output layer sets (OLS) (181, 182) from the video bitstream (14), each OLS indicating a set of layers from the plurality of layers of the video bitstream (14), identifying a first access unit of said access units (22), for the first access unit, when external information for identifying the OLS to be decoded among the one or more OLSs is unavailable, identifying the OLS to be decoded based on the number of layers of each of the one or more OLSs, and The identified OLS is decoded.
2. The decoder (50) according to claim 1, wherein To determine the OLS to decode, the instructions, when executed by the processor, cause the processor to identify a highest value or a lowest value beyond the index of each of the OLSs.
3. The decoder (50) according to claim 1, wherein To determine the OLS to decode, the instructions, when executed by the processor, cause the processor to identify a maximum number of layers of the OLS.
4. The decoder (50) according to claim 1, wherein The external information is external to the video bitstream.
5. The decoder (50) according to claim 1, wherein When executed by the processor, the instructions further cause the processor to: obtaining the extrinsic information separated from the video bitstream; and The OLS to be decoded is identified based on the external information.
6. The decoder (50) according to claim 1, wherein When executed by the processor, the instructions further cause the processor to: determining whether to obtain additional information indicating an operating point; identifying the operating point for decoding in response to obtaining the additional information; as well as In response to not obtaining the additional information, selecting the OLS to be decoded.
7. The decoder (50) according to claim 1, wherein The first access unit is at the beginning of one of the coded video sequences (20) and represents at least one of the first in time order, the first in decoding order, or the first in output order.
8. The decoder (50) according to claim 1, wherein The first access unit is a sequence start access unit type.
9. The decoder (50) according to claim 8, wherein The first access unit is at the beginning of the video bitstream.
10. A method for decoding a video bitstream, the method comprising: Obtaining a video bitstream (12, 14), the video bitstream (12, 14) comprising access units (22), each of the access units (22) comprising one or more pictures (26), and each of the one or more pictures corresponding to one layer of a plurality of layers (24) of the video bitstream (14); deriving one or more output layer sets (OLS) (181, 182) from the video bitstream (14), each OLS indicating a set of layers from the plurality of layers of the video bitstream (14); identifying a first access unit among the access units (22); for the first access unit, when external information for identifying the OLS to be decoded among the one or more OLSs is unavailable, identifying the OLS to be decoded based on the number of layers of each of the one or more OLSs, and The identified OLS is decoded.
11. The method according to claim 10, wherein: Determining the OLS to decode includes identifying a highest value or a lowest value beyond an index of each of the OLSs.
12. The method according to claim 10, wherein: Determining the OLS to be decoded includes identifying a maximum number of layers of the OLS.
13. The method according to claim 10, wherein: Determining the OLS to be decoded is performed based on at least one of: an index of each of the OLSs, a number of output layers of each of the OLSs, or a number of layers of each of the OLSs.
14. The method according to claim 10, further comprising: obtaining the external information separated from the video bitstream; as well as The OLS to be decoded is identified based on the external information.
15. The method according to claim 10, further comprising: determining whether to obtain additional information indicating an operating point; identifying the operating point for decoding in response to obtaining the additional information; as well as In response to not obtaining the additional information, selecting the OLS to be decoded.
16. The method according to claim 10, wherein The first access unit is at the beginning of one of the coded video sequences (20) and represents at least one of the first in time order, the first in decoding order, or the first in output order.
17. The method according to claim 10, wherein The first access unit is a sequence start access unit type.
18. The method according to claim 17, wherein The first access unit is at the beginning of the video bitstream.
19. A non-transitory computer-readable medium storing a computer program which, when executed by a processor of a decoder, causes the decoder to: Obtaining a video bitstream (12, 14), the video bitstream (12, 14) comprising access units (22), each of the access units (22) comprising one or more pictures (26), and each of the one or more pictures corresponding to one layer of a plurality of layers (24) of the video bitstream (14); deriving one or more output layer sets (OLS) (181, 182) from the video bitstream (14), each OLS indicating a set of layers from the plurality of layers of the video bitstream (14); identifying a first access unit among the access units (22); for the first access unit, when external information for identifying the OLS to be decoded among the one or more OLSs is unavailable, identifying the OLS to be decoded based on the number of layers of each of the one or more OLSs, and The identified OLS is decoded.
Citation Information
Patent Citations
Profile, tier, level for the 0-th output layer set in video coding
US20150373361A1