Video decoder, video decoding method, computer-readable storage medium, and computer program
The solution for extracting sub-bitstreams from multi-layer video bitstreams addresses the challenges of efficient extraction and reduced signaling overhead by using output layer sets and syntax elements to identify bitstream portions, enhancing decoder resource utilization and decoding speed.
Patent Information
- Application Number
- JP2025115995
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2020-05-22
- Filing Date
- 2025-07-09
- Publication Date
- 2025-09-25
- Estimated Expiration
- 2041-05-20
AI Technical Summary
Existing video decoding technologies face challenges in efficiently extracting sub-bitstreams from multi-layer video bitstreams while minimizing signaling overhead and ensuring precise definition of extractable sub-bitstreams, decoder resource utilization, and avoiding unnecessary decoding of non-essential parts.
A concept for representing and decoding sub-bitstreams from multi-layer video bitstreams, where the extracted sub-bitstreams include bitstream portions associated with an output layer set, indicated by a reference layer, and utilize syntax elements to identify bitstream portions by their type and inter-layer dependencies, allowing accurate extraction with reduced signaling overhead.
Enables efficient extraction of sub-bitstreams with improved decoder resource utilization and reduced signaling overhead, ensuring that decoders can select appropriate output layer sets without explicit instructions, and provides more sequence start access units for faster decoding.
Smart Images

Figure 2025138881000001_ABST
Abstract
Description
[Technical Field]
[0001] explanation Embodiments of the present invention relate to apparatus for encoding video into a video bitstream, apparatus for decoding a video bitstream, and apparatus for processing a video bitstream, e.g., extracting a bitstream such as a sub-bitstream from a video bitstream. Further embodiments relate to methods for encoding, decoding, and processing (e.g., extracting) video bitstreams. Further embodiments relate to video bitstreams. [Background technology]
[0002] The emerging VVC codec is expected to support hierarchical coding for temporal, fidelity, and spatial scalability from the outset. That is, the coded video bitstream is structured into so-called layers and (temporal) sublayers, and the coded picture data corresponding to time, i.e., so-called access units (AUs), can contain, within each layer, pictures that can be predicted from one another, some of which will be output after decoding. The concept of a so-called output layer set (OLS) indicates to the decoder the reference relationships and which layers will be output when the bitstream is decoded. The OLS can also be used to identify corresponding HRD-related timing / buffer information in the form of SEI messages for buffering period, picture timing, and decode unit information, which are carried in the bitstream encapsulated in so-called scalable nesting SEI messages. Summary of the Invention
[0003] It is desirable to have a concept for processing output layer sets that enables extraction of sub-bitstreams from a video bitstream, which concept provides an improved trade-off between precise definition of extractable sub-bitstreams by the output layer set (in terms of describing exactly which parts of the video bitstream are extracted), efficient utilization of decoder resources (e.g., in terms of avoiding extraction of parts unnecessary for decoding the selected sub-bitstream, or in terms of providing precise information about decoder settings or requirements for decoding the selected sub-bitstream), and low signaling overhead.
[0004] A first aspect according to the present invention provides a concept for representing, extracting, and / or decoding a randomly accessible sub-bitstream from a multi-layer video bitstream, wherein the extracted randomly accessible sub-bitstream selectively includes bitstream portions of an access unit of the multi-layer video bitstream that are associated with an output layer of the randomly accessible sub-bitstream as indicated by an output layer set indication of the randomly accessible sub-bitstream, or bitstream portions required to decode the randomly accessible bitstream portion of the output layer.
[0005] A second aspect of the present invention provides the concept of a multi-layer video bitstream having multiple layers and multiple temporal layers. The multi-layer video bitstream comprises an indication of an output layer set including one or more layers of the multi-layer video bitstream and a reference layer indication indicating inter-layer references for layers of the output layer set. The multi-layer video bitstream includes an indication, e.g., a temporal layer indication or an intra-temporal layer indication, that, in combination with the way the multi-layer video bitstream is encoded, enables identifying bitstream portions of layers of the output layer set that belong to the output layer set. This concept enables identifying bitstream portions of an OLS by their type and / or by inter-layer dependencies of the OLS indicated by the reference layer indication. Thus, embodiments of the second aspect enable accurate extraction of sub-bitstreams while avoiding unnecessarily high signaling overhead.
[0006] A third aspect of the present invention provides a concept that enables a decoder that decodes a video bitstream to determine an output layer set to decode based on attributes of the video bitstream provided to the decoder. Thus, this concept allows the decoder to select an OLS without the decoder being instructed which OLS to decode. A decoder that can select an OLS in the absence of instructions can ensure that the bitstream decoded by the decoder meets level requirements that are known to the decoder, for example, by an indication in the video bitstream.
[0007] A fourth aspect of the present invention provides a concept for extracting a sub-bitstream from a multi-layer video datastream, wherein, within the extracted sub-bitstream, access units constituting only one, e.g., the same, picture or bitstream portion of a set of predetermined bitstream portion types or picture types (e.g., randomly accessible or independently coded bitstream portion types or picture types) are indicated by a sequence start indicator, even if the respective access units are not sequence start access units in the original multi-layer video bitstream from which the sub-bitstream was extracted. Thus, the frequency of sequence start access units in the sub-bitstream may be higher than in the multi-layer video datastream, and accordingly, a decoder may benefit from having more sequence start access units available, thereby avoiding unnecessarily long wait times before it can start decoding a video sequence.
[0008] A fifth aspect of the present invention provides a concept for extracting sub-bitstreams from a multi-layer video bitstream such that the sub-bitstreams consist only of pictures belonging to one or more temporal sublayers associated with the output layer set that describes the extracted sub-bitstream. To this end, syntax elements in the multi-layer video bitstream are used to indicate a given temporal sublayer of an OLS in a manner that identifies different states, including a state in which the given temporal sublayer is below the maximum of the temporal sublayers in an access unit that has at least one picture of a subset of layers. Avoiding transmission of unnecessary sublayers of a multi-layer video bitstream can reduce the size of the sub-bitstream and potentially reduce the decoder requirements for decoding the sub-bitstream.
[0009] According to an embodiment, decoder capability related parameters of a sub-bitstream that exclusively comprises pictures of a temporal sub-layer that belong to an OLS describing the sub-bitstream are signaled in the sub-bitstream and / or the multi-layer video data stream, so that when determining the decoder-related capability parameters, pictures that do not belong to the OLS can be omitted, thereby making it possible to efficiently utilize the decoder capabilities.
[0010] A sixth aspect of the present invention provides a concept for processing temporal sublayers in signaling video parameters for an output layer set of a multi-layer video bitstream. According to an embodiment, an OLS is associated with one of one or more bitstream conformance sets, one of one or more buffer requirement sets, and one of one or more decoder requirement sets signaled in the video bitstream, and each of the bitstream conformance set, buffer requirement set, and decoder requirement set is valid for one or more temporal sublayers indicated by a constraint of maximum temporal sublayers (e.g., hierarchically ordered temporal sublayers). The embodiment provides a concept of a relationship between the bitstream conformance set, buffer requirement set, and decoder requirement set associated with an OLS in terms of the maximum temporal sublayer with which they are associated, thereby enabling a decoder to easily determine the parameters of the OLS associated with the bitstream conformance set, buffer requirement set, and decoder requirement set. For example, the embodiment may enable a decoder to conclude that the parameters given in the bitstream conformance set, buffer requirement set, and decoder requirement set of the OLS are fully valid for the OLS. Another embodiment allows the decoder to conclude how effective the parameters given in the bitstream conformance set, buffer requirement set, and decoder requirement set are for OLS.
[0011] According to an embodiment, an OLS is only valid for as long as the maximum temporal sublayer indicated by the decoder requirement set associated with the OLS is smaller than or equal to the maximum temporal sublayer indicated by each of the buffer requirement set and bitstream adaptation set associated with the OLS, and the parameters in the buffer requirement set and bitstream adaptation set are the same for the same or lower temporal layer as the maximum temporal sublayer indicated by the decoder requirement set associated with the OLS. As a result, if the maximum temporal sublayer indicated by the decoder requirement set associated with the OLS is smaller than or equal to the maximum temporal sublayer indicated by each of the buffer requirement set and bitstream adaptation set associated with the OLS, the decoder can infer that the parameters of the buffer requirement set and bitstream adaptation set associated with the OLS are equal to the maximum temporal sublayer indicated by the decoder requirement set associated with the OLS and are valid for the OLS only as long as they relate to the lower temporal layer. Thus, an embodiment may enable a decoder to determine the video parameters of the OLS based on the indication regarding the maximum temporal sublayer constraint signaled for each parameter set, thereby avoiding complex analysis of the OLS and the video parameters. Furthermore, since the concept allows the association of an OLS with a buffer requirement set and a bitstream adaptation set, a maximum temporal sublayer constraint can be signaled that is greater than the constraint of the decoder requirement set associated with the OLS, and since that maximum temporal sublayer constraint is greater than the constraint of the decoder requirement set associated with the OLS, signaling of dedicated buffer requirement sets and dedicated bitstream adaptation sets associated with the same maximum temporal sublayer as the decoder requirement set can be omitted, thereby reducing the overhead of signaling video parameter sets.
[0012] A seventh aspect of the present invention provides a concept for handling picture loss in a multi-layer video bitstream, for example, due to a bitstream error or transmission loss, where pictures are encoded using inter-layer prediction. If a picture that is part of a first layer is lost, the picture is replaced with another picture in a second layer, and the picture in the second layer is used for inter-layer prediction of the picture in the first layer. This concept includes replacing a picture with a further picture based on a match between a scaling window defined for the picture and a picture boundary of the picture, and a match between a scaling window defined for the further picture and a picture boundary of the further picture. If a scaling window defined for the picture and a picture boundary of the picture match, and if a scaling window defined for the picture and a picture boundary of the further picture match, replacing the picture with the further picture may not result in a change in the display window of the presented content, for example, a change from a detailed view to an overview. Further embodiments and advantageous implementations of the present disclosure are explained in more detail below with reference to the figures. [Brief explanation of the drawings]
[0013] [Figure 1] FIG. 1 shows an example of an encoder, extractor, decoder, and multi-layer video bitstream according to an embodiment. [Figure 2] FIG. 2 shows an example of an output layer set of a randomly accessible sub-bitstream. [Figure 3] FIG. 3 shows an example of an extracted randomly accessible sub-bitstream with unused pictures. [Figure 4] FIG. 4 shows an example of a three-layer bitstream that aligns independently coded pictures across three layers. [Figure 5] FIG. 5 shows an example of a three-layer bitstream with unaligned independently coded pictures. [Figure 6] FIG. 6 shows an example of a four-layer bitstream in which pictures of an access unit containing independently coded pictures include pictures that reference different temporal sub-layers. [Figure 7] FIG. 7 illustrates an example of a decoder according to an embodiment. [Figure 8] FIG. 8 shows an example of a multi-layer video data stream having access units with randomly accessible and non-randomly accessible pictures. [Figure 9] FIG. 9 shows an example of a sub-bitstream of the multi-layer video data stream of FIG. [Figure 10] FIG. 10 shows an example of a multi-layer video bitstream having two layers with different picture rates. [Figure 11] FIG. 11 shows an encoder, extractor, multi-layer video bitstream and sub-bitstreams according to an embodiment. [Figure 12] FIG. 12 shows an example of video parameter sets and their mapping to output layer sets. [Figure 13] FIG. 13 shows an example of sharing video parameters between different output layer sets. [Figure 14] FIG. 14 illustrates an example of sharing video parameters between different OLSs according to an embodiment. [Figure 15] FIG. 15 illustrates sharing of video parameters between different OLSs according to another embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0014] Although the following describes embodiments in detail, it should be understood that the embodiments provide many applicable concepts that can be embodied in a wide variety of video coding concepts. The specific embodiments described are merely illustrative of specific ways to implement and use the concepts and do not limit the scope of the embodiments. In the following description, numerous details are set forth to provide a more thorough description of embodiments of the present invention. However, it will be apparent to one skilled in the art that other embodiments can be practiced without these specific details. In other instances, well-known structures and devices are shown in block diagram form, rather than in detail, in order to avoid obscuring the examples described herein. Furthermore, features of different embodiments described herein can be combined with each other unless otherwise noted.
[0015] In the following description of the embodiments, identical or similar elements or elements having the same functions are given the same reference numerals or identified by the same names, and repeated descriptions of elements given the same reference numerals or identified by the same names are generally omitted. Therefore, the descriptions provided for elements having the same or similar reference numerals or identified by the same names can be mutually interchangeable or applied to each other in different embodiments.
[0016] 0. Encoder 10, extractor 30, decoder 50 and video bitstreams 12, 14 according to FIG. The embodiments described in this section provide examples of frameworks into which embodiments of the present invention may be incorporated. Below, a description of embodiments of the present invention's concepts is presented, along with a description of how such concepts may be incorporated into the encoder and extractor of FIG. 1. However, the embodiments described with respect to FIG. 2 and subsequent ones may be used to form encoders and extractors that do not operate according to the framework described with respect to FIG. 1. Furthermore, it should be noted that the encoder, extractor, and decoder, although illustrated together in FIG. 1 for illustrative purposes, may be implemented separately from one another. It should be further noted that the extractor and decoder may be combined within a single device, or one of the two may be implemented as part of the other.
[0017] FIG. 1 illustrates an example of an encoder 10, an extractor 30, a decoder 50, a video bitstream 14 (also referred to as a video data stream or data stream), and a sub-bitstream 12. The encoder 10 is for encoding a video sequence 20 into the video bitstream 14. The encoder 10 encodes the video sequence 20 into the video bitstream 14 in units of pictures 26, where each picture 26 belongs to a time instant, e.g., a frame of the video sequence. Encoded video data belonging to a common time instant may be referred to as an access unit (AU) 22. FIG. 1 illustrates exemplary operations of three access units 221, 222, and 223 of the video sequence 20. Note that any description referring to an access unit 22 may refer to any of the exemplary access units 221, 222, and 223. Each of the access units 22 includes or encodes one or more pictures 26, each of which is associated with one of multiple layers of the video bitstream 14. In FIG. 1 , examples of pictures 26 are represented by picture 261 and picture 262. Picture 261 is associated with a first layer 241 of video bitstream 14, and picture 262 is associated with a second layer 242 of video bitstream 14. Note that in FIG. 1 , each of access units 22 includes a picture for both first layer 241 and second layer 242, but video bitstream 14 may include access units 22 that do not necessarily include a picture 26 for each of layers 24 of video bitstream 14. Furthermore, video bitstream 14 may include additional layers in addition to the first and second layers shown in FIG. 1 . Encoder 10 is configured to encode each of pictures 26 into one or more bitstream portions, e.g., NAL units, of video sequence 14. For example, each bitstream portion 16 into which picture 26 is encoded may encode a portion of picture 26, such as a slice of picture 26.The bitstream portions 16 in which pictures 26 are actually encoded are sometimes referred to as video coding layer (VCL) NAL units. The bitstream 14 may further include description data, e.g., non-VCL NAL units, indicating information describing the encoded video data. For example, each of the access units 22 may include a bitstream portion signaling description data for the respective access unit in addition to a bitstream portion signaling the decoded video data. The video bitstream 14 may further include description data that refers to multiple access units or portions of one or more access units. For example, the video bitstream 14 may encode an output layer set (OLS) indication 18 that indicates one or more output layer sets.
[0018] An OLS may be a representation of a sub-bitstream extractable from video bitstream 14. An OLS may indicate one or more or all of multiple layers of video bitstream 14 as output layers of the sub-bitstream described by the respective OLS. Note that the set of layers indicated by an OLS may not necessarily be a proper subset of the layers of video bitstream 14. In other words, all layers of video bitstream 14 may be included in the OLS. An OLS may optionally further include a description of the sub-bitstream described by the OLS and / or decoder requirements for decoding the sub-bitstream indicated by the OLS. Note that a sub-bitstream described by an OLS may be defined by additional parameters other than layers, such as temporal sublayers or subpictures. For example, picture 26 of layer 24 may be associated with one of one or more temporal sublayers of layer 24. A temporal sublayer may include a picture for a time instant associated with the respective temporal sublayer. For example, pictures of a first temporal sublayer may be associated with time instants that form a sequence with a first frame rate, and pictures of a second temporal sublayer may be associated with time instants located between the time instants to which the pictures of the first temporal sublayer are associated, such that a combination of the first and second temporal sublayers provides a video sequence with a higher frame rate than one of the first and second temporal sublayers. The OLS may optionally indicate a temporal sublayer to describe which bitstream portions or pictures 26 belong to the sub-bitstream described by the OLS. The temporal sublayers of a bitstream or coded video sequence may be hierarchically ordered, for example, by indexing. For example, a hierarchical order may mean that decoding a picture of a bitstream that includes a particular temporal sublayer requires all temporal sublayers lower in the hierarchical order.
[0019] It should be noted that the OLS may include one or more output layers and, optionally, one or more non-output layers. In other words, the OLS may indicate one or more of the layers included in the OLS as output layers of the OLS, and may optionally indicate one or more of the layers of the OLS as non-output layers. For example, a layer including a reference picture for a picture of an output layer of the OLS may be included in the OLS as a non-output layer, because a picture of the non-output layer may be required to decode a picture of the output layer of the OLS.
[0020] The OLS may further include level information for the bitstream described by the OLS, where the level information indicates or is associated with one or more bitstream constraints, such as a maximum value for one or more of the bitrate, picture size, and frame rate.
[0021] Optionally, the bitstream 14 may further include an extractability indication 19 of the OLS. For example, the extractability indication may be part of the OLS indication. The extractability indication may indicate a (not necessarily proper) subset of the bitstream portions 16 that form a decodable sub-bitstream associated with the OLS. That is, the extractability indication may indicate which of the bitstream portions 16 belong to the OLS.
[0022] Pictures 26 may be encoded into video bitstream 14 with reference to other pictures, e.g., for residual, motion vector, and / or syntax element prediction. For example, a picture may reference another picture in the same access unit (called the picture's reference picture), or a reference picture associated with another layer, which may be called an inter-layer reference picture. Additionally or alternatively, a picture may reference a reference picture that is part of the same layer but in a different access unit than the picture.
[0023] The extractor 30 may receive the video bitstream 14 and, for example, may select an OLS from one or more OLSs indicated in the video bitstream 14 based on an indication 32 provided to the extractor 30. The extractor 30 may provide the sub-bitstream 12 indicated by the selected OLS by forwarding at least the bitstream portions 16 belonging to the selected OLS in the sub-bitstream 12. Note that the extractor 30 may modify or adapt one or more of the bitstream portions 16, so the forwarded bitstream portions do not necessarily correspond exactly to the bitstream portions 16 signaled in the video bitstream 14. In FIG. 1, an apostrophe is used around a bitstream portion of the sub-bitstream 12, e.g., reference numeral 16′, to indicate a potential change in the bitstream portion of the video bitstream 14 when forwarded to the sub-bitstream 12.
[0024] Sub-bitstream 12 may be decoded by decoder 50 to obtain a decoded video sequence represented by sub-bitstream 12. It should be noted that besides the fact that the decoded video sequence may differ from video sequence 20 in terms of resolution, fidelity, frame rate, picture size and video content (when focusing on sub-picture extraction), in the sense that the decoded video sequence may optionally represent only a portion of video sequence 20, the decoded video sequence may have distortions due to quantization losses.
[0025] Pictures 26 of video sequence 20 may include independently coded pictures that do not reference pictures of other access units. That is, for example, independently coded pictures are encoded into video bitstream 14 without using inter prediction (although with respect to temporal prediction, independently coded pictures may optionally be encoded using inter-layer prediction). Independent coding allows a decoder to start decoding the video sequence at the independently coded picture's access unit. Independently coded pictures are sometimes referred to as instantaneous random access points (IRAP). Examples of IRAP pictures are IDR and CRA pictures. In contrast, a trailing (TRAIL) picture may refer to a picture that references a picture of another access unit that may precede the trailing picture encoding order (the order in which picture 26 is coded into video stream 14). The bitstream portion 16 into which an independently coded picture is encoded may be referred to as an independently coded bitstream portion, e.g., an IRAP NAL unit, while the bitstream portion 16 into which a dependently coded picture 26 is encoded may be referred to as a dependent bitstream portion, e.g., a non-IRAP NAL unit. Furthermore, it should be noted that not all bitstream portions of a picture are necessarily encoded in the same manner, including independent and dependent coding. For example, a first portion of picture 26 in a first access unit, e.g., access unit 221, may be independently coded, and a second portion of picture 26 in the first access unit may be dependently coded. In this case, for picture 26 in a second access unit, e.g., access unit 222, the first portion of picture 26 in the second access unit may be dependently coded, and the second portion may be independently coded. In this way, the higher data rate of independent coding relative to dependent coding may be distributed across multiple access units.Such encoding may be referred to as general decoder refresh (GDR) because decoder 50 may have to decode several access units before decoding an entire picture independent of the access unit preceding the GDR cycle, i.e., a sequence of pictures interspersed with independently coded portions covering the entire picture.
[0026] In the following, some concepts and embodiments will be described with reference to Fig. 1. It is pointed out that features described with reference to an encoder, a video bitstream, an extractor, or a decoder shall be understood to be descriptions of other of these entities as well. For example, a feature described as being present in a video data stream shall be understood as a description of an encoder configured to encode this feature into a video bitstream, and a decoder or extractor configured to read the feature from the video bitstream. It is further pointed out that the estimation of information based on the representation coded in the video bitstream can be performed equally on the encoder side and on the decoder side. It is further noted that the aspects described in the following sections can be combined with each other.
[0027] 1. Randomly accessible sub-bitstream display In this section, embodiments according to the first aspect are described with reference to Figure 1. Details described in Section 0 may optionally be applied to embodiments according to the first aspect. Also, details described with respect to further aspects may optionally be implemented in embodiments described in this section.
[0028] A randomly accessible bitstream portion may refer to an independently coded bitstream portion as described with respect to Figure 1. Accordingly, a randomly accessible picture (or bitstream portion) may be referred to as an independently coded picture (or bitstream portion) as described with respect to Figure 1.
[0029] Some embodiments according to the first aspect may refer to the full IRAP level indication for unaligned IRAPs. Embodiments may refer to the impact of IRAP alignment where max_tid_il_ref_pics_plus1==0 (e.g., references to IDRs only, or references only to either IDRs, CRAs, or GRDs with ph_recovery_poc_cnt equal to 0).
[0030] Figure 2 illustrates an example of an output layer set for a video sequence, such as video sequence 20. In other words, the video sequence of Figure 2 may represent a video sequence formed by layers of an OLS, including an output layer L1 and a non-output layer L0. In access unit 22*, the multi-layer OLS bitstream of Figure 2 includes an example of a non-aligned IRAP. That is, access unit 22* includes a non-randomly accessible bitstream portion in one of the layers of the OLS, i.e., L1 in Figure 2.
[0031] An indication of level information for all IRAP sub-bitstreams, i.e., a bitstream containing the result of dropping all non-IRAP NAL units from the bitstream, can be useful for trick mode playback, such as fast-forward based on IRAP pictures only. This level information references the level_idc indication, which points to a list of defined constraints on parameters such as maximum picture size, maximum picture rate, maximum bitrate, maximum buffer size, maximum slices / tiles / subpictures per picture, and minimum compression ratio. However, in multi-layered cases, it is not uncommon for IRAP pictures to be misaligned across layers. For example, IRAPs are less frequent in higher (dependent) layers than in lower (reference) layers due to the longer IRAP distance in higher (dependent) layers. This is illustrated in Figure 2. Here, the higher layer L1 of the illustrated OLS with two layers contains a trailing NAL unit at POC (Picture Order Count) 3, i.e., within access unit 22*, while the lower layer L0 contains an IRAP NAL unit at the same location.
[0032] Figure 3 shows an example of an extracted full IRAP sub-bitstream that can be extracted from the multi-layer video bitstream of Figure 2 and executed conventionally. In light of the fact that not all layers of an OLS are output by a decoder, for example, because only the upper layer L1 is marked as an output layer in the example of Figure 2, maintaining the lower (non-output) layer IDRs in the bitstream would be a waste of decoder resources unless there is an IDR in the corresponding position of the upper (output) layer, as in the case of picture 260*. This is because the decoder output is the same regardless of whether the L0 IDR of POC3 is present in the full IRAP sub-bitstream, so the decoder would not output any pictures when decoding the full IRAP sub-bitstream of POC3. Furthermore, if the full IRAP level indication (e.g., relative to the level and playback rate of the complete OLS bitstream) is used to estimate the maximum playback speed of the full IRAP presentation, decoding the above L0 IRAP of POC3 would reduce the maximum achievable playback speed of the full IRAP sub-bitstream.
[0033] Therefore, it is part of an embodiment of the present invention to omit decoding / dropping from such bitstreams, and thereby also exclude from consideration by the respective level indication all IRAP NAL units of non-output layers in access units that do not have IRAP NAL units in all corresponding output layers in the OLS of such full IRAP sub-bitstreams.
[0034] According to an embodiment of the first aspect, the video bitstream 14 represents an encoded video sequence 20, e.g., as described with respect to FIG. 1 . The video bitstream 14 includes a sequence of access units 22, each of which includes one or more bitstream portions 16, each of which is associated with one of the multiple layers 24 of the video bitstream 14. Each of the bitstream portions 16 is one of the bitstream portion types, including independently coded bitstream portion types such as a randomly accessible bitstream portion type, e.g., an IRAP type. The video bitstream 14 includes, for example, an OLS indication 18 of the OLS of the video bitstream 14 to be encoded by the encoder 10 and detected by the extractor 30, and an extractability indication 19 of the randomly accessible sub-bitstream described by the OLS, the OLS including one or more output layers and one or more non-output layers. For example, the randomly accessible sub-bitstream may be the entire IRAP sub-bitstream. For example, the OLS indication may include a level indication for the entire IRAP sub-bitstream. It should be noted that the term full IRAP sub-bitstream should not be understood to mean that the full IRAP sub-bitstream necessarily includes a randomly accessible bitstream portion or an IRAP bitstream portion exclusively. Rather, a randomly accessible sub-bitstream or a full IRAP sub-bitstream may optionally include a non-IRAP bitstream portion or a non-randomly accessible bitstream portion, for example, a bitstream portion of a reference picture of a randomly accessible bitstream portion. In other examples, a randomly accessible sub-bitstream may include only a randomly accessible bitstream portion.
[0035] According to an embodiment, the encoder 10 provides a video bitstream 14 such that, for each layer of the OLS, for each access unit 22 that exceeds the bitstream portion of the respective access unit 22, the bitstream portion of all output layers is a randomly accessible bitstream portion if the respective access unit 22 includes one of the randomly accessible bitstream portions.
[0036] In other words, in one embodiment, a requirement for bitstream conformance is that the bitstream indicated by the full IRAP level indication does not contain access units without output pictures in all output layers.
[0037] There are use cases where not all output layers have pictures in all access units of the original bitstream, such as stereo video where the frame rates differ for each eye. In such cases, the bitstream requirements are less strict.
[0038] In another embodiment, a requirement for bitstream conformance is that the bitstream indicated by the full IRAP level indication does not contain access units without output pictures in all output layers.
[0039] Thus, according to an embodiment, the encoder 10 according to the first aspect provides a video bitstream 14 such that, for each of the layers indicated by the OLS of the randomly accessible sub-bitstream, for each of the access units 22 beyond the bitstream portion of the respective access unit 22, the bitstream portion 16 of at least one of the output layers is a randomly accessible bitstream portion if the respective access unit includes one of the randomly accessible bitstream portions.
[0040] The consequence of both (randomly accessible bitstream portions of at least one or all output layers) is that access units that do not satisfy the bitstream constraints are either not created at the encoder side or are dropped during extraction. In other words, a bitstream containing a representation of an IRAP-only level is such that there are no AUs that have IRAPs in non-output layers and non-IRAP NAL units in temporally co-located pictures in the output layers.
[0041] Alternatively, the randomly accessible sub-bitstream described by OLS indication 18 and extractability indication 19 (which are encoded into video bitstream 14 by encoder 10) selectively includes, for each access unit, a respective bitstream portion if, for each bitstream portion of the respective access unit, one of the following two conditions is met: The first condition is met if the respective bitstream portion is a randomly accessible bitstream portion and if the respective bitstream portion is associated with one of one or more output layers, e.g., a bitstream portion of picture 26* in FIG. The second condition is met if each bitstream portion is associated with a reference layer of one of the output layers and each bitstream portion is associated with one of the one or more non-output layers, and further, beyond the bitstream portions of the respective access units, at least one bitstream portion of the output layers is a randomly accessible bitstream portion (which is the case for the bitstream portion of picture 26** in Figure 2, and the corresponding access unit includes the bitstream portion of picture 26* that is part of the output layer of the OLS), or alternatively, beyond the bitstream portions of the respective access units, the bitstream portions of all output layers are randomly accessible bitstream portions (which is also the case for picture 26** in Figure 2).
[0042] Thus, an apparatus for extracting a sub-bitstream 12 from a video bitstream 14, such as apparatus 30 of FIG. 1 according to a preceding embodiment of the first aspect, is configured to provide the sub-bitstream 12 as indicated by the OLS representation of the randomly accessible sub-bitstream.
[0043] In other words, as an optional alternative to the bitstream constraint that the bitstream indicated by the all IRAP level indication does not include access units without output pictures in all output layers, according to an embodiment, the indicated level does not include AUs in which this IRAP NAL unit is "mixed" with non-IRAP NAL units, and therefore such AUs need to be dropped if a bitstream having only IRAPs of the indicated level is desired.
[0044] A similar case is considered when an output layer has an IRAP NAL unit but a reference layer does not. In a further embodiment, as long as the output layer has an IRAP NAL unit, NAL units in the co-reference layer are not dropped and are considered for the indicated level. For this to work, there is a bitstream constraint that temporal references in co-located reference layers that do not have IRAP NAL units refer only to pictures that are also contemporaneous with the IRAP NAL units of the output layer. Alternatively, the indicated level applies only to AUs whose NAL units are IRAP NAL units; if such a bitstream at such a level (IRAP only) is considered, all others are discarded.
[0045] In a further embodiment, instead of referring to the layer selected for OLS, the requirement that only AUs with all NAL unit types of IRAP type are considered for IRAP level indication is applied to the entire bitstream. In such a case, such AUs contain access unit delimiters with aud_irap_or_gdr_au_flag equal to 1, so that when there are IRAP NAL units (i.e., the GDR case is not considered), the presence of an access unit delimiter with aud_irap_or_gdr_au_flag equal to 1 is used to determine whether the AU is subject to IRAP-only level indication.
[0046] According to an example of the first aspect, video bitstream 14 includes level indications of randomly accessible sub-bitstreams, e.g., randomly accessible sub-bitstreams that are described as extractable according to extractability information 19. The level indications (also referred to as level information) may indicate levels associated with bitstream constraints, as described in Section 0, e.g., by pointing to a list of levels. For example, the level indications are associated with one or more of CPD size, DPB size, picture size, picture rate, minimum compression ratio, picture splitting restrictions (e.g., tile / slice / subpicture), HRD timing (e.g., access unit / DU removal time, DPB output time, etc.).
[0047] In other words, in addition to the level indication, further parameters become relevant when an extracted bitstream with only IRAP access units is considered: such parameters are the DPB parameter and the HRD parameter.
[0048] According to an embodiment of the first aspect, a decoder such as the decoder 50 is configured to check whether a picture buffer conforms to the randomly accessible sub-bitstream according to the extractability information 19. For example, the decoder 50 may check the extractability information, the HRD parameter, and the level indication in the DPB parameter in addition to the above-mentioned parameters. The picture buffer may refer to the coded picture buffer and / or the decoded picture buffer of the decoder. Optionally, the decoder 50 may be configured to derive timing information for the picture buffer, for example, timing information for the randomly accessible sub-bitstream indicated by the OLS indication 18, from the video bitstream 12. The decoder 50 may decode the randomly accessible sub-bitstream based on the timing information.
[0049] In other words, in practice, all IRAP variants that omit IRAPs in non-output layers without IRAPs in output layers within the same access unit also omit decoding of non-output pictures that are not used for the aforementioned reference, thereby enabling a reduction in DPB requirements (i.e., DPB size in units of picture slots). Notably, this is a separate part of the bitstream's level restriction and is not directly related to the restriction defined by the bitstream's level_idc. Together with the bitstream's picture size, the level sets a limit on the maximum number of pictures that can be held in the DPB. However, DPB parameters also contain more information, such as the maximum reordering of pictures when outputting them, i.e., the number of pictures that can precede another picture in output order but follow it in output order. Such information may be different if the extracted bitstream contains only IRAP pictures. Therefore, signaling additional DPB parameters for this representation is part of the present invention, so that decoders can more efficiently utilize their resources. One embodiment of the present invention is shown in Table 1 below. [Table 1-1] [Table 1-2] In Table 1, level_indication_for_all_irap_present indicates the presence of level information for all IRAP representations excluding non-output IRAP pictures without an output layer IRAP in the respective access unit.
[0050] For example, vps_ols_dpb_params_all_irap_idx[i] specifies the index into the list of dpb_parameters() syntax structures in the VPS of the dpb_parameters() syntax structure that applies to the ith multi-layer OLS when only IRAP sub-bitstreams are considered. If present, the value of vps_ols_dpb_params_idx[i] shall be in the range from 0 to VpsNumDpbParams-1, inclusive. If vps_ols_dpb_params_all_irap_idx[i] is not present, it is inferred to be equal to vps_ols_dpb_params_idx[i]. For a single-layer OLS, the corresponding dpb_parameters() syntax structure is present in the SPS referenced by the layer in the OLS. Each dpb_parameters() syntax structure in the VPS shall be referenced by at least one value of vps_ols_dpb_params_idx[i] or vps_ols_dpb_params_all_irap_idx[i] for i in the range 0 to NumMultiLayerOlss-1 inclusive.
[0051] As noted, further additional information that may be needed for an extracted bitstream containing only IRAP NAL units are HRD parameters, which may include, for example, one or more or all of the required CPB size, the time at which access units are removed from the CPB, the bitrate at which the CPB is provided, or whether the resulting bitstream after extraction corresponds to a constant bitrate representation.
[0052] 2. Reference Picture Alignment Section 2 describes an embodiment according to the second aspect with reference to Figure 1, and details described in Section 0 may optionally be applied to embodiments according to the second aspect. Also, details described with respect to further aspects may optionally be implemented in embodiments described in this section.
[0053] In VVC, the output layer set defines the prediction dependency between layers of the bitstream. The syntax element vps_max_tid_il_ref_pics_plus1[i][j], signaled for all direct reference layers of a given layer, allows to further limit the amount of pictures of the reference layer used for prediction, as follows:
[0054] vps_max_tid_il_ref_pics_plus1[i][j] equal to 0 specifies that a picture in the jth layer that is neither an IRAP picture nor a GDR picture with ph_recovery_poc_cnt equal to 0 is not used as an ILRP to decode a picture in the ith layer. vps_max_tid_il_ref_pics_plus1[i][j] greater than 0 specifies that a picture from the jth layer whose TemporalId is greater than vps_max_tid_il_ref_pics_plus1[i][j]-1 is not used as an ILRP to decode a picture in the ith layer. If not present, the value of vps_max_tid_il_ref_pics_plus1[i][j] is inferred to be equal to vps_max_sublayers_minus1+1.
[0055] If not present, the value is inferred to vps_max_sublayer_minus1+1, where vps_max_sublayer_minus1 is the maximum number of sublayers present in any layer in the bitstream. Note, however, that certain layers may have a smaller value for the maximum number of sublayers.
[0056] This syntax element not only indicates that inter-layer referencing is not used for some sublayers, or that some sublayers of a reference layer are not needed for decoding, but also indicates a special mode (vps_max_tid_il_ref_pics_plus1[i][j] equal to 0) in which only IRAP NAL units or GDR NAL units with ph_recovery_poc_cnt equal to 0 are needed from a reference layer for decoding. Furthermore, the output layer set describing the bitstream passed to the decoder does not contain the unnecessary NAL units indicated by this syntax element vps_max_tid_il_ref_pics_plus1[i][j], or such NAL units are dropped in a particular decoder implementation that implements the extraction process defined in the specification.
[0057] The syntax element vps_max_tid_il_ref_pics_plus1[i][j] is only present in direct reference layers. For example, imagine an OLS with three layers as shown in Figure 4. Figure 4 shows three example layers L0, L1, and L2, where vps_max_tid_il_ref_pics_plus1 is equal to 0 for all reference layers. L2 has L1 as its direct reference layer, and L1 uses L0 as its direct reference layer, so L2 has L0 as its indirect reference layer. In such a case, if vps_max_tid_il_ref_pics_plus1[2][1] is equal to 0, only IRAP or GDR NAL units with ph_recovery_poc_cnt equal to 0 are retained from L1 and, consequently, from L0, as indicated in the specification. More specifically, for each layer in the OLS, a variable is derived indicating the number of sublayers retained: NumSubLayersInLayerInOLS[i][j] (where i is the OLS index and j is the layer index). If the layer is an output layer, the value of this variable is set to the maximum temporalId desired for the bitstream. If the layer is not an output layer, but a reference layer for each layer k in the OLS using layer j as a reference, the value of NumSubLayersInLayerInOLS[i][j] is set to the maximum of min(NumSubLayersInLayerInOLS[i][k],vps_max_tid_il_ref_pics_plus1[k][j]). That is, for each layer k, check what is the minimum between the number of sublayers required for layer k (NumSubLayersInLayerInOLS[i][k]) and the number of sublayers required for layer j (vps_max_tid_il_ref_pics_plus1[k][j]), taking all sublayers of layer k into account. If layer k requires fewer sublayers than indicated by vps_max_tid_il_ref_pics_plus1[k][j], then the minimum of the two is chosen, since layer j only requires the same amount of sublayers as layer k.Then further layers k that also use j as a reference are checked, and if other layers indicate that more sublayers are needed, a higher value is taken, i.e. the maximum value needed after all layers that use layer j as a reference have been checked.
[0058] A problem occurs when IRAP NAL units are not aligned between L0 and L1. For example, as shown in Figure 5, which shows a three-layer example with a non-aligned IRAP in the lower layer, imagine that there is an IRAP AU at some point in L1 but not in L0, and the IRAP AU in L1 uses the non-IRAP AU in L0 as a reference. In such a case, the IRAP-based extraction process would discard the non-IRAP in L0 (picture 260* in Figure 5), and therefore would not be able to decode the IRAP AU in L1 (picture 261 in Figure 5).
[0059] In an embodiment, when vps_max_tid_il_ref_pics_plus1[i][j] is 0 for layer i, it is required that any direct or indirect layer of such layer i has an aligned IRAP or GDR NAL unit with ph_recovery_poc_cnt equal to 0. In other words, for any indirect reference layer, the NAL units of its indirect reference layer whose synchronizing NAL units of its direct reference layer that depend on it are either IRAP or GDR NAL units with ph_recovery_poc_cnt equal to 0 must also be either IRAP or GDR NAL units with ph_recovery_poc_cnt equal to 0.
[0060] According to an embodiment of the second aspect, the video bitstream 14 includes a sequence of access units 22, each of which includes one or more bitstream portions 16. Each of the bitstream portions 16 is associated with one of a plurality of layers 24 of the video bitstream 14 and one of a plurality of temporal layers of the video bitstream, e.g., the temporal sub-layers described with reference to FIG. 1 . Bitstream portions 16 within the same access unit 22 are associated with the same temporal layer. Furthermore, each of the bitstream portions is one of a bitstream portion type including a set of predetermined bitstream portion types. For example, the set of predetermined bitstream portion types may include independently coded bitstream portion types, such as IDR, and, optionally, types that depend only on other access units of the set of predetermined bitstream portion types.
[0061] According to an embodiment of the second aspect, the encoder 10 is configured to provide, in the video bitstream 14, an OLS indication of the OLS of the video bitstream 14, where the OLS includes one or more layers of the video bitstream. Furthermore, the encoder 10 provides, in the video bitstream, for each layer of the OLS, a reference layer indication indicating a set of reference layers on which the respective layer depends. Furthermore, in the video bitstream 14, for each layer (e.g., i) of the OLS, for each reference layer (e.g., j) of the respective layer, the encoder 10 provides, in the video bitstream 14, a temporal layer indication (e.g., vps_max_tid_il_ref_pics_plus1[i][j]) indicating whether the entire bitstream portion of each reference layer on which the respective layer depends is one of a set of predetermined bitstream portion types, or if not, whether the respective layer is a bitstream portion up to which temporal layer the respective layer depends (e.g., a temporal layer indication indicating the bitstream portion up to the maximum index indexing the temporal layer on which the respective layer depends).
[0062] The encoder 10 according to this embodiment is configured to provide a video bitstream such that, for each layer (e.g., i) of the OLS for which the temporal layer indication indicates that all bitstream portions of a given reference layer (of the reference layer of the given layer) on which the respective layer depends are one (e.g., the same) of a set of given bitstream portion types, and an access unit including bitstream portions of the given reference layer that are one of the set of given bitstream portion types does not include bitstream portions other than the set of given bitstream portion types for each further reference layer on which the given reference layer directly or indirectly depends (e.g., direct dependency or direct reference is dependency between a (dependent) layer and its reference layer, e.g., as shown in the reference layer indication, and indirect dependency or reference is dependency between a (dependent) layer and a direct or indirect reference layer of a reference layer of a (dependent) layer that is not shown in the reference layer indication).
[0063] In practice, this is more restrictive than necessary: such indirect reference layers (L0) can also be reference layers of other layers, as in Figure 6, which shows an example of four layers directly referencing a sublayer with Tid1 (imagine the case of a fourth L3 layer, where sublayers 0 and 1 of L0 are needed, and vps_max_tid_il_ref_pics_plus1[3][0] is 2).
[0064] In such a case, the IRAP alignment constraint described in the previous embodiment would be unnecessary, since the non-IRAP NAL units of layer 0 required for the IRAP NAL units of L1 would be kept in the OLS bitstream corresponding to L0+L1+L2+L3. Therefore, to express the constraint, we can use the variable NumSubLayersInLayerInOLS[i][j] instead, which indicates the number of sublayers (with their respective temporal IDs) kept in the ith OLS for the jth layer (0 means that only IRAPs or GDRs with ph_recovery_poc_cnt equal to 0 are kept).
[0065] In an embodiment, within the i-th OLS, for two layers k and j, where k>j, when NumSubLayersInLayerInOLS[i][j] and NumSubLayersInLayerInOLS[i][k] are equal to 0, an IRAP NAL unit or a GDR with ph_recovery_poc_cnt equal to 0 is aligned if j is a reference layer of k (direct or indirect).
[0066] 4 and 5, the encoder 10 may provide the bitstream 14 for each layer (e.g., i) 20B of the OLS in which the temporal layer indication indicates that all bitstream portions of a given reference layer 20C (of the reference layer of the respective layer) on which the respective layer 20B depends are one (e.g., the same) of a set of predetermined bitstream portion types for a further reference layer 20D on which the given reference layer 20C depends directly or indirectly. The following two criteria are satisfied: first, access units 40A, 40B that include a bitstream portion of one of the set of predetermined bitstream portion types among the bitstream portions of the given reference layer either have no bitstream portions other than the set of predetermined bitstream portion types (40A) or do not (40B). second, each further reference layer 20D is a reference layer of the direct reference layer 20A that depends on the respective layer 20B according to the reference layer indication.
[0067] According to an alternative embodiment, the encoder 10 is configured to provide in the video bitstream 14, in addition to the OLS indication and reference layer indication described with respect to the previous embodiment, for each layer (e.g., j) of an OLS (e.g., i), an intra-layer temporal layer indication (e.g., NumSubLayersInLayerInOLS[i][j]) indicating whether the OLS requires only bitstream portions of the respective layer that are one of a set of predetermined bitstream portion types (e.g., indicated by NumSubLayersInLayerInOLS[i][j] = 0), or otherwise, a subset of temporal layers (e.g., maximum temporal layer index) that includes bitstream portions of the respective layer that the OLS requires.
[0068] According to this embodiment, the encoder 10 is configured to provide a respective bitstream 14 for each bitstream portion of an access unit that includes bitstream portions of one of a set of predetermined bitstream portion types for each layer 24 of the OLS whose intra-layer temporal layer indication indicates that the OLS requires only bitstream portions of one of a set of predetermined bitstream portion types, where the following conditions are met: if the respective bitstream portion belongs to a layer of the OLS and its intra-layer temporal indication indicates that the OLS requires only bitstream portions of the respective layer that are one of a set of predetermined bitstream portion types, the respective bitstream portion is one of a set of predetermined bitstream portion types, or according to the reference layer indication, the respective layer is independent of the layer of the respective bitstream portion.
[0069] According to a further embodiment, if an IRAP is not aligned, inter-layer prediction is not used for such IRAP NAL units for layers in which non-RAP NAL units reside in the same AU.
[0070] Therefore, according to another embodiment according to the second aspect, the encoder 10 is configured to provide, in the video bitstream 14, an OLS representation, a reference layer representation and a temporal layer representation, as described with respect to the previous embodiment in section 2. Further, according to this embodiment, the encoder 10 is configured to encode, without using inter-prediction means, bitstream portions of an access unit including bitstream portions of a predetermined reference layer that is one of a set of predetermined bitstream portion types, with respect to bitstream portions belonging to a layer that has direct or indirect reference to one of the further reference layers that is not devoid of bitstream portions other than the set of predetermined bitstream portion types, if the temporal layer indication for each layer (e.g., i) of the OLS indicates that all bitstream portions of a predetermined reference layer (of the reference layers of the respective layer) on which the respective layer depends are one of a set of predetermined bitstream portion types, and if an access unit including bitstream portions of a predetermined reference layer that is one of a set of predetermined bitstream portion types is not devoid of bitstream portions other than the set of predetermined bitstream portion types, for each reference layer on which the predetermined reference layer directly or indirectly depends.
[0071] For example, the set of predetermined bitstream portion types may include one or more or all of the IRAP types and the GDR types with ph_recovery_poc_cnt equal to zero.
[0072] An embodiment of the encoder 10 according to the second aspect may be configured to provide, within the video bitstream 14, a level indication of the bitstream 12 that can be extracted from the video bitstream according to the OLS. For example, the level indication includes one or more of the following: encoded picture buffer size, decoded picture buffer size, picture size, picture rate, minimum compression ratio, picture partitioning limit (e.g., tile / slice / subpicture), bit rate, buffer scheduling (e.g., HRD timing (AU / DU removal time, DPB output time)).
[0073] 3. Bitstream-based OLS Decision Section 3 describes an embodiment according to a third aspect of the invention with reference to Figure 1, and details described in Section 0 may optionally be applied to embodiments according to the third aspect. Also, details described with respect to further aspects may optionally be implemented in embodiments described in this section.
[0074]
[0033] Embodiments of the third aspect can provide identification of an OLS corresponding to a bitstream. In other words, embodiments of the third aspect enable estimation of the OLS of a video bitstream from which the OLS is to be decoded or extracted from the video bitstream. A decoder receiving a bitstream to be decoded may be provided with additional information regarding the operation point to be decoded via its API. For example, in the current VVC draft specification, two variables are set externally as follows:
[0075] The variable TargetOlsIdx, which identifies the OLS index of the target OLS to decode, and the variable Htid, which identifies the highest temporal sublayer to decode, are set by external means not specified in this specification. The bitstream BitstreamToDecode does not contain any layers other than those included in the target OLS, and does not contain any NAL units whose TemporalId is greater than Htid.
[0076] This specification is silent on what to do if these variables are not set, since a decoder in such a case is expected to simply decode the entire given bitstream rather than decoding a subset of the bitstream, e.g., for temporal sublayers.
[0077] However, there is a problem with output layer sets: If a decoder is given a bitstream containing multiple layers and a parameter set defines multiple OLSs that include all layers in the bitstream (e.g., variants with different output layers), the decoder cannot simply determine which output layer set it must decode from the bitstream itself. Depending on the characteristics of the OLS, the OLS to be selected may result in different level requirements, due to different DPS parameters, etc. Therefore, it is important to allow the decoder to select an OLS even in the absence of an external signal via the API. In other words, a fallback method is necessary, just as in other cases where no external means exist, such as selecting the highest temporal sublayer in the bitstream to decode.
[0078] In one embodiment, there is a bitstream compatibility constraint that a bitstream only supports a single OLS, so that a decoder can unambiguously determine which OLS to decode from a given bitstream. This property can be instantiated, for example, by a syntax element that indicates that all OLSs are unambiguously determinable by the layers present in the bitstream, i.e., there is a unique mapping from layer numbers to OLSs.
[0079] According to an embodiment of the third aspect, an encoder 10 for providing a multi-layer video bitstream 14 is configured to represent, within the multi-layer video bitstream 14, a plurality of OLSs, e.g., in the OLS representation 18 of FIG. 1 . Each of the OLSs represents a subset of layers of the multi-layer video bitstream 14. Note that the subset of layers is not necessarily a proper subset of layers. That is, the subset of layers may include all layers of the multi-layer video bitstream 14. The encoder 10 according to this embodiment provides the multi-layer video bitstream 14 such that, for each OLS, a sub-bitstream, such as sub-bitstream 12 of the multi-layer video bitstream 14 defined by the respective OLS, is distinguishable from a sub-bitstream of the multi-layer video bitstream defined by any other OLS among the plurality of OLSs. For example, the encoder 10 may provide the multi-layer video bitstream such that the OLSs represent mutually distinct subsets of layers of the multi-layer video bitstream 14, such that the OLSs are distinguishable by their subsets of layers.
[0080] For example, each OLS may be defined by indicating, for the OLS, a subset of layers by layer index in the OLS indication 18. Optionally, the OLS indication may include further parameters that define the subset of bitstream portions of the layers of the OLS whose bitstream portions belong to the OLS. For example, the OLS indication may indicate which temporal sublayers belong to the OLS.
[0081] In an example, encoder 10 may indicate within multi-layer video bitstream 14 that multi-layer video bitstream 14 uniquely belongs to one of the OLSs. For example, encoder 10 may indicate multiple OLSs, such that, for each OLS, the subset of layers in each OLS is different from any of the subsets of layers in the other OLSs. Thus, in an example, encoder 10 may indicate that a set of layers of multi-layer video bitstream 14, which may be indicated by, for example, a set of indices within multi-layer video bitstream 14, uniquely belongs to one of the OLSs.
[0082] According to an embodiment, encoder 10 is configured to check conformance of multi-layer video bitstream 14 by checking, for each OLS of the plurality of OLSs, whether the sub-bitstreams of multi-layer video bitstream 14 defined by the respective OLS are distinguishable or different from the sub-bitstreams of multi-layer video bitstreams defined by any of the others of the OLSs. For example, if not, encoder 10 can deny conformance of the bitstream.
[0083] Thus, an embodiment of a decoder for decoding a video bitstream, e.g., decoder 50 of Figure 1, when decoding a video bitstream, e.g., video bitstream 14 or a video bitstream extracted therefrom, such as video bitstream 12 of Figure 1, may derive one or more OLSs from the video bitstream to be decoded, each OLS indicating a subset of the layers of the video bitstream. The decoder may detect an indication in the video bitstream that indicates that the video bitstream uniquely belongs to one of the OLSs and decode one of the OLSs that belongs to the video bitstream.
[0084] For example, decoder 50 may identify an OLS that belongs to a video bitstream by identifying the layers included in the video bitstream, and decode the OLS that accurately identifies the layers included in the video bitstream. Thus, an indication that a video bitstream uniquely belongs to one of the OLSs may indicate that a set of layers in the video bitstream uniquely belongs to one of the OLSs.
[0085] For example, the decoder 50 may determine one of the OLSs belonging to a video bitstream by examining the first access unit of a coded video sequence. The first access unit may refer to the first received, the first in temporal order, the first in decoding order, or the first in output order. Alternatively, the decoder 50 may determine one of the OLSs by examining the first of the access units as being a sequence-start access unit type, e.g., a CVSS access unit, where the first is defined, for example, by the receiving order, temporal order, or decoding order. For example, the decoder 50 may examine the first of the access units of a coded video sequence or the first of the access units as being a sequence-start access unit type for the layer contained in the respective access unit.
[0086] For example, decoder 50 may determine one of the OLSs when the first access unit of a coded video sequence, or the first of the access units, is of a sequence start access unit type, such that each access unit contains pictures of exactly one layer of the OLS.
[0087] Some of the above embodiments may have the drawback that certain combinations of OLSs are prohibited, for example in a multiview two-layer scenario with an OLS that outputs both views and an OLS that outputs only the dependently coded view. To mitigate this limitation, another embodiment of the present invention is to have a selection algorithm among the OLSs corresponding to the bitstream or its first access unit or its CVSS AU, for example by a combination of one or more of the following:
[0088] Selecting the OLS with the highest or lowest index Select the OLS with the most output layers. FIG. 7 illustrates a decoder 50 according to one embodiment, which may optionally follow the latter embodiment with a selection algorithm. The decoder 50 according to FIG. 7 may optionally correspond to the decoder 50 according to FIG. 1. The decoder 50 according to FIG. 7 is configured to decode a video bitstream 12, e.g., a video bitstream 12 extracted from a multi-layer video bitstream 14. Alternatively, the video bitstream 12 of FIG. 7 may correspond to the video bitstream 14 of FIG. 1. The video bitstream 12 of FIG. 7 may be a multi-layer video bitstream, but this is not necessarily the case, i.e., in the example, it may be a single-layer video bitstream. The video bitstream 12 includes access units 22 of a coded video sequence 20, e.g., access units 221 and 222, each of which includes one or more pictures of the coded video sequence, e.g., pictures 261 and 262. Each of the pictures 26 belongs to one of one or more layers 24 of the video bitstream 12, e.g., as described in Section 0. Decoder 50 according to FIG. 7 is configured to derive one or more OLSs from video bitstream 12. For example, video bitstream 12 includes OLS indication 18, e.g., as described with respect to FIG. 1, which indicates one or more OLSs, such as OLS 181 and OLS 182 as shown in FIG. 7. Each of the one or more OLSs indicates a (not necessarily appropriate) set of one or more layers 24 of video bitstream 12. In other words, each of the OLSs indicates one or more of the layers that should be part of the respective OLS. Decoder 50 determines one of the OLSs based on one or more attributes of each of the OLSs and decodes the determined one OLS from the OLSs.
[0089] As described with respect to the previous embodiment, the decoder 50 may determine a subset of OLSs and one OLS by examining whether the first of the access units in the coded video sequence, or the first of the access units, is a sequence-start access unit type. For example, the access unit 221 shown in FIG. 7 may be the first of the access units in the coded video sequence of the video bitstream 12 (e.g., the first received one, or the first in encoding order, the first in temporal order, or the first in output order). In another example, the coded video sequence 20 may include an additional access unit preceding the access unit 221, and because the preceding access unit is not a sequence-start access unit, the access unit 221 is the first sequence-start access unit of the sequence 20. The decoder 50 may examine the access unit 221 to detect the picture 261 in the first layer 241 and the picture 262 in the second layer 242. Based on this knowledge, the decoder 50 may conclude that the video bitstream 12 comprises the first layer 241 and the second layer 242.
[0090] According to the embodiment of FIG. 7, the decoder 50 determines one OLS to decode based on one or more attributes of the OLSs. The one or more attributes may include one or more of the index of the respective OLS (i.e., the OLS index), the number of layers of the OLS, and / or the number of output layers of the respective OLS. For example, the decoder determines the OLS with the highest or lowest OLS index and / or the largest number of layers and / or the most output layers as one OLS. In other words, the decoder 50 may evaluate which OLS has the highest or lowest OLS index and / or which OLS has the most layers. Additionally or alternatively, the decoder 50 may evaluate which of the OLSs includes the output layer with the highest or lowest index. By selecting one OLS based on the most number of output layers and / or the most number of layers, a bitstream that provides the highest quality output for the video sequence may be selected for decoding.
[0091] For example, the decoder 50 may select, as one OLS, the OLS with the largest number of output layers, the OLS with the largest number of layers beyond the OLS with the largest number of output layers, or the OLS with the lowest OLS index beyond the OLS with the largest number of output layers beyond the OLS with the largest number of layers.
[0092] According to an embodiment, the decoder 50 determines one OLS by evaluating which OLS has the largest number of layers among the OLSs. If there are multiple OLSs with the largest number of layers, the decoder 50 may evaluate the OLS with the largest number of layers that has the largest number of output layers, or may select the OLS with the largest number of output layers as one OLS.
[0093] In other words, if there are no OLS instructions to encode, decoder 50 can decode the OLS indicated by OLS indication 18 in which all required layers are present in the bitstream and which uses most of the layers present, providing a high fidelity video output.
[0094] In the following, a further embodiment of the decoder 50 is described with reference to FIG.
[0095] Also, decoder 50 of this further embodiment may optionally follow the selection algorithm described above and may optionally correspond to decoder 50 according to Fig. 1. According to this further embodiment, decoder 50 is configured to decode video bitstream 12, e.g., video bitstream 12 extracted from multi-layer video bitstream 14, as shown in Fig. 1. Alternatively, video bitstream 12 of Fig. 7 may correspond to video bitstream 14 of Fig. 1. Video bitstream 12 of Fig. 7 may be a multi-layer video bitstream, but this is not necessarily the case, i.e., in the example, it may be a single-layer video bitstream. Video bitstream 12 includes access units 22 of coded video sequence 20, e.g., access units 221, 222, each of which includes one or more pictures of the coded video sequence, e.g., pictures 261, 262. Each of pictures 26 belongs to one of one or more layers 24 of video bitstream 12, e.g., as described in Section 0. A decoder 50 according to this further embodiment is configured to derive one or more OLSs from a video bitstream 12. For example, the video bitstream 12 includes an OLS indication 18, e.g., as described with respect to FIG. 1, which indicates one or more OLSs, such as OLS 181 and OLS 182 as shown in FIG. 7. Each of the one or more OLSs indicates a (not necessarily proper) set of one or more layers 24 of the video bitstream 12. In other words, each OLS indicates one or more layers that should be part of the respective OLS. The decoder 50 according to this further embodiment determines a subset of OLSs from (or among) the OLSs, such that each OLS in the subset of OLSs belongs to the video bitstream 12. The decoder 50 according to this further embodiment further determines one of the subset of OLSs based on one or more attributes of each of the subset of OLSs, and decodes the determined one OLS from the subset of OLSs.
[0096] For example, an OLS attributed to a video bitstream may mean that a set of one or more layers present in video bitstream 12 corresponds to the set of layers indicated in the respective OLS. In other words, decoder 50 may determine a subset of OLSs attributed to video bitstream 12 based on the set of layers present in video bitstream 12. That is, decoder 50 may determine a subset of OLSs such that, for each subset of OLSs, the subset of layers indicated by the respective OLS corresponds to the set of layers in video bitstream 12. For example, in FIG. 7, video bitstream 12 illustratively includes picture 261 of first layer 241 and picture 262 of second layer 242. OLS181 indicates that first layer 241 and second layer 242 are part of OLS181. Furthermore, OLS182 indicates that both the first layer and the second layer are part of OLS182. Thus, according to the example of FIG. 7, decoder 50 can attribute both OLSs 181, 182 to the subscript portions of the OLSs attributed to video bitstream 12.
[0097] Optionally, the decoder 50 may consider only those OLSs for decoding that are decodable by the decoder 50 according to the level information of the OLSs.
[0098] As described with respect to the previous embodiment, the decoder 50 may determine a subset of the attributed OLSs and one OLS by examining whether the first of the access units in the coded video sequence, or the first of the access units, is a sequence-start access unit type. For example, the access unit 221 shown in FIG. 7 may be the first of the access units in the coded video sequence of the video bitstream 12 (e.g., the first received one, or the first in encoding order, the first in temporal order, or the first in output order). In another example, the coded video sequence 20 may include an additional access unit preceding the access unit 221, and because the preceding access unit is not a sequence-start access unit, the access unit 221 is the first sequence-start access unit of the sequence 20. The decoder 50 may examine the access unit 221 to detect the picture 261 in the first layer 241 and the picture 262 in the second layer 242. Based on this knowledge, the decoder 50 may conclude that the video bitstream 12 comprises the first layer 241 and the second layer 242.
[0099] According to an embodiment, decoder 50 can determine the subset of OLSs such that the first of the access units in the coded video sequence, or if the first of the access units is a sequence start access unit type, e.g., access unit 221, each access unit contains exactly one picture of each layer of the subset of OLSs.
[0100] According to the embodiment of FIG. 7, decoder 50 determines one OLS to decode based on one or more attributes of the subset of OLSs. The one or more attributes may include one or more of the index and / or number of output layers of the respective OLS, the highest or lowest index, i.e., the highest or lowest layer index, and the highest number of layers. In other words, decoder 50 may evaluate which OLS of the subset of OLSs attributable to video bitstream 12 includes layers indexed with the highest or lowest layer index and / or which OLS has the most layers. Additionally or alternatively, decoder 50 may evaluate which OLS of the subset of OLSs has the greatest or smallest number of output layers and / or which OLS of the subset of OLSs constitutes the output layer with the highest or lowest index.
[0101] According to an embodiment, the decoder 50 determines one OLS by evaluating which OLS among the subset of OLSs has the largest number of layers. If there are multiple OLSs with the largest number of layers, the decoder 50 may evaluate the OLS with the largest number of layers that has the largest number of output layers, or may select the OLS with the largest number of output layers as the one OLS.
[0102] In other words, if there are no OLS instructions to encode, decoder 50 can decode the OLS indicated by OLS indication 18 in which all required layers are present in the bitstream and which uses most of the layers present, providing a high fidelity video output.
[0103] In other words, an embodiment of the third aspect includes a decoder 50 for decoding a video bitstream 12, 14, where the video bitstream 14 includes access units 22 of a coded video sequence 20, each access unit 22 including one or more pictures 26 of the coded video sequence, each of the pictures belonging to one of one or more layers 24 of the video bitstream 14. The decoder is configured to: derive one or more output layer sets (OLSs) 181, 182 from the video bitstream 14, each indicating a set of one or more layers of the video bitstream 14; determine from the OLSs 181, 182 a subset of the OLSs 181, 182, where each of the subsets of the OLSs belongs to the video bitstream 14; determine one of the subsets of the OLSs based on one or more attributes of each subset of the OLSs; and decode the OLSs.
[0104] According to an embodiment, decoder 50 is configured to determine a subset of OLSs such that, for each subset of OLSs, the subset of layers indicated by the respective OLS corresponds to a set of layers in video bitstream 14.
[0105] According to an embodiment, the decoder 50 is configured to determine the subset of OLSs and one OLS by checking that the first of the access units of the encoded video sequence, or the first of the access units 22, is of the sequence start access unit 22 type.
[0106] According to an embodiment, the decoder 50 is configured to determine a subset of the OLS such that the first of the access units 22 of the coded video sequence, or if the first of the access units 22 is of a sequence start access unit type, each access unit contains exactly one picture of each layer of the subset of the OLS.
[0107] According to an embodiment, the decoder 50 is configured to determine one of the OLSs by evaluating a respective criterion for one or more attributes of a subset of the OLSs.
[0108] 4. Sub-bitstream sequence start access unit Section 4 describes embodiments of the fourth aspect of the present invention with reference to Figure 1. The descriptions provided in Section 0 may optionally be applied to embodiments of the fourth aspect, and details described with respect to further aspects may optionally be implemented in embodiments described in this section.
[0109] Some embodiments of the fourth aspect may relate to access unit delimiters (AUDs) in supplemental enhancement information (SEI) to enable coded video sequence start access units (CVSS AUs) in a video bitstream that was not originally a CVSS AU, such as video bitstream 14 from which video bitstream 12 was extracted. For example, a CVSS AU may be a randomly accessible AU, such as an AU with randomly accessible or independently coded pictures in each layer of the video bitstream, or an AU that is independently decodable from a previous AU in the video bitstream.
[0110] The current specification mandates that a Coded Video Sequence Start (CVSS) AU has either an IRAP or GDR NAL unit type in each layer, that the IRAP NAL unit type be the same within a CVSS AU, and that an AUD (Access Unit Delimiter) be present to indicate that the CVSS AU is an IRAP or GDR AU.
[0111] Figure 8 shows an example of a multi-layer bitstream where the IRAPs are not aligned at different layers, i.e., not all IRAPs are aligned across layers, and the identification of CVSS AUs.
[0112] Although AU2, AU4, and AU6 have IRAP type NAL unit types in the two lowest layers, not all layers have the same IRAP type for these AUs, so these AUs are not CVSS AUs. To easily identify CVSS AUs without having to parse all NAL units in an AU, the AUD NAL unit is used and easily identified. That is, AU0 and AU8 contain AUDs with a flag indicating that these AUs are CVSS AUs (IRAP AUs).
[0113] 9 shows an example of a video bitstream after extraction, e.g., video bitstream 12. However, if the bitstream is extracted only at L0 and L1, new CVSS AUs will exist in the extracted bitstream, as indicated by reference numeral 22*. In other words, after extraction, some AUs (e.g., AU22*) will be changed to CVSS AUs.
[0114] AU2 and AU6 can be CVSS AUs or IRAP AUs, so there must be an AUD present in the bitstream of such AUs that indicates the IRAP AU property.
[0115] An embodiment according to a fourth aspect includes an apparatus for extracting sub-bitstreams from a multi-layer video bitstream, such as an extractor 30 for extracting sub-bitstream 12 from multi-layer video bitstream 14 described with reference to FIG. 1. According to the fourth aspect, multi-layer video bitstream 14 represents a coded video sequence, such as coded video sequence 20, and the multi-layer video bitstream includes access units 22 of the coded video sequence. Each of the access units 22 includes one or more bitstream portions 16 of multi-layer video bitstream 14, each bitstream portion belonging to one of layers 24 of the multi-layer video bitstream. According to the fourth aspect, extractor 30 is configured to derive one or more optically significant segments (OLSs) from multi-layer video bitstream 14, each of which indicates a (not necessarily appropriate) subset of layers of multi-layer video bitstream 14. For example, multi-layer video bitstream includes OLS indications 18 including information about or descriptions of one or more OLSs, such as OLS 181, OLS 182, as shown in FIG. 7.
[0116] The extractor 30 according to the fourth aspect is configured to provide, within the sub-bitstream 12, layers 24 of the multi-layer video bitstream 14 that are indicated by a predetermined one of the OLSs, i.e., that are indicated as being part of the predetermined OLS. In other words, the extractor 30 can provide, within the sub-bitstream 12, bitstream portions 16 that belong to each layer of the predetermined OLS. For example, the predetermined OLS may be provided to the extractor 30 by external means, such as an OLS instruction 32 as shown in FIG. 1. According to an embodiment of the fourth aspect, the extractor 30 provides, within the sub-bitstream 12, for each of the access units of the sub-bitstream 12, a sequence start indication indicating that the sub-bitstream access unit 22* including L0 and L1 described with reference to FIGS. 8 and 9 is the start access unit of a sub-sequence of a coded video sequence when all bitstream portions of the respective access units are bitstream portions of the same type from a set of predetermined bitstream portion types.
[0117] In other words, the extractor 30 may provide access units that the extractor 30 includes or provides in the sub-bitstream 12 exclusively contains bitstream portions of the same type from a set of predetermined bitstream portion types that have sequence start indications.
[0118] For example, the set of predefined bitstream portion types may include one or more IRAP NAL unit types and / or GDR NAL unit types, e.g., the set of predefined bitstream portion types may include IDR_NUT, CRA_NUT, or GDR_NUT NAL unit types.
[0119] For example, extractor 30 may determine, for each access unit that does not have a sequence start indication in the multi-layer video bitstream, or alternatively, for each access unit of sub-bitstream 12, whether all bitstream portions of the respective access unit are bitstream portions of the same one of a set of predetermined bitstream portion types. For example, extractor 30 may analyze relevant information within each access unit or multi-layer video bitstream to determine whether all bitstream portions of the respective access unit are bitstream portions of the same one of a set of predetermined bitstream portion types.
[0120] In other words, in one embodiment, the bitstream extraction process removes unnecessary layers if they are not present in the AUs that require AUDs after extraction, and adds AUD NUTs for such uses.
[0121] According to an embodiment, the extractor 30 is configured to infer from an indication in the multi-layer video bitstream 14 that, for a given OLS, one of the access units is the starting access unit of a sub-sequence of the coded video sequence represented by the given OLS, and to provide a sequence start indication in the sub-bitstream indicating that one access unit is the starting access unit.
[0122] In other words, in another embodiment, there is an indication in the bitstream for such AUs (i.e., those that become CVSS AUs), and such AUs become CVSS AUs or IRAP AUs when the layers are extracted, making the insertion (addition) of AUDs simpler and less analysis-intensive.
[0123] According to a further embodiment, extractor 30 is configured to extract nested information, e.g., nested SEI, that indicates, for a given OLS, that one or more access units, e.g., access units that are not the start access unit of multi-layer video bitstream 14, are the start access units for the OLS. According to this embodiment, extractor 30 provides, within the sub-bitstream, a sequence start indication that indicates that one or more access units indicated in the nested information are the start access units. For example, extractor 30 may provide a sequence start indication for each of the indicated access units, as described above. Alternatively, the device may provide a common indication for the indicated access units within the sub-bitstream.
[0124] In other words, in another embodiment, there is a nesting SEI that can encapsulate an AUD such that when a specific OLS, such as turning the AU into an IRAP / GDR AU, is extracted, the encapsulated AUD is de-encapsulated and added to the bitstream. Currently, the specification only includes nesting of other SEI messages. Therefore, non-VCL non-SEI payloads need to be allowed within the nested SEI. One way is to extend the existing nested SEI and indicate that a non-SEI is included. Another is to add a new SEI that includes other non-VCL payloads within the SEI.
[0125] The first option for implementation is shown in Table 2. [Table 2-1] [Table 2-2] A nesting SEI will encapsulate other SEIs (sei_message()) and a given number of non-VCL NAL units that are not SEIs (nesting_num_nonVclNuts_minus1). Such encapsulated non-VCL NAL units have their length (length_minus1) written immediately before the nested SEI (nonVclNut) so that boundaries within the nested SEI can be found.
[0126] Table 3 shows another option, option 2. [Table 3] In Option 2, if other non-VCL NAL units need to be included in the nesting SEI in the future, additional types that allow other non-VCL NAL units can be added so that the nonVclNutPayloadSEI message can be used.
[0127] In such cases, a new SEI message is defined that directly contains a single non-VCL NAL unit (in this case, the access unit delimiter (AUD_rbsp())), and thus such an encapsulating SEI can be added directly without modifying the nesting SEI (which already contains other SEIs within itself).
[0128] According to a further embodiment of the fourth aspect, the extractor 30 is configured to provide, for each access unit of the sub-bitstream, a sequence start indication indicating that the respective access unit is the start access unit of the coded video sequence if all bitstream portions of the respective access unit are bitstream portions of the same one of a set of predetermined bitstream portion types and the respective access unit includes bitstream portions of two or more layers.
[0129] In another embodiment, AUDs are also mandatory for access units that may become IRAP or GDR AUs in case of OLS extraction (layer drop), and the extractor can ensure they are present when needed and easily rewrite aud_irap_or_gdr_au_flag to 1 when appropriate (i.e., the AU turns into an IRAP or GDR AU after extraction). One way to express this constraint in the specification is to extend the existing text that makes AUDs mandatory for IRAP or GDR AUs in the current bitstream.
[0130] There can be at most one EOB NAL unit in an AU. If vps_max_layers_minus1 is greater than 0, there shall be exactly one AUD NAL unit in each IRAP or GDR AU.
[0131] This is changed to:
[0132] There can be at most one EOB NAL unit in an AU. If vps_max_layers_minus1 is greater than 0, there shall be exactly one AUD NAL unit in each AU that contains at least two layers containing only IRAP or GDR NAL units.
[0133] In other words, instead of using all NAL units, the words IRAP or GDR picture are used.
[0134] There can be at most one EOB NAL unit in an AU. If vps_max_layers_minus1 is greater than 0, there shall be exactly one AUD NAL unit in each AU that contains at least two IRAP or GDR pictures.
[0135] Accordingly, a further embodiment according to a fourth aspect includes an encoder 10 for providing a multi-layer video bitstream, such as the encoder 10 described with reference to FIG. 1. The embodiment of the encoder 10 according to the fourth aspect is configured to provide a multi-layer video bitstream representing an encoded video sequence, such as video sequence 20. The multi-layer video bitstream 14 provided by the encoder 10 includes a sequence of access units 22, each of which includes one or more pictures, each of which is associated with one of a plurality of layers 24 of the video bitstream. The encoder 10 is configured to provide, for each of the access units of the multi-layer video bitstream 14, within the sub-bitstream 12, a sequence start indicator that indicates whether all pictures of the respective access unit are of one of a set of predetermined picture types, provided that the respective access unit includes at least two pictures of one of the set of predetermined picture types. Optionally, the encoder 10 is configured to provide an output layer set (OLS) representation 18 of an OLS in the multi-layer video bitstream 14, where the OLS includes one or more layers of the multi-layer video bitstream 14, and where each access unit includes at least two pictures of one of a set of predetermined access unit types, the encoder 10 provides a sequence start indicator for each access unit.
[0136] For example, picture types include one or more IRAP picture types and / or GDR picture types. A picture being a picture type may mean that one or more bitstream portions in which the picture is encoded are each one of a set of predetermined bitstream portion types, such as one or more IRAP NAL unit types or GDR NAL units.
[0137] For example, the sequence start indicators referred to in the embodiments throughout Section 4 may be provided in the form of or as part of an access unit delimiter (AUD), which may be a bitstream portion provided within a respective access unit. For example, an indication that an access unit is a sequence start access unit may be indicated by setting a flag in the AUD, e.g., aud_irap_or_gdr_au_flag, to indicate that the respective access unit is a sequence start access unit, e.g., to a value of 1.
[0138] Alternatively, the encoder 10 may not necessarily provide a sequence start indicator for each access unit that includes at least two pictures of one of a set of predetermined picture types, but may provide a sequence start indicator for each access unit of the multi-layer video bitstream 14 that includes at least two pictures of one of a set of predetermined picture types, each picture belonging to one of the layers of the OLS indicated in the OLS indication provided in the multi-layer video bitstream 14 by the encoder 10.
[0139] As an alternative to the above embodiment where AUD is mandated even for access units that could become IRAP or GDR AUs during OLS extraction (layer drop), the extractor can ensure they are present when needed and easily rewrite aud_irap_or_gdr_au_flag to 1 when appropriate (i.e., the AU becomes an IRAP or GDR AU after extraction). The encoder needs to write an AUD NAL unit only if the access unit corresponds to at least one CVSS AU in OLS. An example specification is as follows:
[0140] There can be at most one EOB NAL unit in an AU. If vps_max_layers_minus1 is greater than 0, there shall be exactly one AUD NAL unit in each AU that contains only IRAP or GDR NAL units in all layers of at least one multi-layer OLS.
[0141] That is, if we use the term IRAP or GDR picture:
[0142] There can be at most one EOB NAL unit in an AU. If vps_max_layers_minus1 is greater than 0, there shall be exactly one AUD NAL unit in each AU that contains only IRAP or GDR pictures in all layers of at least one multi-layer OLS.
[0143] In this embodiment, the associated OLS extraction process will be extended through the following steps:
[0144] […] The output sub-bitstream OutBitstream is derived as follows. - The bitstream outBitstream is set to be the same as the bitstream inBitstream. - Remove all NAL units with TemporalId greater than tIdTarget from outBitstream. - Remove all NAL units from outBitstream whose nal_unit_type is not equal to any of VPS_NUT, DCI_NUT, and EOB_NUT and whose nuh_layer_id is not included in LayerIdInOls[targetOlsIdx]. - If an AU has two or more layers and contains only NAL units whose nal_unit_type is equal to a single type of IDR_NUT, CRA_NUT, or GDR_NUT, rewrite the flag aud_irap_or_gdr_au_flag in the AUD of the AU to 1.
[0145] Thus, the extractor 30 described above according to the fourth aspect may provide a sequence start indicator by setting a value of the sequence start indicator for each access unit 22*. For example, the sequence start indicator may be a syntax element signaled in an access unit of the multi-layer video bitstream 14, and the extractor 30 may modify or preserve the value of the sequence start indicator when forwarding the access unit in the sub-bitstream 12.
[0146] In other words, according to an embodiment, when the extractor 30 according to the fourth aspect transfers each access unit 22* of the multi-layer video bitstream 14 in the sub-bitstream 12, for each of the sub-bitstream access units 22, if all bitstream portions of the respective access units are bitstream portions of the same type of a predetermined set of bitstream portion types, e.g., access unit 22*, the extractor 30 may provide a sequence start indication in the sub-bitstream 12 indicating that the respective access unit is the start access unit of a coded video sequence by setting the value of a sequence start indicator, e.g., aud_irap_or_gdr_flag, present in the respective access unit of the multi-layer video bitstream (e.g., present in the AUD NAL unit of the respective access unit 22*) to a predetermined value, e.g., 1. The predetermined value indicates that the respective access unit is the first access unit of the coded video sequence. For example, if the value of the sequence start indicator does not have a predetermined value in the multi-layer video bitstream 14, the extractor 30 may change the value of the sequence start indicator to a predetermined value.
[0147] Thus, an embodiment of encoder 10 for providing a multi-layer video bitstream 14 according to the fourth aspect provides, in the multi-layer video bitstream 14, an OLS indication 18 that indicates an OLS that includes layers (i.e., at least two layers) of the video bitstream 14. For each access unit that includes pictures of one of the predetermined picture types (e.g., the same type, or not necessarily the same type) for the layers of the OLS, encoder 10 may provide a sequence start indicator in the multi-layer video bitstream 14, the sequence start indicator indicating whether all pictures of the access unit, i.e., even those that are not part of the OLS, are of one of the predetermined picture types, e.g., depending on the value of the sequence start indicator.
[0148] In other words, encoder 10 may signal a sequence start indicator for an access unit, e.g., access unit 22*, whose pictures belong to one of the layers of the OLS and are of one of the predetermined types.
[0149] 5. Handling of temporal sublayers in the process of extracting output layer sets Section 5 describes an embodiment according to a fifth aspect of the invention, which may optionally follow the embodiments of the encoder 10 and extractor 30 described with reference to Figure 1. Also, details described with reference to further aspects may optionally be implemented in the embodiments described in this section.
[0150] Some embodiments according to the fifth aspect relate to the extraction process of OLS and vps_ptl_max_temporal_id[i][j]. Some embodiments according to the fifth aspect may relate to the derivation of NumSublayerInLayer[i][j].
[0151] To extract the output layer set (OLS), layers that do not belong to the OLS must be dropped or removed from the bitstream. Note, however, that layers that belong to the OLS may have a different number of sublayers (temporal layers TLx in Figure 10).
[0152] Figure 10 shows an example of a two-layer bitstream in which each of the two layers has a different frame rate. The bitstream in Figure 10 includes an access unit 221 associated with a first temporal layer TL0 and an access unit 222 associated with a second temporal layer TL1. The first layer 241 includes pictures from both temporal sub-layers TL0 and TL1, while the second layer 242 includes pictures from only TL0. Thus, the first layer 241 has twice the frame rate or picture rate of the second layer 242.
[0153] The bitstream in Figure 10 may have two OLSs: one consisting of only L0 and two operating points (e.g., TL0 at 30 fps and TL0+TL1 at 60 fps), and the other OLS may consist of L0 and L1, or it may consist of only TL0.
[0154] The current specification allows signaling of OLS profile and level using TL0 and TL1 of L0, but the extraction process does not allow for the generation of a TL0-only bitstream.
[0155] NumSublayerInLayer[i][j], which represents the maximum sublayer currently included in the ith OLS of layer j, is set to vps_max_sublayers_minus1+1 if vps_max_tid_il_ref_pics_plus1[m][k] does not exist or if layer j is the output layer of the ith OLS.
[0156] FIG. 11 shows an encoder 10 and an extractor 30 according to an embodiment of the fifth aspect. The encoder 10 and the extractor 30 according to FIG. 11 may optionally correspond to the encoder 10 and the extractor 30 described with reference to FIG. 1. Thus, the description of FIG. 1 may also optionally apply to the elements shown in FIG. 11. The encoder 10 according to FIG. 11 is configured to encode a coded video sequence 20 into a multi-layer video bitstream 14. The multi-layer video bitstream 14 includes access units 22, e.g., access units 221, 222, and 223 of FIG. 11 (or also of FIG. 1), each of which constitutes one or more pictures 26 of the coded video sequence. For example, each of the access units 22 includes one or more pictures related to one common instant or frame of the coded video sequence, as described with reference to FIG. 1. Each of the pictures 26 belongs to one of the layers 24 of the multi-layer video bitstream. For example, in the illustrative example of FIG. 11 , the multi-layer video bitstream 14 includes a first layer 241 and a second layer 242, where the first layer includes picture 261 and the second layer includes picture 262. According to the fifth aspect, each of the access units 22 belongs to a temporal sublayer of the set of temporal sublayers TL0, TL1 of the coded video sequence 20. For example, the access units 221 and 223 of FIG. 11 may belong to the first temporal sublayer TL0, and the access unit 222 may belong to the second temporal sublayer TL1, e.g., as described with respect to FIG. 10 . The temporal sublayers are sometimes also referred to as temporal subsets or temporal layers. For example, each of the temporal sublayers may be indicated or indexed with a temporal identifier and characterized by, e.g., a temporal relationship with respect to a frame rate and / or other temporal subsets. For example, each of the access units may include or be associated with a temporal identifier that associates the respective access unit with one of the temporal sublayers.
[0157] According to a fifth aspect, the encoder 10 is configured to provide a syntax element, e.g., the above-mentioned max_tid_within_ols, in the multi-layer video bitstream 14, which indicates a predetermined temporal sublayer for the OLS, where the OLS includes or indicates a (not necessarily proper) subset of the layers of the multi-layer video bitstream 14. The syntax element indicates the predetermined temporal sublayer for the OLS in a manner that distinguishes between different states, including states where the predetermined temporal sublayer is less than the maximum temporal sublayer in an access unit that is at least one picture of the subset of layers. For example, the predetermined temporal sublayer is the maximum temporal sublayer included in the OLS.
[0158] For example, encoder 10 may provide, within multi-layer video data stream 14, OLS indication 18, e.g., as described with respect to Figure 1. OLS indication 18 may include a description or indication of one or more OLSs, e.g., OLS 181 shown in Figure 11. Each of the OLSs may be associated with a set of layers of multi-layer video data stream 14. In the example shown in Figure 11, OLS 181 is associated with first layer 241 and second layer 242. It should be noted that multi-layer video bitstream 14 may optionally include additional layers.
[0159] For example, with respect to the example shown in FIG. 11 , OLS 181 may include a first temporal sublayer (to which access units 221 and 223 may belong), but a second temporal sublayer, to which access unit 222 may belong, is not included in the OLS in this example. Therefore, according to this example, TL0 may be the largest temporal sublayer included in the OLS. Note that temporal sublayers may have a hierarchical order. In the example of FIG. 11 , among the temporal sublayers in access unit 22, at least one picture of a subset of layers (layer 241 that is part of the OLS), e.g., picture 261 in the case of access unit 221, is the largest, which is TL1. Thus, the largest temporal sublayer included in the OLS is below the largest temporal sublayer TL1. Thus, a given temporal sublayer may be indicated, for example, by indicating that it is below the largest temporal sublayer present in an access unit at least partially included in the OLS. In a further example, a given temporal sublayer may be indicated by indicating an index identifying the given temporal sublayer.
[0160] Thus, in the example of OLS 181, the given temporal sublayer may be the first temporal sublayer. Syntax elements provided in the multi-layer video bitstream 14, such as max_tid_within_ols or vps_ptl_max_temporal_id, indicate the given temporal sublayer for the OLS. The syntax elements distinguish different states. According to one of the states, the given temporal sublayer is below the largest temporal sublayer in the access unit in which at least one picture of the subset of layers is included. For example, in FIG. 11, the example of OLS 181 described above includes layers 241 and 242. The largest temporal sublayer in the access unit of the subset of layers of OLS 181 is the second temporal sublayer to which access unit 222 belongs. In examples in which the second temporal sublayer does not belong to OLS 181, a syntax element can indicate this state.
[0161] The extractor 30 according to the fifth aspect can derive syntax elements from the multi-layer video bitstream 14 and can provide the sub-bitstream 12 by selectively forwarding pictures of the multi-layer video bitstream 14 in the sub-bitstream 12, where each picture belongs to one of the layers of the OLS 181 and the picture belongs to an access unit 221, 223 that belongs to a temporal sublayer equal to or below a given temporal sublayer.
[0162] That is, the extractor 30 can provide a bitstream portion of each picture 26 in the sub-bitstream 12 if the picture belongs to a temporal sublayer equal to or below a given temporal sublayer, and can drop, i.e., not forward, the picture otherwise.
[0163] That is, the extractor 30 can use syntax elements in constructing the sub-bitstream 12 to exclude pictures that are not included in the OLS to be decoded but belong to a temporal sublayer that is included in one of the layers indicated by the OLS from being transmitted in the sub-bitstream 12.
[0164] According to an embodiment, multi-layer video bitstream 14 indicates, for each layer of an OLS, a syntax element indicating a predetermined temporal sublayer, e.g., a maximum temporal sublayer, included in the respective OLS. Based on the syntax element for a layer of the OLS, extractor 30 may identify whether bitstream portions belonging to a temporal sublayer of the OLS belong to a temporal sublayer that does not belong to the OLS, and consider the bitstream portions belonging to a temporal sublayer of the OLS for forwarding to sub-bitstream 12.
[0165] For example, a syntax element may be part of an OLS representation 18, e.g., a syntax element may be part of the OLS 181 to which it refers. For example, a syntax element may be part of a video parameter set for the respective OLS.
[0166] In one embodiment, signaling (e.g., of maximum temporal sublayers) is provided in the bitstream to indicate that the OLS has a maximum sublayer different from vps_max_sublayers_minus1+1, or is the maximum of all layers present in the OLS. For this purpose, the existing syntax element vps_ptl_max_temporal_id[i][j] may be reused to indicate the maximum sublayers present in the OLS.
[0167] According to some embodiments, furthermore, NumSublayerInLayer[i][j], which represents the maximum sublayer included in the ith OLS for layer j, is changed to vps_ptl_max_temporal_id[i][j] if vps_max_tid_il_ref_pics_plus1[m][k] does not exist or if layer j is the output layer of the ith OLS.
[0168] Alternatively, a new syntax element can be added to indicate the maximum sublayer within an OLS, e.g., max_tid_within_ols.
[0169] According to an embodiment, the encoder 10 and / or the extractor 30 are configured to derive decoder capability-related parameters for each of the pictures of the multi-layer video bitstream 14, for a sub-stream obtained by selectively inheriting the pictures, e.g., the sub-bitstream 12, if the respective pictures belong to one of the layers of the OLS 181 and if the pictures belong to an access unit equal to or below a predetermined temporal sublayer. That is, the encoder 10 and / or the extractor 30 can derive decoder capability-related parameters for a sub-bitstream that exclusively includes pictures belonging to a temporal sublayer that belongs to the OLS that describes the sub-bitstream. The encoder 30 or the extractor 30 can signal the capability-related parameters in the sub-bitstream 12. Thus, the encoder 10 can signal the capability-related parameters in the multi-layer video bitstream 14. For example, the decoder capability-related parameters may include parameters such as those described in Section 6.
[0170] 6. Handling of Temporal Sublayers in Video Parameter Signaling In Section 6, embodiments according to a sixth aspect of the present invention are described with reference to Figure 11 and Figure 1. Figures 1 and 11 may therefore optionally be applied to embodiments according to the sixth aspect, and details described with respect to further aspects may optionally be implemented in embodiments described in this section.
[0171] Some examples according to the sixth aspect relate to constraints on vps_ptl_max_temporal_id[i], vps_dpb_max_temporal_id[i], vps_hrd_max_tid[i] to be consistent for a given OLS.
[0172] The multi-layer video bitstream 14 and / or sub-bitstreams 12 described with respect to Figures 1 and / or 11 may optionally include a video parameter set 81. The video parameter set 81 may include one or more decoder requirement sets, e.g., profile-tier-level-sets (PTL sets), and / or one or more buffer requirement sets, e.g., DPB parameter sets, and / or one or more bitstream adaptation sets, e.g., hypothetical reference decoder (HRD) parameter sets. For example, each of the OLSs indicated in the OLS indication 18 may be associated with a respective decoder requirement set, buffer requirement set, and bitstream adaptation set that applies to the bitstream described by the respective OLS. For each of the decoder requirement set, buffer requirement set, and bitstream adaptation set, the video parameter set may indicate the maximum temporal sublayer referenced by the respective set, i.e., the maximum temporal sublayer of the video bitstream or video sequence referenced by the respective set.
[0173] Figure 12 shows an example of a video parameter set 81 including a first decoder requirement set 821, a first buffer requirement set 841, and a first bitstream adaptation set 861 associated with a first OLS1 of an OLS representation 18. Furthermore, according to Figure 12, the video parameter set 81 comprises a second decoder requirement set 822, a second buffer requirement set 842, and a second bitstream adaptation set 862 associated with a second output layer set OLS2.
[0174] For example, each of the OLSs described by the OLS indication 18 may be associated with one of the decoder requirement set 82, the buffer requirement set 84, and the bitstream conformance set 86 by associating with the respective OLS a respective index pointing to the decoder requirement set, the buffer requirement set, and the bitstream conformance set. According to an embodiment of the sixth aspect, the multi-layer video bitstream includes access units, each of which belongs to one of the sets of temporal sub-layers of the coded video sequence coded into the multi-layer video bitstream 14. The multi-layer video bitstream 14 according to the sixth aspect further comprises a video parameter set 81 and an OLS indication 18. For each of the bitstream conformance set 86, the buffer requirement set 84, and the decoder requirement set 82, the temporal subset indication indicates a constraint on a maximum temporal sub-layer, e.g., a maximum temporal sub-layer referenced by the respective bitstream conformance set / buffer requirement set / decoder requirement set. For example, each of the bitstream adaptation set 86, the buffer requirement set 84, and the decoder requirement set 82 signals a syntax element indicating the respective temporal subset indication (e.g., vps_ptl_max_temporal_id for the PTL set, vps_dpb_max_temporal_id for the DPB parameter set, and vps_hrd_max_tid for the bitstream adaptation set).
[0175] As shown in Figure 12, each of the bitstream conformance set 86, buffer requirement set 84, and decoder requirement set 82 may include one or more sets of parameters for each temporal sublayer present in the set of layers comprising the bitstream portion of the video bitstream referenced by the respective bitstream conformance set 86, buffer requirement set 84, or decoder requirement set 82. For example, in Figure 12, OLS1 includes layer L0, which comprises the bitstream portion of temporal layer TL0, and OLS2 includes layers L0 and L1, which comprise the bitstream portions of temporal layers TL0 and TL1. The bitstream conformance set 862 and decoder requirement set 822 associated with OLS2 include parameters for L0, and thus, according to this example, include parameters for temporal layer TL0, and further include parameters for L1, and thus, parameters for temporal layer TL1. The buffer requirement set 842 includes a set of parameters DPB0 for temporal sublayer TL0 and DPB1 for temporal sublayer TL1.
[0176] Typically, there are three syntactic structures in a VPS that are generally defined and then mapped to a specific OLS: Profile Hierarchy Level (PTL), e.g. one or more decoder requirement sets DPB parameters, e.g., one or more sets of buffer requirements HRD parameters, e.g., one or more bitstream adaptation sets The PTL to OLS mapping is performed for all VPSs with OLS (single layer or multi-layer). However, the DPB and HRD parameters are only mapped to OLSs with VPSs with multiple layers. As shown in Figure 12, the PTL, DPB, and HRD parameters are first written to the VPS, and then mapped to indicate which parameters the OLS should use.
[0177] In the example shown in Figure 12, there are two OLSs and two of each of these parameters. However, the definitions and mappings are specified so that multiple OLSs can share the same parameters, so there is no need to repeat the same information multiple times, as shown in Figure 13, for example.
[0178] FIG. 13 shows an example where OLS2 and OLS3 have the same PTL and DBP parameters, but different HRD parameters.
[0179] In the examples of Figures 12 and 13, the values of vps_ptl_max_temporal_id[i], vps_dpb_max_temporal_id[i], and vps_hrd_max_tid[i] for a given OLS are aligned (TL0 for OLS1, TL1 for OLS2, and TL1 for OLS3), but this is no longer required. These three values associated with the same OLS are not currently constrained to have the same value. There are currently no constraints on any of these values to match the number of sublayers in the bitstream. For example, in the example above, the values are defined for two sublayers, but the bitstream could have a single sublayer for OLS2 and OLS3. Therefore, as matching becomes more complex, the decoder will no longer be able to easily identify the characteristics of the bitstream.
[0180] In a first embodiment, the bitstream signals the maximum number of sub-layers present in the OLS (not necessarily the bitstream, as some may have been dropped), which can be understood at least as an upper bound, i.e. there cannot be more sub-layers than the signal value, e.g., vps_ptl_max_temporal_id[i], for the OLS in the bitstream. Therefore, the DPB and HRD parameters are also used by the decoder.
[0181] If the values of vps_dpb_max_temporal_id[i], vps_hrd_max_tid[i] are different from vps_ptl_max_temporal_id[i], the decoder needs to perform a more complex mapping. Thus, in one embodiment, if the OLS indexes the PTL structure, DPB structure, and HRD parameter structure by vps_ptl_max_temporal_id[i], vps_dpb_max_temporal_id[i], and vps_hrd_max_tid[i], respectively, there is a bitstream constraint that vps_dpb_max_temporal_id[i], vps_hrd_max_tid[i] must be equal to vps_ptl_max_temporal_id[i].
[0182] According to an embodiment, the encoder 10, for example the encoder 10 of FIG. 1 or FIG. 11, is configured to form the OLS representation 18 such that the maximum temporal sublayers indicated by the bitstream adaptation set 86, buffer requirement set 84, and decoder requirement set 82 associated with the OLS are equal to each other, and the parameters in the bitstream adaptation set 86, buffer requirement set 84, and decoder requirement set 82 are fully valid for the OLS.
[0183] However, looking at the example in Figure 12, the parameters for OLS1, which has a single sublayer (TL0), namely Level 0 (PTL0's L0(TL0) DPB Parameter 0 (DPB0's DPB0) and HRD Parameter 0 (HRD0's HRD0), are also described in PTL1 822, DPB1 842, and HRD1 862. To avoid repeating many parameters, one option is to derive the value of OLS1 from the DPB and HRD parameters that include more sublayers, but not DPB0 841 and HRD0 861. An example of this is shown in Figure 14.
[0184] Figure 14 shows an example of the definition of PTL, DPB, and HRD and sharing between different OLSs with sublayer information that is not relevant to some OLSs. Since PTL0 indicates that there is only one sublayer, only the parameters of TL0 of DPB1 and HRD1 are used.
[0185] Therefore, in another embodiment, when the OLS indexes the PTL structure, DPB structure, and HRD parameter structure by vps_ptl_max_temporal_id[i], vps_dpb_max_temporal_id[i], and vps_hrd_max_tid[i] respectively, there is a bitstream constraint that vps_dpb_max_temporal_id[i] and vps_hrd_max_tid[i] must be greater than or equal to vps_ptl_max_temporal_id[i], and that larger values corresponding to higher sublayers of the DPB and HRD parameters are ignored by the OLS.
[0186] Thus, according to a further embodiment, the encoder 10 is configured to form the OLS representation 18 and / or the video parameter set 81 (or generally, the multi-layer video bitstream 14) to be valid only for the OLS as long as the maximum temporal sublayer indicated by the decoder requirement set 82 associated with the OLS is less than or equal to the maximum temporal sublayer indicated by each of the buffer requirement set 84 and bitstream adaptation set 86 associated with the OLS, and the parameters in the buffer requirement set 84 and bitstream adaptation set 86 are the same for temporal layers at or below the maximum temporal sublayer indicated by the decoder requirement set 82 associated with the OLS.
[0187] In other words, the encoder 10 may provide the OLS indication 18 and / or the video parameter set 81 such that the maximum temporal sublayer indicated by the buffer requirement set 84 associated with the OLS is greater than or equal to the maximum temporal sublayer indicated by the decoder requirement set 82 associated with the OLS, and such that the maximum temporal sublayer indicated by the bitstream conformance set 86 associated with the OLS is greater than or equal to the maximum temporal sublayer indicated by the decoder requirement set 82 associated with the OLS.
[0188] 14 shows an example of a video parameter set 81 including a first decoder requirement set 821 for a first set of layers, such as layer L0, which includes access units of a first temporal sublayer TL0. The video parameter set 81 further includes a second decoder requirement set 822 that references a second set of layers, the second set of layers including layer L0 and layer L1, which includes access units of a first temporal sublayer and a second temporal sublayer, i.e., TL0 and TL1. Thus, the largest temporal sublayer of the layers in the second set is the second temporal sublayer TL1. The video parameter set 81 further includes a DPB parameter set 842 that references the second set of layers, a bitstream conformance set 862 that references the second set of layers, and a bitstream conformance set 863 that references a set of a third layer L2 having access units of a first temporal sublayer and a second temporal sublayer. A first OLS, OLS1, is associated with a first set of layers, and its decoder requirements are described by a first decoder requirement set 821, which indicates that the largest temporal sublayer of the first set of layers is the first temporal sublayer. Since the first temporal sublayer is less than or equal to the largest temporal sublayer indicated by the DPB parameter set 842 associated with OLS1 and the bitstream adaptation set 862 associated with OLS1 (note that temporal sublayers are hierarchically ordered), the DPB parameter set 842 and the bitstream adaptation set 862 contain information about the first set of layers. Thus, the DPB parameter set 842 and the bitstream adaptation set 862 are valid for OLS1 insofar as they relate to the first set of layers, and its access units belong to the first temporal sublayer. For example, as described with respect to FIG. 12 and also shown in FIGS. 13 and 14, the decoder requirement set 822, the buffer requirement set 842, and the bitstream adaptation set 862 include sets of parameters for each temporal sublayer included in the bitstream to which they refer.According to this embodiment, the parameters for the first temporal sublayer TL0 are valid for OLS1 because the first temporal sublayer is less than or equal to the maximum temporal sublayer indicated by the decoder requirement set 822.
[0189] In other words, for the OLS being decoded, the decoder 50 may use (for example) (only) the parameters of the decoder requirement set 82, buffer requirement set 84, and bitstream adaptation set 86 associated with the OLS that relate to a temporal sublayer that is equal to or smaller than the maximum temporal sublayer associated with the decoder requirement set 82 for the OLS.
[0190] FIG. 15 shows another example of a video parameter set 81 and an OLS indication 18. FIG. 15 illustrates possible alternatives when the parameters in TL1 and TL0 are the same, or when the value of TL1 is the maximum allowed for a level. In such a case, instead of including DPB0 and HRD0 in the VPS as shown previously, both can be included without including DPB1 and HRD1. The upper sublayer value of OLS1 can then be derived as equal to the value signaled in TL0 or the maximum allowed for the level. Thus, FIG. 15 illustrates the sublayer information that must be inferred if not present for a given OLS, with respect to the definitions of PTL, DPB, and HRD and their sharing between different OLSs.
[0191] Therefore, in another embodiment, there are no bitstream constraints on the values vps_ptl_max_temporal_id[i], vps_dpb_max_temporal_id[i], vps_hrd_max_tid[i], but for values greater than vps_ptl_max_temporal_id[i] where i>vps_dpb_max_temporal_id[i], the vps_hrd_max_tid[i] DPB and HRD parameters up to vps_ptl_max_temporal_id[i] shall be assumed to be the maximum value specified by the profile lever or equal to the highest signal DPB and HRD parameters.
[0192] Thus, according to another embodiment, encoder 10 is configured to form OLS display and / or video parameter set 81 such that the maximum temporal sublayer indicated by decoder requirement set 82 associated with the OLS is greater than or equal to the maximum temporal sublayer indicated by buffer requirement set 84 and bitstream adaptation set 86, respectively, associated with the OLS. According to these embodiments, parameters for temporal sublayers that are missing in buffer requirement set 84 and bitstream adaptation set 86 associated with an OLS, e.g., OLS2 in FIG. 15 , and that are greater than or equal to the maximum temporal sublayer indicated by buffer requirement set 84 and bitstream adaptation set 86, respectively, will be set equal to a fourth parameter or to a parameter in buffer requirement set 84 and bitstream adaptation set 86 associated with the OLS associated with the maximum temporal sublayer indicated by buffer requirement set 84 and bitstream adaptation set 86, respectively.
[0193] 1 , may be configured to set the OLS-related parameters in buffer requirement set 84 and bitstream conformance set 86 equal to a default value, such as the maximum value of the respective parameter, indicated in decoder requirement set 82, for each parameter in buffer requirement set 84 and bitstream conformance set 86, or to estimate a value of the respective parameter in buffer requirement set 84 or bitstream conformance set 86 associated with the OLS that is related to the maximum temporal sublayer indicated by buffer requirement set 84 and bitstream conformance set 86, respectively, if the maximum temporal sublayer indicated by decoder requirement set 82 associated with the OLS is greater than or equal to the maximum temporal sublayer, i.e., the OLS to be decoded, indicated by buffer requirement set 84 and bitstream conformance set 86, respectively. For example, the selection of whether to use a default value or whether to use a value of the respective parameter in buffer requirement set 84 or bitstream conformance set 86 associated with the OLS that is related to the maximum temporal sublayer indicated by buffer requirement set 84 and bitstream conformance set, respectively, may be made differently for each parameter in buffer requirement set 84 and bitstream conformance set 86.
[0194] 7. Selecting the Output Layer in the Region of Interest Application Section 7 describes an embodiment according to a seventh aspect with reference to Figure 1. Thus, the description of Figure 1 may optionally be applied to an embodiment according to the seventh aspect, and details described with respect to further aspects may optionally be implemented in embodiments described in this section.
[0195] Some embodiments according to the seventh aspect relate to derivation of PicOutputFlag in RoI applications.
[0196] When a multi-layer bitstream such as video bitstream 14 is used and a picture in a specified output layer is unavailable at the decoder side (e.g., due to a bitstream error or transmission loss), not following certain considerations may result in a suboptimal user experience. Typically, if an access unit does not contain a picture in an output layer, it is up to the implementation to select and output a picture from a non-output layer to compensate for the error or loss, as is clear from the notes in the derivation of the variable PicOutputFlag below.
[0197] The variable PictureOutputFlag for the current picture is derived as follows: - PictureOutputFlag is set to 0 if sps_video_parameter_set_id is greater than 0 and the current layer is not an output layer (i.e., nuh_layer_id is not equal to OutputLayerIdInOls[TargetOlsIdx][i] for any value of i in the range 0 to NumOutputLayersInOls[TargetOlsIdx]-1), or if any of the following conditions are true: - The current picture is a RASL picture and the associated IRAP picture's NoOutputBeforeRecoveryFlag is equal to 1. - The current picture is a GDR picture with NoOutputBeforeRecoveryFlag equal to 1 or is a recovery picture of a GDR picture with NoOutputBeforeRecoveryFlag equal to 1. - Otherwise, PictureOutputFlag is set equal to ph_pic_output_flag.
[0198] NOTE - In an implementation, a decoder may output pictures that do not belong to an output layer. For example, if there is only one output layer but in an AU, a picture in the output layer is not available, e.g., due to loss or layer down-switching, the decoder may set PictureOutputFlag equal to 1 for the picture with the highest value of nuh_layer_id and ph_pic_output_flag equal to 1 of all pictures in the AU available to the decoder, and set PictureOutputFlag equal to 0 for all other pictures in the AU available to the decoder.
[0199] However, if the bitstream is created for a region of interest (RoI) application, i.e., the higher layer depicts only a subset of the lower layer picture (through the use of a scaling window), then it is undesirable to switch between layers at the decoder output in a short time frame, since the switch between overview and detail display would be very fast. Therefore, as part of the present invention, in one embodiment, decoder implementations are not allowed to freely select the output layer when a scaling window that does not cover the entire picture plane is used, as follows:
[0200] In one implementation, the decoder can output pictures that do not belong to an output layer as long as the scaling window covers the complete picture plane. For example, if there is only one output layer but a picture in the output layer is unavailable in an AU, for example, due to loss or layer down-switching, the decoder can set PictureOutputFlag equal to 1 for the picture with the highest value of nuh_layer_id and ph_pic_output_flag equal to 1 among all pictures in the AU available to the decoder, and set PictureOutputFlag equal to 0 for all other pictures in the AU available to the decoder.
[0201] According to an embodiment of the seventh aspect, a decoder 50 for decoding a multi-layer video bitstream, e.g., multi-layer video bitstream 14 or sub-bitstream 12, is configured to use vector-based inter-layer prediction of a predicted picture 262 of a first layer 242 from a reference picture 261 of a second layer 241, with the prediction vector scaled and offset according to the relative sizes and positions of the predicted and reference pictures defined in the multi-layer video bitstream 14. For example, picture 262 of layer 242 of FIG. 1 may be encoded into the multi-layer video data stream 14 using inter-layer prediction from picture 261 of layer 241, e.g., picture 261 of the same access unit 221. According to the seventh aspect, the multi-layer video bitstream 14 may include an OLS indication 18 of the OLS indicating a subset of the layers of the multi-layer video bitstream 14, the OLS including one or more output layers including the first layer 241 and one or more non-output layers including the second layer.
[0202] If a given picture 262 of the first layer 242 of the OLS, for example picture 262, is lost, the decoder 50 according to the seventh aspect is configured to replace the given picture by a further given picture of the second layer 241 of the OLS that is in the same access unit 22 as the given picture, if the scaling window defined for the given picture coincides with a picture boundary of the given picture 262 and if the scaling window defined for the further given picture coincides with a picture boundary of the further given picture. If at least one of the scaling pictures defined for the given picture does not coincide with a picture boundary of the given picture, and if the scaling window defined for the given picture does not coincide with a picture boundary of the further given picture, the decoder 50 is configured to replace the given picture by other means or not at all.
[0203] 8. Further Embodiments In the previous section, some aspects have been described as features in the context of an apparatus, but it will be apparent that such description may also be considered as a description of corresponding features of a method. While some aspects have been described as features in the context of a method, it will be apparent that such description may also be considered as a description of corresponding features with respect to the functionality of the apparatus.
[0204] Some or all of the method steps may be performed by (or using) a hardware apparatus, such as, for example, a microprocessor, a programmable computer, or electronic circuitry. In some embodiments, one or more of the most important method steps may be performed by such an apparatus.
[0205] The inventive encoded image signal may be stored on a digital storage medium or transmitted over a transmission medium, such as a wireless or wired transmission medium, such as the Internet.
[0206] Depending on certain implementation requirements, embodiments of the present invention may be implemented in hardware or software, or at least partially in hardware or at least partially in software. This implementation may be performed using a digital storage medium, such as a floppy disk, DVD, Blu-ray, CD, ROM, PROM, EPROM, EEPROM, or FLASH memory, having electronically readable control signals that may be stored thereon and that cooperate (or may be able to cooperate) with a programmable computer system to perform the respective methods. Thus, the digital storage medium may be computer-readable.
[0207] Some embodiments according to the invention include a data carrier having electronically readable control signals that can cooperate with a programmable computer system to perform one of the methods described herein.
[0208] Generally, embodiments of the present invention may be implemented as a computer program product having program code operable to perform one of the methods when the computer program product runs on a computer, which may for example be stored on a machine-readable carrier.
[0209] Other embodiments comprise the computer program for performing one of the methods described herein, stored on a machine readable carrier.
[0210] In other words, an embodiment of the inventive methods is, therefore, a computer program having a program code for performing one of the methods described herein, when the computer program runs on a computer.
[0211] A further embodiment of the inventive method is therefore a data carrier (or digital storage medium, or computer-readable medium) comprising, recorded on it, the computer program for performing one of the methods described herein. The data carrier, digital storage medium, or recorded medium is typically tangible and / or non-transitory.
[0212] A further embodiment of the inventive method is, therefore, a data stream or a sequence of signals representing the computer program for performing one of the methods described herein, The data stream or the sequence of signals can be adapted to be transferred via a data communication connection, such as, for example, via the Internet.
[0213] A further embodiment comprises a processing means, such as for example a computer, or a programmable logic device, configured to or adapted to perform one of the methods described herein.
[0214] A further embodiment comprises a computer having installed thereon the computer program for performing one of the methods described herein.
[0215] Further embodiments according to the invention include an apparatus or system configured to transfer (e.g., electronically or optically) a computer program for performing one of the methods described herein to a receiver. The receiver may be, for example, a computer, a mobile device, a memory device, etc. The apparatus or system may, for example, include a file server for transferring the computer program to the receiver.
[0216] In some embodiments, a programmable logic device (e.g., a field programmable gate array) may be used to perform some or all of the functions of the methods described herein. In some embodiments, a field programmable gate array may cooperate with a microprocessor to perform one of the methods described herein. In general, the methods are preferably performed by any hardware apparatus.
[0217] The devices described herein may be implemented using a hardware device, or using a computer, or using a combination of a hardware device and a computer.
[0218] The methods described herein may be performed using a hardware apparatus, or using a computer, or using a combination of a hardware apparatus and a computer.
[0219] In the foregoing Detailed Description, it will be appreciated that various features are grouped together into examples for the purpose of streamlining the disclosure. This method of disclosure should not be interpreted as reflecting an intention that the claimed examples require more features than are expressly recited in each claim. Rather, as the following claims reflect, subject matter may lie in fewer than all features of a single disclosed example. Accordingly, the following claims are hereby incorporated into the Detailed Description, with each claim standing on its own as a separate example. While each claim may stand on its own as a separate example, and a dependent claim may refer to a specific combination with one or more other claims in the claim, it should be noted that other examples may include combinations of the dependent claim with the subject matter of each other dependent claim, or combinations of each feature with other dependent or independent claims. Such combinations are suggested herein unless it is stated that a particular combination is not intended. Furthermore, it is intended to include features of other independent claims, even if the claims are not directly dependent on those independent claims.
[0220] The above-described embodiments are merely illustrative of the principles of the present disclosure. It is understood that modifications and variations of the arrangements and details described herein will be apparent to those skilled in the art. It is therefore intended to be limited only by the scope of the appended claims and not by the specific details presented by way of description and illustration of the embodiments herein.
Claims
1. 1. A video decoder comprising: a processor; The processor: Obtaining a bitstream including a plurality of layers and access units; decoding an indication that the bitstream includes a plurality of output layer sets (OLS), each of the plurality of OLS having a unique number of the plurality of layers; determining that extrinsic information for identifying an OLS to be decoded from among the plurality of OLSs is unavailable; In response to the determination, for a first access unit among the access units, an OLS having a maximum number of layers among the plurality of layers is selected from the plurality of OLSs as a decoding target. Video decoder.
2. the processor is further configured to decode the selected OLS.
2. The video decoder of claim 1.
3. The processor, in order to select the OLS, performs the following steps from among the plurality of OLSs: has the largest number of layers among the plurality of layers, and having the smallest OLS index value, and further configured to select the OLS for decoding.
2. The video decoder of claim 1.
4. In order to select the OLS to be decoded, the processor Identifying a set of OLSs from the plurality of OLSs that has the largest number of layers from the plurality of layers; Identifying the OLS from the set that has the smallest OLS index value; and selecting the OLS having the identified smallest OLS index value as the OLS to be decoded.
4. A video decoder according to claim 3.
5. The external information is external information that is not included in the bitstream.
2. The video decoder of claim 1.
6. 1. A video decoding method, comprising: obtaining a bitstream including a plurality of layers and access units; decoding an indication that the bitstream includes a plurality of Output Layer Sets (OLS), each of the plurality of OLS having a unique number of the plurality of layers; determining that extrinsic information is unavailable to identify an OLS to decode from among the plurality of OLSs; and In response to the determination, selecting, for a first access unit among the access units, an OLS having a maximum number of layers among the plurality of layers from the plurality of OLSs as a decoding target; Including, Video decoding methods.
7. further comprising the step of decoding the selected OLS. The video decoding method of claim 6.
8. When selecting the OLS, from among the plurality of OLSs, has the largest number of layers among the plurality of layers, and having the smallest OLS index value, further comprising the step of selecting the OLS to be decoded; The video decoding method of claim 6.
9. The step of selecting the OLS as a decoding target includes: Identifying a set of OLSs from the plurality of OLSs that has the largest number of layers among the plurality of layers; identifying an OLS from the set that has the smallest OLS index value; selecting the OLS having the identified smallest OLS index value as the OLS to be decoded; Including, The video decoding method of claim 8.
10. The external information is external information that is not included in the bitstream. The video decoding method of claim 6.
11. 11. A method for implementing a method according to any one of claims 6 to 10, wherein said method comprises: A computer-readable storage medium.
12. When run on a computer or signal processor, causes the method of any one of claims 6 to 10 to be carried out. Computer program.
Citation Information
Patent Citations
Image decoder and image encoder
JP2015195543A
Conformance parameters for bitstream segmentation
JP2017522779A