Processing of the output layer set of the symbolized video
The OLS concept enables efficient extraction of sub-bitstreams from multi-layer video bitstreams, addressing inefficiencies in existing technologies by optimizing decoder resource utilization and reducing signaling overhead.
Patent Information
- Application Number
- JP2022570633
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-05-22
- Filing Date
- 2021-05-20
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2041-05-20
AI Technical Summary
Existing video encoding technologies struggle with inefficient extraction of sub-bitstreams from multi-layer video bitstreams, leading to unnecessary resource utilization and high signaling overhead, particularly in scenarios involving hierarchical coding for temporal, fidelity, and spatial scalability.
The concept of an output layer set (OLS) is introduced to enable precise extraction of sub-bitstreams, ensuring efficient decoder resource utilization by accurately defining which parts of the video bitstream are extracted, and reducing signaling overhead through improved trade-offs.
This approach allows for accurate extraction of sub-bitstreams, optimizing decoder resource usage and reducing unnecessary decoding times by ensuring only necessary parts are processed, thereby enhancing efficiency and reducing signaling overhead.
Smart Images

Figure 0007711102000006 
Figure 0007711102000007 
Figure 0007711102000008
Abstract
Description
Technical Field
[0001] Description Embodiments of the present invention relate to an apparatus for encoding a video into a video bitstream, an apparatus for decoding a video bitstream, and an apparatus for processing a video bitstream, for example, an apparatus for extracting a bitstream such as a sub-bitstream from a video bitstream. Further embodiments relate to an encoding method, a decoding method, a processing method (e.g., an extraction method), and a video bitstream. Further embodiments relate to a video bitstream.
Background Art
[0002] In the newly emerged VVC codec, it is assumed that hierarchical coding for temporal, fidelity, and spatial scalability is supported from the beginning. That is, an encoded video bitstream structured into so-called layers and (temporal) sub-layers, and encoded picture data corresponding to a time, that is, a so-called access unit (AU), can include pictures that can predict each other within each layer, and some of which are output after decoding. The concept of the so-called output layer set (OLS) indicates to the decoder the reference relationship and which layer is output when the bitstream is decoded. The OLS can also be used to identify the corresponding HRD-related timing / buffer information in the form of SEI messages of buffering period, picture timing, and decode unit information carried in a bitstream encapsulated in a so-called scalable nesting SEI message.
Summary of the Invention
[0003] It is desirable to have a concept for processing an output layer set that enables extraction of a sub-bitstream from a video bitstream, the concept providing an accurate definition of the sub-bitstream extractable by the output layer set (in terms of precisely describing which parts of the video bitstream are extracted), efficient utilization of decoder resources (in terms of avoiding extraction of unnecessary parts or providing accurate information regarding decoder settings or requirements for decoding the selected sub-bitstream), and an improved trade-off between small signaling overhead.
[0004] A first aspect according to the present invention provides a concept for displaying, extracting, and / or decoding a random-accessible sub-bitstream from a multi-layer video bitstream. According to the first aspect, the random-accessible sub-bitstream to be extracted selectively includes a bitstream portion associated with the output layer of the random-accessible sub-bitstream indicated by the output layer set display of the random-accessible sub-bitstream, among the bitstream portions of the access units of the multi-layer video bitstream, or a bitstream portion necessary for decoding a random-accessible bitstream portion of the output layer.
[0005] A second aspect of the present invention provides a concept of a multi-layer video bitstream having a plurality of layers and a plurality of temporal layers. The multi-layer video bitstream comprises a display of an output layer set including one or more layers of the multi-layer video bitstream and a reference layer display indicating layer-to-layer references of the layers of the output layer set. The multi-layer video bitstream includes a display, such as a temporal layer display or an in-temporal-layer display, that enables identification of the bitstream portions of the layers of the output layer set belonging to the output layer set, in combination with the way the multi-layer video bitstream is encoded. This concept enables the bitstream portions of the OLS to be identified by the type of the bitstream portion and / or by the layer dependencies between the OLS indicated by the reference layer display. Thus, embodiments of the second aspect enable accurate extraction of sub-bitstreams while avoiding unnecessarily high signaling overhead.
[0006] A third aspect of the present invention provides a concept that enables a decoder that decodes a video bitstream to determine an output layer set to be decoded based on the attributes of the video bitstream provided to the decoder. Thus, this concept enables the decoder to select an OLS without an instruction to the decoder as to which OLS to decode. A decoder that can select an OLS without an instruction can ensure that the bitstream decoded by the decoder meets the level requirements known to the decoder, for example, by a display within the video bitstream.
[0007] A fourth aspect of the present invention provides a concept for extracting a sub-bitstream from a multi-layer video data stream, and within the extracted sub-bitstream, one of a set of predetermined bitstream partial types or picture types (e.g., randomly accessible or independently encoded bitstream partial type picture types), e.g., the same, access units that consist of only pictures or bitstream parts are indicated by a sequence start indicator even if each access unit is not a sequence start access unit in the original multi-layer video bitstream from which the sub-bitstream was extracted. Thus, the frequency of sequence start access units in the sub-bitstream may be higher than that of the multi-layer video data stream, and accordingly, the decoder may benefit from being able to avoid an unnecessarily long waiting time until the decoding of the video sequence can start by making more sequence start access units available.
[0008] A fifth aspect of the present invention provides a concept for extracting a sub-bitstream from a multi-layer video bitstream such that the sub-bitstream consists of only pictures belonging to one or more temporal sub-layers associated with an output layer set that describes the sub-bitstream to be extracted. For this purpose, syntax elements in the multi-layer video bitstream are used to indicate a predetermined temporal sub-layer of the OLS in a way that identifies different states including a state where the predetermined temporal sub-layer is below the maximum of the temporal sub-layers within an access unit in which at least one picture of a subset of the layers exists. By not transferring unnecessary sub-layers of the multi-layer video bitstream, the size of the sub-bitstream may be reduced, and the requirements of the decoder for decoding the sub-bitstream may be reduced.
[0009] According to the embodiment, decoder function-related parameters of a picture of a temporal sublayer belonging to an OLS that describes a sub-bitstream are signaled in the sub-bitstream and / or a multi-layer video data stream. Therefore, when determining decoder-related functional parameters, pictures that do not belong to the OLS can be omitted, so that the decoder function can be utilized efficiently.
[0010] A sixth aspect of the present invention provides a concept for processing a temporal sublayer in signaling video parameters for an output layer set of a multi-layer video bitstream. According to the embodiment, the OLS is associated with one of one or more bitstream-conformant sets signaled in the video bitstream, one of one or more buffer requirement sets, and one of one or more decoder requirement sets, and each of the bitstream-conformant set, the buffer requirement set, and the decoder requirement set is valid for one or more temporal sublayers indicated by constraints of a maximum temporal sublayer (e.g., hierarchically ordered temporal sublayers). The embodiment provides a concept of the relationship among the bitstream-conformant set, the buffer requirement set, and the decoder requirement set associated with the OLS with respect to the maximum temporal sublayer with which they are associated, so that a decoder can easily determine the parameters of the OLS associated with the bitstream-conformant set, the buffer requirement set, and the decoder requirement set. For example, the embodiment may enable the decoder to conclude that the parameters given in the bitstream-conformant set, the buffer requirement set, and the decoder requirement set of the OLS are completely valid for the OLS. In other embodiments, it enables the decoder to conclude to what extent the parameters given in the bitstream-conformant set, the buffer requirement set, and the decoder requirement set are valid for the OLS.
[0011] According to an embodiment, the maximum temporal sublayer indicated by the decoder requirement set associated with the OLS is less than or equal to the maximum temporal sublayer indicated by each of the buffer requirement set and the bitstream compliance set associated with the OLS, and the parameters within the buffer requirement set and the bitstream compliance set are the same for the temporal layer equal to or lower than the maximum temporal sublayer indicated by the decoder requirement set associated with the OLS, and are only valid for the OLS. As a result, when the maximum temporal sublayer indicated by the decoder requirement set associated with the OLS is less than or equal to the maximum temporal sublayer indicated by each of the buffer requirement set and the bitstream compliance set associated with the OLS, the decoder can presume that the parameters of the buffer requirement set and the bitstream compliance set associated with the OLS are valid for the OLS only as long as they are equal to the maximum temporal sublayer indicated by the decoder requirement set associated with the OLS and relate to the temporal layers below it. Therefore, the embodiment enables the decoder to determine the video parameters of the OLS based on the indication regarding the constraint of the maximum temporal sublayer signaled for each set of parameters, and can avoid a complex analysis of the OLS and the video parameters. Further, in this concept, since the OLS can be associated with the buffer requirement set and the bitstream compliance set, it is possible to signal a constraint of a maximum temporal sublayer that is greater than the constraint of the decoder requirement set associated with the OLS. Since the constraint of the maximum temporal sublayer is greater than the constraint of the decoder requirement set associated with the OLS, the signaling of the dedicated buffer requirement set and the dedicated bitstream compliance set related to the same maximum temporal sublayer as the decoder requirement set is omitted, and the overhead of the signaling of the video parameter set can be reduced.
[0012] The seventh aspect of the present invention provides a concept for handling loss of pictures in a multi-layer video bitstream, for example due to bitstream errors or transmission losses, where pictures are encoded using inter-layer prediction. If a picture that is part of the first layer is lost, that picture is replaced by another picture in the second layer, and the picture in the second layer is used for inter-layer prediction of the picture in the first layer. This concept includes replacing a picture with a further picture based on the coincidence of the scaling window defined for the picture and the picture boundary of the picture, and the coincidence of the scaling window defined for a further picture and the picture boundary of the further picture. When the scaling window defined for a picture coincides with the picture boundary of the picture, and further when the scaling window defined for the picture coincides with the picture boundary of a further picture, replacing the picture with a further picture may not cause a change in the display window of the presented content, for example a change from a detailed view to an overview. Further embodiments and advantageous aspects of the present disclosure are described in more detail below with respect to the figures.
Brief Description of the Drawings
[0013]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
[0014] The following describes embodiments in detail, but it should be understood that the embodiments provide many applicable concepts that can be embodied in a variety of video encoding concepts. The specific embodiments described are merely illustrative of specific ways to implement and use the present concept and do not limit the scope of the embodiments. In the following description, a number of details are presented to provide a more complete description of the embodiments of the present invention. However, it will be apparent to those skilled in the art that other embodiments can be practiced without these specific details. In other instances, well-known structures and devices are shown in block diagram form rather than in detail to avoid obscuring the examples described herein. Further, the features of the different embodiments described herein can be combined with each other unless otherwise specified.
[0015] In the following description of the embodiments, elements that are the same or similar, or have the same function, are either given the same reference numerals or are identified by the same names, and the repeated description of elements that are given the same reference numerals or are identified by the same names is usually omitted. Thus, the descriptions provided for elements having the same or similar reference numerals, or identified by the same name, are interchangeable with each other or can be applied to each other in different embodiments.
[0016] 0. Encoder 10, Extractor 30, Decoder 50, and Video Bitstreams 12, 14 According to FIG. 1 The embodiments described in this section provide examples of frameworks that can incorporate embodiments of the present invention. Hereinafter, an explanation of embodiments of the concepts of the present invention will be presented together with an explanation of how such concepts can be incorporated into the encoder and extractor of FIG. 1. However, the embodiments described with respect to FIG. 2 and those that follow may be used to form encoders and extractors that do not operate according to the framework described with respect to FIG. 1. Furthermore, it should be noted that the encoder, extractor, and decoder, although illustrated together in FIG. 1 for example purposes, may be implemented separately from each other. It should be further noted that the extractor and decoder can be combined within one device, or one of the two can be implemented as part of the other.
[0017] Figure 1 shows an example of an encoder 10, an extractor 30, a decoder 50, a video bitstream 14 (also referred to as a video data stream or a data stream), and a sub-bitstream 12. The encoder 10 is for encoding a video sequence 20 into the video bitstream 14. The encoder 10 encodes the video sequence 20 into the video bitstream 14 in units of pictures 26, and each of the pictures 26 belongs to a time, for example, a frame of the video sequence. Encoded video data belonging to a common time may be called an access unit (AU) 22. In Figure 1, exemplary operations of three access units 221, 222, 223 of the video sequence 20 are shown. Note that the description referring to the access unit 22 may refer to any of the exemplary access units 221, 222, 223. Each of the access units 22 contains or encodes one or more pictures 26, and each of the pictures 26 is associated with one of a plurality of layers of the video bitstream 14. In Figure 1, examples of the pictures 26 are represented by the picture 261 and the picture 262. The picture 261 is associated with the first layer 241 of the video bitstream 14, and the picture 262 is associated with the second layer 242 of the video bitstream 14. In Figure 1, each of the access units 22 contains pictures for both the first layer 241 and the second layer 242, but note that the video bitstream 14 can contain access units 22 that do not necessarily contain pictures 26 for each of the layers 24 of the video bitstream 14. Further, the video bitstream 14 can contain additional layers in addition to the first and second layers shown in Figure 1. The encoder 10 is configured to encode each of the pictures 26 into one or more bitstream portions of the video sequence 14, for example, NAL units. For example, each of the bitstream portions 16 into which the picture 26 is encoded can encode a part of the picture 26, such as a slice of the picture 26.The bitstream portion 16 where picture 26 is actually encoded is sometimes called a video coding layer (VCL) NAL unit. The bitstream 14 can further include descriptive data indicating information describing the encoded video data, such as non-VCL NAL units. For example, each of the access units 22 can include a bitstream portion signaling the descriptive data of each access unit in addition to the bitstream portion signaling the decoded video data. The video bitstream 14 can further include descriptive data referring to a plurality of access units, or a part of one or more access units. For example, an output layer set (OLS) indication 18 indicating one or more output layer sets can be encoded in the video bitstream 14.
[0018] The OLS may be a display of sub-bitstreams extractable from the video bitstream 14. The OLS can indicate one or more or all of the multiple layers of the video bitstream 14 as the output layer of the sub-bitstream described by each OLS. Note that the set of layers indicated by the OLS may not necessarily be a proper subset of the layers of the video bitstream 14. In other words, all layers of the video bitstream 14 may be included in the OLS. The OLS may optionally further include a description of the sub-bitstream described by the OLS and / or decoder requirements for decoding the sub-bitstream indicated by the OLS. Note that the sub-bitstream described by the OLS may be defined by additional parameters other than layers, such as a temporal sublayer or a subpicture. For example, picture 26 of layer 24 may be associated with one of one or more temporal sublayers of layer 24. A temporal sublayer can include pictures at times associated with each temporal sublayer. For example, the pictures of the first temporal sublayer may be associated with times that form a sequence at the first frame rate, and the pictures of the second temporal sublayer may be associated with times located between the times with which the pictures of the first temporal sublayer are associated such that the combination of the first and second temporal sublayers can provide a video sequence having a higher frame rate than in the case of one of the first and second temporal sublayers. The OLS may optionally indicate a temporal sublayer for describing which bitstream portion or picture 26 belongs to the sub-bitstream described by the OLS. The temporal sublayers of the bitstream or the encoded video sequence can be hierarchically ordered, for example, by indexing. For example, the hierarchical order may mean that decoding of the pictures of the bitstream including a particular temporal sublayer requires all temporal sublayers lower in the hierarchical order.
[0019] Note that the OLS can include one or more output layers, and optionally one or more non-output layers as well. In other words, the OLS can designate one or more of the layers included in the OLS as the output layer(s) of the OLS, and optionally designate one or more of the layers of the OLS as non-output layer(s). For example, a layer that includes a reference picture for a picture of the output layer of the OLS may be included in the OLS as a non-output layer because a picture of the non-output layer may be required to decode the picture of the output layer of the OLS.
[0020] The OLS can further include level information regarding the bitstream described by the OLS, and the level information indicates or is associated with one or more bitstream constraints such as the maximum value of one or more of bitrate, picture size, and frame rate.
[0021] Optionally, the bitstream 14 can further include an extractability indication 19 of the OLS. For example, the extractability indication can be part of the OLS indication. The extractability indication can indicate a (not necessarily appropriate) subset of the bitstream portion 16 that forms a decodable sub-bitstream associated with the OLS. That is, the extractability indication can indicate which of the bitstream portion 16 belongs to the OLS.
[0022] The picture 26 can be encoded in the video bitstream 14 with reference to other pictures, for example, for prediction of residuals, motion vectors, and / or syntax elements. For example, the picture can refer to another picture in the same access unit (referred to as the reference picture of the picture), and the reference picture can be associated with another layer (referred to as an inter-layer reference picture). Additionally or alternatively, the picture can refer to a reference picture that is part of the same layer but a different access unit from the picture.
[0023] Extractor 30 can receive video bitstream 14 and can select one or more OLSs shown in video bitstream 14, for example, based on display 32 provided to extractor 30. Extractor 30 can provide sub-bitstream 12 indicated by the selected OLS by transferring at least bitstream portion 16 belonging to the selected OLS in sub-bitstream 12. Note that since extractor 30 can modify or adapt one or more of bitstream portions 16, the transferred bitstream portion does not necessarily exactly correspond to bitstream portion 16 signaled in video bitstream 14. In FIG. 1, an apostrophe is used for the bitstream portion of sub-bitstream 12, for example reference numeral 16', to indicate a potential change in the bitstream portion of video bitstream 14 when transferred to sub-bitstream 12.
[0024] Sub-bitstream 12 can be decoded by decoder 50 to obtain the decoded video sequence represented by sub-bitstream 12. Note that the decoded video sequence can be different from video sequence 20 in that the decoded video sequence can optionally represent only a part of video sequence 20 in terms of resolution, fidelity, frame rate, picture size, and video content (when focusing on the extraction of sub-pictures), and in addition to this fact, the decoded video sequence can have distortion due to quantization loss.
[0025] Picture 26 of video sequence 20 may include a separately encoded picture that does not reference pictures of other access units. That is, for example, a separately encoded picture (with respect to temporal prediction, a separately encoded picture may be encoded using layer - to - layer prediction optionally) is encoded into video bitstream 14 without using inter - prediction. Due to separate encoding, the decoder can start decoding the video sequence at the access unit of the separately encoded picture. A separately encoded picture may be called an instantaneous random access point (IRAP). Examples of IRAP pictures are IDR and CRA pictures. In contrast, a trailing picture may refer to a picture that references a picture of another access unit that may precede the trailing picture encoding order (the order in which picture 26 is encoded into video stream 14). The bitstream portion 16 in which a separately encoded picture is encoded may be called a separately encoded bitstream portion, for example, an IRAP NAL unit, while the bitstream portion 16 in which a dependently encoded picture 26 is encoded may be called a dependent bitstream portion, for example, a non - IRAP NAL unit. Further, it should be noted that not all bitstream portions of a picture are necessarily encoded in the same way, beyond separate and dependent encoding. For example, the first portion of picture 26 of the first access unit, for example, access unit 221, may be separately encoded, and the second portion of picture 26 of the first access unit may be dependently encoded. In this case, in picture 26 of the second access unit, such as access unit 222, the first portion of picture 26 of the second access unit may be dependently encoded, and the second portion may be separately encoded. In this way, the higher data rate of separate encoding over dependent encoding can be distributed across multiple access units.Such encoding may be called general decoder refresh (GDR) because the decoder 50 may have to decode some access units before decoding an entire picture that is independent of the access units preceding the GDR cycle, i.e., a sequence of pictures in which independently encoded portions covering the entire picture are distributed.
[0026] In the following, some concepts and embodiments will be described with respect to FIG. 1. It is pointed out that the features described with respect to the encoder, video bitstream, extractor, or decoder are to be understood as also being descriptions of other ones of these entities. For example, a feature described as being present in a video data stream is to be understood as a description of an encoder configured to encode this feature into a video bitstream and a decoder or extractor configured to read the feature from the video bitstream. It is further pointed out that the estimation of information based on the representation encoded in the video bitstream can be performed equally on the encoder side and the decoder side. It is further noted that the aspects described in the following sections can be combined with each other.
[0027] 1. Random-accessible sub-bitstream representation In this section, embodiments according to the first aspect will be described with reference to FIG. 1. The details described in Section 0 can be optionally applied to the embodiments according to the first aspect. Also, the details described with respect to further aspects can be optionally implemented in the embodiments described in this section.
[0028] The randomly accessible bitstream part may refer to an independently encoded bitstream part as described with respect to FIG. 1. Accordingly, a randomly accessible picture (or bitstream part) can be referred to as an independently encoded picture (or bitstream part) as described with respect to FIG. 1.
[0029] Some embodiments according to the first aspect can refer to the full IRAP level display for non-aligned IRAP. The embodiments can refer to the impact of IRAP alignment where max_tid_il_ref_pics_plus1 == 0 (e.g., reference to only IDR, or IDR, CRA, or GRD where ph_recovery_poc_cnt is equal to 0, or reference to only one of them), ph_recovery_poc_cnt is equal to 0).
[0030] FIG. 2 shows an example of an output layer set of a video sequence such as video sequence 20. In other words, the video sequence of FIG. 2 can represent a video sequence formed by the layers of an OLS including output layer L1 and non-output layer L0. In access unit 22*, the multi-layer OLS bitstream of FIG. 2 includes an example of non-aligned IRAP. That is, access unit 22* includes a non-randomly accessible bitstream part in one of the layers of the OLS, i.e., L1 in FIG. 2.
[0031] Display of level information for the entire IRAP sub-bitstream, i.e., the bitstream containing the result of dropping all non-IRAP NAL units from the bitstream, can be useful for trick mode playback such as fast forward based on IRAP pictures only. This level information refers to the level_idc display that refers to a list of restrictions defined for parameters such as level_idc, maximum picture size, maximum picture rate, maximum bitrate, maximum buffer size, maximum slices / tiles / sub-pictures per picture, and minimum compression ratio. However, in the case of multi-layer, it is not uncommon for IRAP pictures not to be aligned across layers. For example, since the IRAP distance is longer for the upper (dependent) layer, the IRAP frequency is not as high in the upper layer as in the lower (reference) layer. This is shown in Figure 2. Here, the upper layer L1 of the illustrated OLS with two layers contains a trailing NAL unit at the position of POC (Picture Order Count) 3, i.e., within access unit 22*, while the lower layer L0 contains an IRAP NAL unit at the same position.
[0032] FIG. 3 shows an example of all the extracted IRAP sub-bitstreams that can be executed in the conventional manner and are extracted from the multi-layer video bitstream of FIG. 2. In light of the fact that not all layers of the OLS are output by the decoder, for example, in the example of FIG. 2, only the upper layer L1 is marked as the output layer. Therefore, unless there is an IDR at the corresponding position of the upper (output) layer, maintaining the lower (non-output) layer IDR in the bitstream, as in the case of picture 260*, would be a waste of decoder resources. The reason is that regardless of whether the L0 IDR of POC3 is in the all-IRAP sub-bitstream, the decoder output is the same, so when the decoder decodes the all-IRAP sub-bitstream of POC3, it will not output any pictures at all. Furthermore, when the all-IRAP level display is used to estimate the maximum playback speed of the all-IRAP presentation (for example, in relation to the level and playback speed of the complete OLS bitstream), decoding the above-mentioned L0 IRAP of POC3 will result in a decrease in the achievable maximum playback speed of the all-IRAP sub-bitstream.
[0033] Therefore, omitting the decoding / dropping from such a bitstream and thereby also excluding the all-IRAP NAL units of the non-output layers in the access units that do not have IRAP NAL units in all the corresponding output layers in the OLS of such all-IRAP sub-bitstreams from consideration based on their respective level displays is part of an embodiment of the present invention.
[0034] According to an embodiment of the first aspect, the video bitstream 14 represents an encoded video sequence 20, as described with respect to FIG. 1 for example. The video bitstream 14 includes a sequence of access units 22, each of which includes one or more bitstream portions 16, and each of them is associated with one of a plurality of layers 24 of the video bitstream 14. Each of the bitstream portions 16 is one of the bitstream portion types including a randomly accessible bitstream portion type, such as an independently encoded bitstream portion type like the IRAP type. The video bitstream 14 includes, for example, an OLS display 18 of the OLS of the video bitstream 14 to be detected by the encoder 10 and an extractability display 19 of a randomly accessible sub-bitstream described by the OLS, where the OLS includes one or more output layers and one or more non-output layers. For example, the randomly accessible sub-bitstream can be an all-IRAP sub-bitstream. For example, the OLS display can include a level display for the all-IRAP sub-bitstream. It should be noted that the term all-IRAP sub-bitstream should not be understood as necessarily exclusively including randomly accessible bitstream portions or IRAP bitstream portions. Rather, the randomly accessible sub-bitstream, or the all-IRAP sub-bitstream, may in some cases include non-IRAP bitstream portions or non-randomly accessible bitstream portions, such as the bitstream portion of a reference picture of a randomly accessible bitstream portion. In other examples, the randomly accessible sub-bitstream may include only randomly accessible bitstream portions.
[0035] According to the embodiment, since the encoder 10 provides the video bitstream 14, for each layer of the OLS, for each access unit 22 beyond the bitstream portion of each access unit 22, if the bitstream portion of all output layers includes one of the bitstream portions that each access unit can randomly access, it is a bitstream portion that can be randomly accessed.
[0036] In other words, in one embodiment, it is a requirement for bitstream compliance that the bitstream indicated by all IRAP level displays does not include access units without output pictures in all output layers.
[0037] There are use cases where not all output layers have pictures for all access units of the original bitstream. For example, it can be stereoscopic video with different frame rates for the eyes. In such cases, the bitstream requirements are not so strict.
[0038] In another embodiment, it is a requirement for bitstream compliance that the bitstream indicated by all IRAP level displays does not include access units without output pictures in all output layers.
[0039] Therefore, according to the embodiment, since the encoder 10 according to the first aspect provides the video bitstream 14, for each layer indicated by the OLS of the randomly accessible sub-bitstream, for each access unit 22 beyond the bitstream portion of each access unit 22, if the bitstream portion 16 of at least one of the output layers among the output layers includes one of the bitstream portions that each access unit can randomly access, it is a bitstream portion that can be randomly accessed.
[0040] As both results (the randomly accessible bitstream portions of at least one or all output layers), access units that do not satisfy the constraints of the bitstream are not created on the encoder side or are dropped during extraction. In other words, a bitstream containing only IRAP-level representations has the constraint that there are no AUs having IRAP in non-output layers and non-IRAP NAL units in pictures at the same temporal position in the output layers.
[0041] Alternatively, for the randomly accessible sub-bitstreams described by the OLS representation 18 and the extractability representation 19 (the OLS representation and the extractability representation are encoded in the video bitstream 14 by the encoder 10), for each access unit, for each bitstream portion of each access unit, if one of the following two conditions is satisfied, the respective bitstream portion is selectively included. The first condition is satisfied when each bitstream portion is a randomly accessible bitstream portion and each bitstream portion is associated with one of one or more output layers, for example, the bitstream portion of picture 26* in FIG. 2. The second condition is satisfied when each bitstream portion is associated with a reference layer of one of the output layers and each bitstream portion is associated with one of one or more non-output layers, and furthermore, beyond the bitstream portion of each access unit, at least one bitstream portion of the output layer is a randomly accessible bitstream portion (applies to the bitstream portion of picture 26** in FIG. 2, and the corresponding access unit includes the bitstream portion of picture 26* which is part of the output layer of OLS), or instead, beyond the bitstream portion of each access unit, all bitstream portions of the output layers are randomly accessible bitstream portions (also applies to the case of picture 26** in FIG. 2).
[0042] Thus, an apparatus for extracting a sub-bitstream 12 from a video bitstream 14, such as apparatus 30 of FIG. 1 according to a prior embodiment of the first aspect, is configured to provide the sub-bitstream 12 as indicated by an OLS representation of a randomly accessible sub-bitstream.
[0043] In other words, as an alternative to an optional bitstream constraint that the bitstream indicated by the full IRAP level representation does not contain access units without output pictures in all output layers, according to an embodiment, since the indicated level does not contain AUs in which this IRAP NAL unit and non-IRAP NAL units are "mixed", if a bitstream having only IRAP at the indicated level is desired, such AUs need to be dropped.
[0044] A similar case is considered when there is an IRAP NAL unit in the output layer but not in the reference layer. In a further embodiment, as long as the output layer has an IRAP NAL unit, NAL units within the simultaneous reference layer are not dropped and are considered for the indicated level. For this to work, there is a bitstream constraint that the temporal reference of the reference layer located in the same place without an IRAP NAL unit only references pictures that are also contemporaneous with the IRAP NAL unit of the output layer. Alternatively, the indicated level applies only to AUs in which all NAL units are IRAP NAL units, and all others are discarded if such a bitstream at such a level (IRAP only) is considered.
[0045] In a further embodiment, instead of referring to the layer selected for OLS, the requirement that only AUs having all NAL unit types of the IRAP type are considered for IRAP level indication is applied to the entire bitstream. In such a case, since such an AU contains an access unit delimiter with aud_irap_or_gdr_au_flag equal to 1, in the case where there is an IRAP NAL unit (i.e., not considering the case of GDR), the presence of an access unit delimiter with aud_irap_or_gdr_au_flag equal to 1 is used to determine whether the AU is subject to IRAP-only level indication.
[0046] According to an example of the first aspect, the video bitstream 14 includes a level indication of a randomly accessible sub-bitstream, for example, a randomly accessible sub-bitstream described as extractable according to the extraction possibility information 19. The level indication (also referred to as level information) can indicate a level related to the bitstream constraint, for example, by pointing to a list of levels, as described in section 0. For example, the level indication is associated with one or more of the CPD size, DPB size, picture size, picture rate, minimum compression ratio, picture segmentation limit (e.g., tile / slice / sub-picture), (e.g., HRD timing such as access unit / DU removal time, DPB output time).
[0047] In other words, in addition to the level indication, when considering an extracted bitstream having only IRAP access units, additional parameters are relevant. Such parameters are DPB parameters and HRD parameters.
[0048] According to an embodiment of the first aspect, a decoder such as decoder 50 is configured to check whether the picture buffer conforms to a sub-bitstream that is randomly accessible according to the extractability information 19. For example, in addition to the above-described parameters, decoder 50 can check the level display in the extractability information, HRD parameters, and DPB parameters. The picture buffer may refer to an encoded picture buffer and / or a decoded picture buffer of the decoder. Optionally, decoder 50 can be configured to derive, from video bitstream 12, for example, timing information for the picture buffer, timing information for a randomly accessible sub-bitstream indicated by OLS display 18. Decoder 50 can decode a randomly accessible sub-bitstream based on the timing information.
[0049] In other words, in practice, all such IRAP variants that omit IRAPs in non-output layers without IRAPs in the output layer within the same access unit also omit the decoding of non-output pictures that are not used for the aforementioned references, thus enabling a reduction in DPB requirements (i.e., DPB size in terms of picture slots). In particular, this is a separate part of the bitstream level limitation and has no direct relation to the limitation defined by the level_idc of the bitstream. Together with the picture size of the bitstream, the level sets a limit on the maximum number of pictures that can be held in the DPB. On the other hand, the DPB parameters include more information such as the maximum number of picture reorderings when outputting a picture, i.e., the number of pictures that can be before another picture in the output order but can follow it in the output order. Such information may be different when the extracted bitstream contains only IRAP pictures. Therefore, signaling additional DPB parameters for this representation is part of the present invention to enable the decoder to utilize its resources more effectively. An embodiment of the present invention is shown in Table 1 below.
Table 1-1
Table 1-2
[0050] For example, vps_ols_dpb_params_all_irap_idx[i] specifies the index to the list of dpb_parameters() syntax structures in the VPS for the dpb_parameters() syntax structure applied to the i-th multi-layer OLS when considering only the IRAP sub-bitstream. If it exists, the value of vps_ols_dpb_params_idx[i] shall be in the range from 0 to VpsNumDpbParams-1 (inclusive of both ends). If vps_ols_dpb_params_all_irap_idx[i] does not exist, this is assumed to be equal to vps_ols_dpb_params_idx[i]. For a single-layer OLS, the corresponding dpb_parameters() syntax structure exists in the SPS referenced by the layer within the OLS. Each dpb_parameters() syntax structure in the VPS shall be referenced by at least one of the values of vps_ols_dpb_params_idx[i] or vps_ols_dpb_params_all_irap_idx[i] for i in the range from 0 to NumMultiLayerOlss-1 (inclusive of both ends).
[0051] As pointed out, the additional information that may be required for the extracted bitstream containing only the IRAP NAL unit is the HRD parameter. The HRD parameter can include, for example, one or more or all of the necessary CPB size, the time at which the access unit is removed from the CPB, the bitrate at which the CPB is supplied, or whether the resulting bitstream after extraction corresponds to a constant bitrate representation.
[0052] 2. Reference Picture Alignment In Section 2, an embodiment according to a second aspect will be described with reference to FIG. 1, and the details described in Section 0 can be optionally applied to the embodiment according to the second aspect. Also, the details described regarding further aspects can be optionally implemented in the embodiment described in this section.
[0053] In VVC, the output layer set defines the prediction dependencies between the layers of the bitstream. The syntax element vps_max_tid_il_ref_pics_plus1[i][j] signaled for all the direct reference layers of a given layer can further limit the amount of pictures of the reference layers used for prediction as follows.
[0054] A vps_max_tid_il_ref_pics_plus1[i][j] equal to 0 specifies that for decoding the picture of the i-th layer, the picture of the j-th layer which is neither an IRAP picture nor a GDR picture with ph_recovery_poc_cnt equal to 0 is not used as an ILRP. A vps_max_tid_il_ref_pics_plus1[i][j] greater than 0 specifies that for decoding the picture of the i-th layer, the picture from the j-th layer with a TemporalId greater than vps_max_tid_il_ref_pics_plus1[i][j] - 1 is not used as an ILRP. If it does not exist, the value of vps_max_tid_il_ref_pics_plus1[i][j] is assumed to be equal to vps_max_sublayers_minus1 + 1.
[0055] Note that if it does not exist, the value is assumed to be vps_max_sublayer_minus1 + 1, where vps_max_sublayer_minus1 is the maximum number of sublayers present in any layer within the bitstream. However, for a particular layer, the value of the maximum number of sublayers may be smaller.
[0056] This syntax element not only indicates that inter-layer references are not used for some sublayers, or that some sublayers of the reference layer are not needed for decoding, but also indicates a special mode (vps_max_tid_il_ref_pics_plus1[i][j] equal to 0) where only IRAP NAL units or GDR NAL units with ph_recovery_poc_cnt equal to 0 are required from the reference layer for decoding. Furthermore, the output layer set describing the bitstream passed to the decoder does not include the unnecessary NAL units indicated by this syntax element vps_max_tid_il_ref_pics_plus1[i][j], or such NAL units are dropped in a particular decoder implementation that performs the extraction process defined in the specification.
[0057] The syntax element vps_max_tid_il_ref_pics_plus1[i][j] exists only in the direct reference layer. For example, imagine an OLS with three layers as shown in Figure 4. Figure 4 shows an example L0, L1, L2 of three layers where vps_max_tid_il_ref_pics_plus1 is equal to 0 in all reference layers. Since L2 has L1 as the direct reference layer and L1 uses L0 as the direct reference layer, L2 has L0 as the indirect reference layer. In such a case, if vps_max_tid_il_ref_pics_plus1[2][1] is equal to 0, only the IRAP NAL unit or GDR NAL unit where ph_recovery_poc_cnt is equal to 0 is retained from L1 and, as a result, from L0 as specified in the specification. More specifically, for each layer of the OLS, a variable indicating the number of retained sublayers, i.e., NumSubLayersInLayerInOLS[i][j] (where i is the OLS index and j is the layer index), is derived. If the layer is an output layer, the value of this variable is set to the maximum temporalId required in the bitstream. If it is not an output layer but a reference layer of each layer k of the OLS that uses layer j as a reference, the value of NumSubLayersInLayerInOLS[i][j] is set to the maximum of min(NumSubLayersInLayerInOLS[i][k], vps_max_tid_il_ref_pics_plus1[k][j]). That is, for each layer k, check what is the minimum value between the number of sublayers required for layer k (NumSubLayersInLayerInOLS[i][k]) and the number of sublayers required for layer j (vps_max_tid_il_ref_pics_plus1[k][j]) considering all sublayers of layer k. If fewer sublayers are required in layer k than indicated by vps_max_tid_il_ref_pics_plus1[k][j], then in layer j only the same amount of sublayers as in layer k are required, so the minimum of the two values is selected.Then, a further layer k that also uses j as a reference is checked, and if other layers indicate that more sub-layers are required, a higher value is obtained, i.e., the maximum value required after all layers that use layer j as a reference have been checked.
[0058] Problems occur when the IRAP NAL unit is not aligned between L0 and L1. For example, as shown in FIG. 5, which shows an example of a three-layer with a misaligned IRAP in the lower layer, imagine that there is an IRAP AU at a certain point in L1 but not in L0, and the IRAP AU in L1 uses the non-IRAP AU in L0 as a reference. In such a case, in the IRAP-based extraction process, since the non-IRAP in L0 (picture 260* in FIG. 5) will be discarded, the IRAP AU in L1 (picture 261 in FIG. 5) cannot be decoded.
[0059] In an embodiment, when vps_max_tid_il_ref_pics_plus1[i][j] is 0 for layer i, it is required that any direct or indirect layer of such layer i has an aligned IRAP or GDR NAL unit with ph_recovery_poc_cnt equal to 0. In other words, in any indirect reference layer, the co-temporal NAL unit of the direct reference layer that depends on that indirect reference layer is either an IRAP NAL unit or a GDR NAL unit with ph_recovery_poc_cnt equal to 0. The NAL unit of that indirect reference layer must similarly be either an IRAP NAL unit or a GDR NAL unit with ph_recovery_poc_cnt equal to 0.
[0060] According to an embodiment of the second aspect, the video bitstream 14 includes a sequence of access units 22, each of which includes one or more bitstream portions 16. Each of the bitstream portions 16 is associated with one of a plurality of layers 24 of the video bitstream 14 and one of a plurality of temporal layers of the video bitstream, for example, the temporal sublayer described with respect to FIG. 1. Bitstream portions 16 within the same access unit 22 are associated with the same temporal layer. Further, each of the bitstream portions is one of a set of bitstream portion types that includes a set of predetermined bitstream portion types. For example, the set of predetermined bitstream portion types can include independently encoded bitstream portion types such as IDR and, optionally, types that depend only on other access units of the set of predetermined bitstream portion types.
[0061] According to an embodiment of the second aspect, the encoder 10 is configured to provide an OLS representation of the OLS of the video bitstream 14 in the video bitstream 14, where the OLS includes one or more layers of the video bitstream. Further, the encoder 10 provides a reference layer representation in the video bitstream that indicates, for each layer of the OLS, the set of reference layers on which each layer depends. Further, the encoder 10 provides, in the video bitstream 14, for each layer (e.g., i) of the OLS, for each reference layer (e.g., j) of each layer, whether all bitstream portions of each reference layer on which each layer depends are one of a set of predetermined bitstream portion types, or, if not, whether it is a bitstream portion up to which temporal layer each layer depends (e.g., vps_max_tid_il_ref_pics_plus1[i][j]), a temporal layer representation (e.g., the maximum index indexing the temporal layer up to which each layer depends).).
[0062] The encoder 10 according to the present embodiment, for each layer of the OLS (for example, i), the time layer display for this indicates that the entire bitstream portion of a predetermined reference layer (the reference layer on which each layer depends) is one of a set of predetermined bitstream portion types (for example, the same), and an access unit including the bitstream portion of the predetermined reference layer that is one of the set of predetermined bitstream portion types is provided for each further reference layer on which the predetermined reference layer directly or indirectly depends, so as not to include a bitstream portion other than the set of predetermined bitstream portion types (for example, direct dependency or direct reference is the dependency between the (dependent) layer and its reference layer, as shown in the reference layer display, and indirect dependency or reference is the dependency between the (dependent) layer and the direct or indirect reference layer of the reference layer of the (dependent) layer not shown in the reference layer display).
[0063] In fact, this is overly restrictive. Such an indirect reference layer (L0) can also be a reference layer of another layer, as shown in FIG. 6 showing an example of a four-layer that directly references a sublayer with Tid1 (imagine the case of the fourth L3 layer where sublayers 0 and 1 of L0 are required and vps_max_tid_il_ref_pics_plus1[3][0] is 2).
[0064] In such a case, since the non-IRAP NAL units of layer 0 required for the IRAP NAL units of L1 are held in the OLS bitstream corresponding to L0+L1+L2+L3, the IRAP alignment constraints described in the previous embodiments should become unnecessary. Therefore, in order to represent the constraints, the variable NumSubLayersInLayerInOLS[i][j] can be used instead. This variable indicates the number of sub-layers (each with its own temporary ID) held in the i-th OLS for the j-th layer (0 means that only IRAP or GDR with ph_recovery_poc_cnt equal to 0 is held).
[0065] In the embodiment, within the i-th OLS, for two layers k and j where k>j, when NumSubLayersInLayerInOLS[i][j] and NumSubLayersInLayerInOLS[i][k] are equal to 0, if j is a reference layer of k (directly or indirectly), the IRAP NAL unit or the GDR with ph_recovery_poc_cnt equal to 0 is aligned.
[0066] Thus, as an alternative to the embodiments described with respect to FIGS. 4 and 5, encoder 10 may supply bitstream 14 to a further reference layer 20D upon which a given reference layer 20C, upon which each layer 20B depends (the reference layer of each layer), directly or indirectly depends, for each layer (e.g., i) 20B of the OLS for which the time layer indication indicates that all bitstream portions of the given reference layer 20C upon which each layer 20B depends are one (e.g., the same) of a set of given bitstream portion types. The following two criteria are met. As a first criterion, access units 40A, 40B, which include a bitstream portion of one of the set of given bitstream portion types among the bitstream portions of the given reference layer, have no bitstream portions outside the set of given bitstream portion types (40A) or do not (40B). As a second criterion, each further reference layer 20D is, according to the reference layer indication, a reference layer of a direct reference layer 20A upon which each layer 20B depends.
[0067] According to an alternative embodiment, encoder 10 is configured to provide, within video bitstream 14, in addition to the OLS representation and the reference layer representation described with respect to the previous embodiments, for each layer (e.g., j) of an OLS (e.g., i), a within-layer time layer representation [e.g., NumSubLayersInLayerInOLS[i][j]] indicating a subset of time layers [e.g., maximum time layer index] that includes the bitstream portion of each layer that the OLS requires if the OLS requires only the bitstream portion of each layer that is one of a set of given bitstream portion types [e.g., indicated by NumSubLayersInLayerInOLS[i][j]=0], or otherwise.
[0068] According to this embodiment, the encoder 10 is configured to provide each bitstream 14 for each bitstream portion of an access unit including one bitstream portion of a set of predetermined bitstream portion types, for each layer 24 of the OLS indicating that the in-layer time layer display requires only one bitstream portion of one of a series of predetermined bitstream portion types. The following conditions are satisfied. That is, when each bitstream portion belongs to a layer of the OLS and its in-layer time display indicates that the OLS requires only the bitstream portion of each layer that is one of a set of predetermined bitstream portion types, each bitstream portion is one of a set of predetermined bitstream portion types, or according to the reference layer display, each layer does not depend on the layer of each bitstream portion.
[0069] According to a further embodiment, when the IRAP is not aligned, the inter-layer prediction is not used for such IRAP NAL units for the layers where the non-RAP NAL units are present in the same AU.
[0070] Thus, according to another embodiment according to the second aspect, encoder 10 is configured to provide OLS display, reference layer display, and temporal layer display in video bitstream 14 as described for the previous embodiment of section 2. Further, according to this embodiment, for each layer (e.g., i) of the OLS, when the temporal layer display for this has all bitstream portions of a predetermined reference layer (the reference layer of each layer) on which each layer depends being a bitstream portion of a predetermined reference layer that is one of a set of types of predetermined bitstream portions, and there is no bitstream portion other than the set of types of predetermined bitstream portions for each reference layer on which the predetermined reference layer directly or indirectly depends, and it indicates that all bitstream portions of the predetermined reference layer on which the predetermined reference layer depends are one of a set of types of predetermined bitstream portions, the bitstream portion of the access unit including the bitstream portion of the predetermined reference layer that is one of a set of predetermined bitstream portion types is encoded without using the inter-prediction means for the bitstream portion belonging to the layer having a direct or indirect reference to one of the further reference layers where there is no bitstream portion other than the set of predetermined bitstream portion types.
[0071] For example, the set of types of predetermined bitstream portions can include one or more or all of the IRAP type and the type of GDR where ph_recovery_poc_cnt is equal to zero.
[0072] An embodiment of the encoder 10 according to the second aspect may be configured to provide a level representation of the bitstream 12 that can be extracted from the video bitstream 14 according to the OLS within the video bitstream. For example, the level representation may include one or more of the coded picture buffer size, decoded picture buffer size, picture size, picture rate, minimum compression rate, picture segmentation limit (e.g., tile / slice / sub-picture), bitrate, buffer scheduling (e.g., HRD timing (AU / DU removal time, DPB output time)).
[0073] 3. Bitstream-based OLS determination In Section 3, embodiments according to the third aspect of the present invention will be described with reference to FIG. 1, and the details described in Section 0 may optionally be applied to the embodiments according to the third aspect. Also, the details described regarding further aspects can optionally be implemented in the embodiments described in this section.
[0074] Embodiments of the third aspect can provide the identification of the OLS corresponding to the bitstream. In other words, according to the embodiments of the third aspect, it becomes possible to estimate the OLS of the video bitstream from which the OLS is to be decoded or extracted from the video bitstream. The decoder that receives the bitstream to be decoded may be given additional information regarding the operation point to be decoded via its API. For example, in the current VVC draft specification, two variables are set by external means as follows.
[0075] The variable TargetOlsIdx that identifies the OLS index of the target OLS to be decoded and the variable Htid that identifies the top temporal sublayer to be decoded are set by external means not specified in this specification. The bitstream BitstreamToDecode does not include NAL units other than those included in the target OLS and with a TemporalId greater than Htid.
[0076] This specification does not mention what to do when these variables are not set. In such cases, the decoder is expected to simply decode the entire given bitstream, rather than decoding a subset of the bitstream, for example, with respect to the time sublayer.
[0077] However, with respect to the output layer set, there is the following problem. When a decoder is given a bitstream containing multiple layers and the parameter set defines multiple OLSs that include all layers within the bitstream (for example, variants with different output layers), the decoder cannot simply determine which output layer set it must decode from the bitstream itself. Depending on the characteristics of the OLS, the selected OLS may result in various level requirements due to various DPS parameters and the like. Therefore, it is important to enable the decoder to select the OLS even when there is no external signal via the API. In other words, a fallback method is required, similar to other cases where there is no external means, such as the selection of the topmost time sublayer within the bitstream to be decoded.
[0078] In one embodiment, there is a bitstream compliance constraint that the bitstream corresponds to only a single OLS so that the decoder can clearly determine which OLS to decode from the given bitstream. This characteristic can be instantiated by a syntax element indicating that, for example, all OLSs can be clearly determined by the layers present in the bitstream, that is, there is a unique mapping from the number of layers to the OLS.
[0079] According to an embodiment of the third aspect, an encoder 10 for providing a multi-layer video bitstream 14 configures a plurality of OLSs within the multi-layer video bitstream 14, for example, as shown in the OLS display 18 of FIG. 1. Each of the OLSs represents a subset of the layers of the multi-layer video bitstream 14. It should be noted that the subset of layers is not necessarily a proper subset of the layers. That is, the subset of layers may include all the layers of the multi-layer video bitstream 14. The encoder 10 according to this embodiment provides the multi-layer video bitstream 14 such that for each of the OLSs, a sub-bitstream, such as the sub-bitstream 12 of the multi-layer video bitstream 14 defined by each OLS, is distinguishable from the sub-bitstreams of the multi-layer video bitstream defined by any other of the plurality of OLSs. For example, since the encoder 10 can provide the multi-layer video bitstream such that the OLSs represent different subsets of the layers of the multi-layer video bitstream 14, the OLSs are distinguishable by the subsets of layers.
[0080] For example, each of the OLSs can be defined for the OLS by indicating the subset of layers by layer indices in the OLS display 18. Optionally, the OLS display can include additional parameters that define a subset of the bitstream portion of the layers of the OLS to which the bitstream portion belongs. For example, the OLS display can indicate which temporal sub-layers belong to the OLS.
[0081] In an example, the encoder 10 can indicate within the multi-layer video bitstream 14 that the multi-layer video bitstream 14 uniquely belongs to one of the OLSs. For example, the encoder 10 can indicate a plurality of OLSs such that, for each of the OLSs, a subset of the layers of each OLS is different from any subset of the layers of the other OLSs. Thus, in the example, the encoder 10 can indicate within the multi-layer video bitstream 14 that a set of layers of the multi-layer video bitstream 14, which can be indicated, for example, by a set of indexes, uniquely belongs to one of the OLSs.
[0082] According to an embodiment, the encoder 10 is configured to check the conformity of the multi-layer video bitstream 14 by checking, for each of the OLSs among the plurality of OLSs, whether the sub-bitstream of the multi-layer video bitstream 14 defined by each OLS is distinguishable or different from the sub-bitstream of the multi-layer video bitstream defined by any of the other OLSs. For example, if not, the encoder 10 can negate the conformity of the bitstream.
[0083] Thus, an embodiment of a decoder for decoding a video bitstream, for example, the decoder 50 of FIG. 1, can derive from the video bitstream to be decoded, such as the video bitstream 14 or the video bitstream 12 of FIG. 1 extracted therefrom, one or more OLSs each of which indicates a subset of the layers of the video bitstream. The decoder can detect an indication within the video bitstream that the video bitstream uniquely belongs to one of the OLSs and can decode one of the OLSs to which the video bitstream belongs.
[0084] For example, decoder 50 can identify the OLS belonging to the video bitstream by identifying the layer included in the video bitstream, and can decode the OLS that accurately identifies the layer included in the video bitstream. Therefore, an indication that the video bitstream unambiguously belongs to one of the OLSs can indicate that the set of layers within the video bitstream unambiguously belongs to one of the OLSs.
[0085] For example, decoder 50 can determine one of the OLSs belonging to the video bitstream by inspecting the first access unit of the encoded video sequence. The first access unit can refer to the first received one, the first in temporal order, the first in decode order, or the first in output order. Alternatively, decoder 50 can determine one of the OLSs by inspecting that the first of the access units is of the sequence start access unit type, e.g., a CVSS access unit, where the first is defined by receiving, for example, the reception order, the temporal order, the decode order. For example, decoder 50 can inspect the first of the access units of the encoded video sequence, or that the first of the access units is of the sequence start access unit type with respect to the layers included in each access unit.
[0086] For example, decoder 50 can determine one of the OLSs such that each access unit contains exactly one OLS layer picture when the first access unit of the encoded video sequence, or the first of the access units, is of the sequence start access unit type.
[0087] Some of the above embodiments, for example, in a multi-view 2-layer scenario having both an OLS that outputs both views and an OLS that outputs only the dependently encoded view, may have the drawback that certain combinations of OLSs are prohibited. To mitigate this limitation, another embodiment of the present invention is to have a selection algorithm by, for example, one or more of the following combinations from among the OLSs corresponding to the bitstream or its first access unit or its CVSS AU.
[0088] - Selecting the OLS with the highest or lowest index - Selecting the OLS with the largest number of output layers FIG. 7 shows a decoder 50 according to one embodiment, which can optionally follow the latter embodiment having a selection algorithm. The decoder 50 according to FIG. 7 can optionally correspond to the decoder 50 according to FIG. 1. The decoder 50 according to FIG. 7 is configured to decode a video bitstream 12, for example, a video bitstream 12 extracted from a multi-layer video bitstream 14. Alternatively, the video bitstream 12 of FIG. 7 can correspond to the video bitstream 14 of FIG. 1. The video bitstream 12 of FIG. 7 may or may not be a multi-layer video bitstream, i.e., in an example, it may be a single-layer video bitstream. The video bitstream 12 includes access units 22 of an encoded video sequence 20, such as access units 221, 222, and each access unit includes one or more pictures of the encoded video sequence, such as pictures 261, 262. Each of the pictures 26 belongs to one of one or more layers 24 of the video bitstream 12, as described, for example, in Section 0. The decoder 50 according to FIG. 7 is configured to derive one or more OLSs from the video bitstream 12. For example, the video bitstream 12 includes an OLS display 18, as described, for example, with respect to FIG. 1, and the OLS display 18 shows one or more OLSs such as OLS181 and OLS182 as shown in FIG. 7. Each of the one or more OLSs shows a (not necessarily appropriate) set of one or more layers 24 of the video bitstream 12. In other words, each of the OLSs shows one or more of the layers that should be part of the respective OLS. The decoder 50 determines one of the OLSs based on one or more attributes of each of the OLSs and decodes one OLS determined from the OLSs.
[0089] As described with respect to the previous embodiment, the decoder 50 may determine a subset of OLSs and one OLS by examining the first access unit of the encoded video sequence, or by examining that the first access unit is of the sequence start access unit type. For example, the access unit 221 shown in FIG. 7 may be the first access unit of the encoded video sequence of the video bitstream 12 (e.g., the first received, or the first in the encoding order, or the first in the temporal order, or the first in the output order). In other examples, the encoded video sequence 20 may include additional access units that precede the access unit 221, and since the preceding access unit is not a sequence start access unit, the access unit 221 is the first sequence start access unit of the sequence 20. The decoder 50 can examine the access unit 221 to detect the picture 261 of the first layer 241 and the picture 262 of the second layer 242. Based on this finding, the decoder 50 can conclude that the video bitstream 12 comprises the first layer 241 and the second layer 242.
[0090] According to the embodiment of FIG. 7, the decoder 50 determines one OLS to be decoded based on one or more attributes of the OLSs. The one or more attributes may include one or more of the index of each OLS (i.e., the OLS index) and / or the number of layers of the OLS and / or the number of output layers of each OLS. For example, the decoder determines, as one OLS, the OLS having the highest or lowest OLS index and / or the largest number of layers and / or the largest number of output layers among the OLSs. In other words, the decoder 50 can evaluate which OLS has the highest or lowest OLS index and / or which OLS has the largest number of layers. Additionally or alternatively, the decoder 50 can evaluate which of the OLSs includes an output layer having the highest or lowest index. By selecting one OLS based on the largest number of output layers and / or the largest number of layers, a bitstream that provides the highest quality output for the video sequence can be selected for decoding.
[0091] For example, the decoder 50 can select, as one OLS, the OLS having the largest number of output layers among the OLSs, the OLS having the largest number of layers beyond the OLS having the largest number of output layers, and the OLS having the lowest OLS index beyond the OLS having the largest number of layers beyond the OLS having the largest number of output layers.
[0092] According to an embodiment, the decoder 50 determines one OLS by evaluating which of the OLSs has the largest number of layers. If there are multiple OLSs having the largest number of layers, the decoder 50 may evaluate the one having the largest number of output layers among the OLSs having the largest number of layers, and may select the OLS having the largest number of output layers among the OLSs having the largest number of layers as one OLS.
[0093] In other words, if there is no OLS instruction to be encoded, the decoder 50 can decode the OLS shown in the OLS display 18, where all necessary layers are present in the bitstream and most of the existing layers are used, and provide a highly faithful video output.
[0094] In the following, a further embodiment of the decoder 50 will be described with reference to FIG. 7.
[0095] Also, the decoder 50 of this further embodiment may optionally conform to the selection algorithm described above and may optionally correspond to the decoder 50 according to FIG. 1. According to this further embodiment, the decoder 50 is configured to decode the video bitstream 12, such as the video bitstream 14 extracted from the multi-layer video bitstream 12, as shown in FIG. 1. Alternatively, the video bitstream 12 of FIG. 7 may correspond to the video bitstream 14 of FIG. 1. The video bitstream 12 of FIG. 7 may be a multi-layer video bitstream, but it does not necessarily have to be, i.e., in an example, it may be a single-layer video bitstream. The video bitstream 12 includes access units 22 of the encoded video sequence 20, such as access units 221, 222, and each access unit includes one or more pictures of the encoded video sequence, such as pictures 261, 262. Each of the pictures 26 belongs to one of the one or more layers 24 of the video bitstream 12, as described, for example, in section 0. The decoder 50 according to this further embodiment is configured to derive one or more OLSs from the video bitstream 12. For example, the video bitstream 12 includes an OLS display 18, as described, for example, with respect to FIG. 1, and the OLS display 18 shows one or more OLSs such as OLS181 and OLS182 as shown in FIG. 7. Each of the one or more OLSs indicates a (not necessarily appropriate) set of one or more layers 24 of the video bitstream 12. In other words, each of the OLSs indicates one or more of the layers that should be part of the respective OLS. The decoder 50 according to this further embodiment determines a subset of OLSs from (or among) the OLSs such that each of the OLSs in the subset of OLSs belongs to the video bitstream 12. The decoder 50 according to this further embodiment further determines one of the OLSs in the subset of OLSs based on one or more attributes of each of the OLSs in the subset of OLSs and decodes the one OLS determined from the subset of OLSs.
[0096] For example, the OLSs belonging to the video bitstream may mean that a set of one or more layers present in the video bitstream 12 corresponds to the set of layers indicated by each OLS. In other words, the decoder 50 may determine a subset of the OLSs belonging to the video bitstream 12 based on the set of layers present in the video bitstream 12. That is, for each of the subsets of OLSs, the decoder 50 may determine the subset of OLSs such that the subset of layers indicated by each OLS corresponds to the set of layers of the video bitstream 12. For example, in FIG. 7, the video bitstream 12 illustratively includes the picture 261 of the first layer 241 and the picture 262 of the second layer 242. The OLS 181 indicates that the first layer 241 and the second layer 242 are part of the OLS 181. Also, the OLS 182 indicates that both the first layer and the second layer are part of the OLS 182. Thus, according to the example of FIG. 7, the decoder 50 can ascribe both the OLSs 181 and 182 to a part of the indices of the OLSs belonging to the video bitstream 12.
[0097] Optionally, the decoder 50 may, according to the level information of the OLSs, consider only those OLSs for decoding that are decodable by the decoder 50.
[0098] As described with respect to the previous embodiments, the decoder 50 can determine a subset of the attributed OLSs and one OLS by examining the first access unit of the encoded video sequence, or by examining that the first access unit is of the sequence start access unit type. For example, the access unit 221 shown in FIG. 7 can be the first access unit of the encoded video sequence of the video bitstream 12 (e.g., the first received, or the first in the encoding order, or the first in the temporal order, or the first in the output order). In other examples, the encoded video sequence 20 can include additional access units that precede the access unit 221, and since the preceding access units are not sequence start access units, the access unit 221 is the first sequence start access unit of the sequence 20. The decoder 50 can examine the access unit 221 to detect the picture 261 of the first layer 241 and the picture 262 of the second layer 242. Based on this finding, the decoder 50 can conclude that the video bitstream 12 comprises the first layer 241 and the second layer 242.
[0099] According to an embodiment, the decoder 50 can determine a subset of the OLSs such that each access unit accurately contains the pictures of the respective layers of the subset of the OLSs when the first access unit of the encoded video sequence, or when the first access unit is of the sequence start access unit type, e.g., the access unit 221.
[0100] According to the embodiment of FIG. 7, the decoder 50 determines one OLS to be decoded based on one or more attributes of a subset of the OLSs. The one or more attributes may include one or more of the index of each OLS and / or the number of output layers, the highest or lowest index, i.e., the index of the highest or lowest layer, and the number of the highest layer. In other words, the decoder 50 can evaluate which OLSs among the subset of OLSs belonging to the video bitstream 12 include the layer indexed by the highest or lowest layer index and / or which OLSs have the largest number of layers. Additionally or alternatively, the decoder 50 can evaluate which OLSs among the subset of OLSs have the maximum or minimum number of output layers and / or which OLSs among the subset of OLSs constitute the output layer with the highest or lowest index.
[0101] According to an embodiment, the decoder 50 determines one OLS by evaluating which OLS among the subset of the OLSs has the largest number of layers. If there are multiple OLSs having the largest number of layers, the decoder 50 may evaluate the one having the largest number of output layers among the OLSs having the largest number of layers, or may select the OLS having the largest number of output layers among the OLSs having the largest number of layers as one OLS.
[0102] In other words, when there is no instruction for the OLS to be encoded, the decoder 50 can decode the OLSs shown in the OLS display 18 in which all the necessary layers exist in the bitstream and most of the existing layers are used, and can provide a highly faithful video output.
[0103] In other words, the embodiment of the third aspect includes a decoder 50 for decoding video bitstreams 12, 14, the video bitstream 14 includes an access unit 22 of an encoded video sequence 20, each access unit 22 includes one or more pictures 26 of the encoded video sequence, and each of the pictures belongs to one of one or more layers 24 of the video bitstream 14. The decoder is configured to derive one or more output layer sets (OLSs) 181, 182 from the video bitstream 14, each indicating a set of one or more layers of the video bitstream 14, to determine a subset of the OLSs 181, 182 from the OLSs 181, 182, each subset of the OLSs belonging to the video bitstream 14, to determine one of the subsets of the OLSs based on one or more attributes of each subset of the OLSs, and to decode the OLSs.
[0104] According to an embodiment, the decoder 50 is configured to determine a subset of the OLSs such that, for each subset of the OLSs, the subset of layers indicated by each OLS corresponds to the set of layers of the video bitstream 14.
[0105] According to an embodiment, the decoder 50 is configured to determine a subset of the OLSs and one OLS by inspecting whether the first access unit of the encoded video sequence or the first access unit 22 is of the sequence start access unit 22 type.
[0106] According to an embodiment, when the first access unit 22 of the encoded video sequence or the first access unit 22 is of the sequence start access unit type, the decoder 50 is configured to determine a subset of the OLSs such that each access unit accurately includes the pictures of each layer of the subset of the OLSs.
[0107] According to an embodiment, the decoder 50 is configured to determine one of the OLSs by evaluating one criterion for each of one or more attributes of a subset of the OLSs.
[0108] 4. Sequence Start Access Unit of Sub-bitstream Section 4 describes an embodiment of the fourth aspect of the present invention with reference to FIG. 1. The description provided in Section 0 can be arbitrarily applied to the embodiment of the fourth aspect. Further details regarding additional aspects can be optionally implemented in the embodiment described in this section.
[0109] Some embodiments of the fourth aspect may relate to access unit delimiters (AUDs) within supplemental enhancement information (SEI) to enable a coded video sequence start access unit (CVSS AU) in a video bitstream, such as video bitstream 14 from which video bitstream 12 is originally extracted from the video bitstream and was not originally a CVSS AU. For example, the CVSS AU can be a randomly accessible AU, such as an AU having a randomly accessible or independently encoded picture in each layer of the video bitstream, or an AU that is decodable independently of the previous AU in the video bitstream.
[0110] In the current specification, it is obligatory for the coded video sequence start (CVSS) AU to have an IRAP NAL unit type or a GDR NAL unit type in each layer, and for the IRAP NAL unit type to be the same within the CVSS AU. Furthermore, the presence of an AUD (access unit delimiter) indicating that the CVSS AU is an IRAP AU or a GDR AU is obligatory.
[0111] Figure 8 shows an example of a multi-layer bitstream in which the IRAPs are not aligned across different layers, i.e., not all IRAPs are aligned across the layers, and the identification of the CVSS AUs.
[0112] AU2, AU4, and AU6 have NAL unit types of the IRAP type in two of the bottom layers, but since not all layers have the same IRAP type in these AUs, it can be seen that these AUs are not CVSS AUs. It is not necessary to analyze all the NAL units of an AU. To easily identify the CVSS AUs, the AUD NAL units are used and can be easily identified. That is, AU0 and AU8 will contain an AUD with a flag indicating that these AUs are CVSS AUs (IRAP AUs).
[0113] Figure 9 shows an example of the video bitstream after extraction, e.g., an example of video bitstream 12. However, if the bitstream is extracted with only L0 and L1, as indicated by reference numeral 22*, there will be new CVSS AUs in the extracted bitstream. In other words, after extraction, some AUs (e.g., AU22*) will become CVSS AUs.
[0114] Since AU2 and AU6 will be CVSS AUs or IRAP AUs, there must be an AUD in the bitstream of such AUs indicating the IRAP AU property.
[0115] An embodiment according to the fourth aspect includes an apparatus for extracting a sub-bitstream from a multi-layer video bitstream, for example, an extractor 30 for extracting the sub-bitstream 12 from the multi-layer video bitstream 14 described with reference to FIG. 1. According to the fourth aspect, the multi-layer video bitstream 14 represents an encoded video sequence such as the encoded video sequence 20, and the multi-layer video bitstream includes access units 22 of the encoded video sequence. Each of the access units 22 includes one or more bitstream portions 16 of the multi-layer video bitstream 14, and each of the bitstream portions belongs to one of the layers 24 of the multi-layer video bitstream. According to the fourth aspect, the extractor 30 is configured to derive from the multi-layer video bitstream 14 one or more OLSs each indicating a (not necessarily appropriate) subset of the layers of the multi-layer video bitstream 14. For example, the multi-layer video bitstream includes an OLS representation 18 including information or an explanation about one or more OLSs such as OLS 181 and OLS 182 as shown in FIG. 7.
[0116] The extractor 30 according to the fourth aspect is configured to provide layer 24 of the multi-layer video bitstream 14, which is indicated by a predetermined one of the OLSs, i.e., shown to be part of a predetermined OLS, within the sub-bitstream 12. In other words, the extractor 30 can provide, within the sub-bitstream 12, a bitstream portion 16 belonging to each layer of a predetermined OLS. For example, the predetermined OLS may be provided to the extractor 30 by external means, e.g., by an OLS command 32 as shown in FIG. 1. According to an embodiment of the fourth aspect, the extractor 30 provides, within the sub-bitstream 12, for each access unit of the sub-bitstream 12, a sequence start indication indicating that it is the start access unit of a sub-sequence of the coded video sequence when all bitstream portions of each access unit are bitstream portions of the same one of a set of predetermined bitstream portion types, e.g., the access unit 22* of the sub-bitstream including L0 and L1 described with respect to FIGS. 8 and 9.
[0117] In other words, the extractor 30 may provide an access unit that exclusively includes bitstream portions of the same one of a set of predetermined bitstream portion types with a sequence start indication, where the access unit is included in or provided by the extractor 30 to the sub-bitstream 12.
[0118] For example, the set of predetermined bitstream portion types can include one or more IRAP NAL unit types and / or GDR NAL unit types. For example, the set of predetermined bitstream portion types can include NAL unit types of IDR_NUT, CRA_NUT, or GDR_NUT.
[0119] For example, the extractor 30 may determine, for each access unit in the multi-layer video bitstream that does not have a sequence start display, or alternatively, for each access unit of the sub-bitstream 12, whether all bitstream portions of each access unit are bitstream portions of the same one of a set of predetermined bitstream portion types. For example, the extractor 30 can analyze the relevant information in each access unit or the multi-layer video bitstream to determine whether all bitstream portions of each access unit are bitstream portions of the same one from a set of predetermined bitstream portion types.
[0120] In other words, in one embodiment, the bitstream extraction process removes unnecessary layers and adds an AUD NUT for such uses if they do not exist in AUs that require an AUD after extraction.
[0121] According to an embodiment, the extractor 30 is configured to estimate from the display in the multi-layer video bitstream 14 that one of the access units is the start access unit of a sub-sequence of an encoded video sequence represented by a predetermined OLS, and provide a sequence start display in the sub-bitstream indicating that one access unit is the start access unit.
[0122] In other words, in another embodiment, in such an AU (i.e., the one that becomes a CVSS AU), there is a display in the bitstream, and such an AU becomes a CVSS AU or an IRAP AU when the layer is extracted, and the insertion (addition) of an AUD is simpler and does not require much analysis.
[0123] According to a further embodiment, for a given OLS, the extractor 30 is configured to extract nested information, such as a nested SEI, indicating that one or more access units, e.g., access units that are not the start access unit of the multi-layer video bitstream 14, are the start access unit for the OLS. According to this embodiment, the extractor 30 provides a sequence start indication within the sub-bitstream indicating that one or more access units indicated within the nested information are start access units. For example, the extractor 30 may provide a sequence start indication for each of the indicated access units as described above. Alternatively, the device may provide a common indication for the indicated access units within the sub-bitstream.
[0124] In other words, in another embodiment, when a specific OLS that changes the AU to an IRAP / GDR AU is extracted, there is a nested SEI that can encapsulate the AUD such that the encapsulated AUD is decapsulated and added to the bitstream. Currently, the specification only includes nesting of other SEI messages. Therefore, non-VCL non-SEI payloads need to be permitted within the nested SEI. One way is to extend the existing nested SEI to indicate that non-SEI is included. Another is to add a new SEI that includes other non-VCL payloads within the SEI.
[0125] The first option for implementation is shown in Table 2.
Table 2-1
Table 2-2
[0126] Table 3 shows Option 2, which is another option. [Table 3] In Option 2, if there is a future need to include other non-VCL NAL units in the nesting SEI, types that allow other non-VCL NAL units can be added so that the nonVclNutPayloadSEI message can be used.
[0127] In such a case, a single non-VCL NAL unit (in this case, a new SEI message directly containing the access unit delimiter (AUD_rbsp())) is defined, and thus such an encapsulated SEI can be added directly without changing the nesting SEI (the nesting SEI already contains other SEIs within itself).
[0128] According to a further embodiment of the fourth aspect, for each access unit of the sub-bitstream, the extractor 30 is configured to provide a sequence start indication indicating that each access unit is the start access unit of the encoded video sequence, where all bitstream portions of each access unit are bitstream portions of the same one of a set of predetermined bitstream portion types, and each access unit contains bitstream portions of two or more layers.
[0129] In another embodiment, the AUD is also obligated to access units that could become IRAP AU or GDR AU in the case of OLS extraction (layer drop), and the extractor can confirm their presence when needed and easily rewrite the aud_irap_or_gdr_au_flag to 1 when appropriate (i.e., the AU becomes an IRAP or GDR AU after extraction). One way to express this constraint in the specification is to extend the existing text that makes the AUD mandatory for IRAP or GDR AU of the current bitstream.
[0130] An AU can have at most one EOB NAL unit. When vps_max_layers_minus1 is greater than 0, each IRAP or GDR AU shall have only one AUD NAL unit.
[0131] This is changed as follows.
[0132] An AU can have at most one EOB NAL unit. When vps_max_layers_minus1 is greater than 0, each AU containing at least two layers including only IRAP or GDR NAL units shall have only one AUD NAL unit.
[0133] In other words, instead of using all NAL units, the words IRAP or GDR picture are used.
[0134] An AU can have at most one EOB NAL unit. When vps_max_layers_minus1 is greater than 0, each AU containing at least two IRAP or GDR pictures shall have only one AUD NAL unit.
[0135] Accordingly, a further embodiment according to the fourth aspect includes an encoder 10 for providing a multi-layer video bitstream, for example, the encoder 10 described with respect to FIG. 1. An embodiment of the encoder 10 according to the fourth aspect is configured to provide a multi-layer video bitstream representing an encoded video sequence such as video sequence 20. The multi-layer video bitstream 14 provided by the encoder 10 includes a sequence of access units 22, each of which includes one or more pictures, and each of which is associated with one of a plurality of layers 24 of the video bitstream. For each access unit of the multi-layer video bitstream 14, the encoder 10 is configured to provide, within the sub-bitstream 12, a sequence start indicator indicating whether all pictures of each access unit are of one picture of a set of predetermined picture types when each access unit includes at least two pictures of one of the set of predetermined picture types. Optionally, the encoder 10 is configured to provide an output layer set (OLS) representation 18 in the multi-layer video bitstream 14, where the OLS includes one or more layers of the multi-layer video bitstream 14, and when each access unit includes at least two pictures of one of a set of predetermined access unit types, the encoder 10 provides a sequence start indicator for each access unit.
[0136] For example, the picture type includes one or more IRAP picture types and / or GDR picture types. A picture that is a picture type may mean that each of one or more bitstream portions in which the picture is encoded is one of a set of predetermined bitstream portion types, such as one or more IRAP NAL unit types or GDR NAL units.
[0137] For example, the sequence start indicator referred to in the embodiments of Section 4 as a whole may be provided in the form of, or as part of, an access unit delimiter (AUD) which may be a part of the bitstream provided within each access unit. For example, the indication that an access unit is a sequence start access unit may be indicated by setting a flag of the AUD, for example, aud_irap_or_gdr_au_flag, to a value of 1, for example, to indicate that each access unit is a sequence start access unit.
[0138] Alternatively, the encoder 10 does not necessarily have to provide a sequence start indicator for each access unit that includes at least two pictures of one of a set of predetermined picture types, but the encoder 10 may provide a sequence start indicator for each access unit of the multi-layer video bitstream 14, and each such access unit includes at least two pictures of one of a set of predetermined picture types, and each of the pictures belongs to one of the OLS layers indicated by the OLS representation provided within the multi-layer video bitstream 14 by the encoder 10.
[0139] As an alternative to the above-described embodiment where an AUD is required for access units that may become IRAP AUs or GDR AUs during OLS extraction (layer dropping), when the extractor needs them, it can be confirmed that they exist, and the aud_irap_or_gdr_au_flag can be easily rewritten to 1 when appropriate (i.e., the AU becomes an IRAP or GDR AU after extraction). The encoder only needs to write an AUD NAL unit when the access unit corresponds to at least one OLS CVSS AU. An example of the specification is as follows.
[0140] There can be at most one EOB NAL unit in an AU. When vps_max_layers_minus1 is greater than 0, assume that there is only one AUD NAL unit in each AU that contains only IRAP or GDR NAL units in all layers of at least one multi-layer OLS.
[0141] That is, using the terms IRAP or GDR picture, it is as follows.
[0142] There can be at most one EOB NAL unit in an AU. When vps_max_layers_minus1 is greater than 0, assume that there is only one AUD NAL unit in each AU that contains only IRAP or GDR pictures in all layers of at least one multi-layer OLS.
[0143] In this embodiment, the related OLS extraction process will be extended through the following steps.
[0144] […] The output sub-bitstream OutBitstream is derived as follows. - The bitstream outBitstream is set to be the same as the bitstream inBitstream. - Delete all NAL units in outBitstream that have a TemporalId greater than tIdTarget. - Remove from outBitstream all NAL units whose nal_unit_type is not equal to any of VPS_NUT, DCI_NUT, and EOB_NUT and whose nuh_layer_id is not included in LayerIdInOls[targetOlsIdx]. - If the AU has two or more layers and contains only NAL units whose nal_unit_type is equal to a single type of IDR_NUT, CRA_NUT, or GDR_NUT, rewrite the flag aud_irap_or_gdr_au_flag of the AUD of the AU to 1.
[0145] Thus, the extractor 30 according to the fourth aspect can provide a sequence start indicator by setting the value of the sequence start indicator for each access unit 22*. For example, the sequence start indicator may be a syntax element signaled in an access unit of the multi-layer video bitstream 14, and the extractor 30 may modify or hold the value of the sequence start indicator when transferring the access unit in the sub-bitstream 12.
[0146] In other words, according to the embodiment, the above-described extractor 30 according to the fourth aspect, for each access unit 22 of the sub-bitstream, when all bitstream parts of each access unit are bitstream parts of the same one of a set of predetermined bitstream part types, for example, when it is the access unit 22*, when transferring each access unit 22* of the multi-layer video bitstream 14 in the sub-bitstream 12, by setting the value of the sequence start indicator existing in each access unit of the multi-layer video bitstream (for example, existing in the AUD NAL unit of each access unit 22*), for example, the aud_irap_or_gdr_flag, to a predetermined value, for example, 1, a sequence start indication indicating that each access unit is the start access unit of the encoded video sequence can be provided in the sub-bitstream 12. The predetermined value is a value indicating that each access unit is the access unit at the head of the encoded video sequence. For example, the extractor 30 may change the value of the sequence start indicator to a predetermined value when the value of the sequence start indicator does not have a predetermined value in the multi-layer video bitstream 14.
[0147] Accordingly, an embodiment of the encoder 10 for providing a multi-layer video bitstream 14 according to the fourth aspect provides an OLS representation 18 indicating an OLS including layers of the video bitstream 14 (i.e., at least two layers) in the multi-layer video bitstream 14. For each access unit including a picture of one of the predetermined picture types (e.g., the same type or not necessarily the same type) for the layers of the OLS, the encoder 10 may provide a sequence start indicator within the multi-layer video bitstream 14, and the sequence start indicator indicates whether all pictures of the access unit, i.e., those that are not part of the OLS, are of one of the predetermined picture types, e.g., by the value of the sequence start indicator.
[0148] In other words, the encoder 10 may signal a sequence start indicator for an access unit, e.g., access unit 22*, the pictures of which are pictures belonging to one of the layers of the OLS and are of one of the predetermined types.
[0149] 5. Handling of Temporal Sub-layers in the Extraction Process of the Output Layer Set In Section 5, embodiments according to the fifth aspect of the present invention are described. Embodiments according to the fifth aspect may optionally follow the embodiments of the encoder 10 and the extractor 30 described with respect to FIG. 1. Also, the details described with respect to further aspects can be optionally implemented in the embodiments described in this section.
[0150] Some embodiments according to the fifth aspect relate to the extraction process of the OLS and vps_ptl_max_temporal_id[i][j]. Some embodiments according to the fifth aspect may be related to the derivation of NumSublayerInLayer[i][j].
[0151] To extract the output layer set (OLS), it is necessary to drop or remove the layers that do not belong to the OLS from the bitstream. However, note that the number of sublayers in the layers belonging to the OLS may be different (the time layer TLx in Figure 10).
[0152] Figure 10 shows an example of a two-layer bitstream where each of the two layers has a different frame rate. The bitstream in Figure 10 includes an access unit 221 associated with the first time layer TL0 and an access unit 222 associated with the second time layer TL1. The first layer 241 includes pictures of both time sublayers TL0 and TL1, and the second layer 242 includes pictures of TL0 only. Therefore, the first layer 241 has a frame rate or picture rate that is twice that of the second layer 242.
[0153] The bitstream in Figure 10 may have two OLSs. One is composed of L0 only and has two operation points (for example, TL0 at 30fps and TL0+TL1 at 60fps). The other OLS may be composed of L0 and L1, or may also be composed of TL0 only.
[0154] In the current specification, the profile and level of the OLS can be signaled using TL0 and TL1 of L0, but in the extraction process, a bitstream of TL0 only cannot be generated.
[0155] Currently, NumSublayerInLayer[i][j], which represents the maximum number of sublayers included in the i-th OLS of layer j, is set to vps_max_sublayers_minus1 + 1 when vps_max_tid_il_ref_pics_plus1[m][k] does not exist, or when layer j is the output layer of the i-th OLS.
[0156] FIG. 11 shows an encoder 10 and an extractor 30 according to an embodiment of the fifth aspect. The encoder 10 and the extractor 30 according to FIG. 11 may optionally correspond to the encoder 10 and the extractor 30 described with respect to FIG. 1. Accordingly, the description of FIG. 1 may optionally be applied to the elements shown in FIG. 11. The encoder 10 according to FIG. 11 is configured to encode an encoded video sequence 20 into a multi-layer video bitstream 14. The multi-layer video bitstream 14 includes access units 22, such as access units 221, 222, 223 in FIG. 11 (or also in FIG. 1), each of which constitutes one or more pictures 26 of the encoded video sequence. For example, each of the access units 22 includes one or more pictures related to one common instant or frame of the encoded video sequence, as described with respect to FIG. 1. Each of the pictures 26 belongs to one of the layers 24 of the multi-layer video bitstream. For example, in the exemplary example of FIG. 11, the multi-layer video bitstream 14 includes a first layer 241 and a second layer 242, the first layer including picture 261 and the second layer including picture 262. According to the fifth aspect, each of the access units 22 belongs to a set of temporal sub-layers TL0, TL1 of the encoded video sequence 20. For example, access units 221 and 223 in FIG. 11 may belong to the first temporal sub-layer TL0, and access unit 222 may belong to the second temporal sub-layer TL1, for example, as described with respect to FIG. 10. Temporal sub-layers may also be referred to as temporal subsets or temporal layers. For example, each of the temporal sub-layers may be indicated or indexed by a temporal identifier and may be characterized, for example, by a temporal relationship with respect to the frame rate and / or other temporal subsets. For example, each of the access units may include or be associated with a temporal identifier that associates each access unit with one of the temporal sub-layers.
[0157] According to a fifth aspect, the encoder 10 is configured to provide a syntax element, such as the above-described max_tid_within_ols, to the multi-layer video bitstream 14, the syntax element indicating a predetermined temporal sublayer for an OLS, where the OLS includes or indicates a (not necessarily proper) subset of the layers of the multi-layer video bitstream 14. The syntax element indicates the predetermined temporal sublayer for the OLS in such a way as to distinguish between different states, including a state in which the predetermined temporal sublayer is below the maximum temporal sublayer within an access unit that is at least one picture of the subset of layers. For example, the predetermined temporal sublayer is the maximum temporal sublayer included in the OLS.
[0158] For example, the encoder 10 can provide an OLS indication 18 within the multi-layer video data stream 14, for example, as described with respect to FIG. 1. The OLS indication 18 can include a description or indication of one or more OLSs, such as the OLS 181 shown in FIG. 11. Each of the OLSs can be associated with a set of layers of the multi-layer video data stream 14. In the example shown in FIG. 11, the OLS 181 is associated with the first layer 241 and the second layer 242. Note that the multi-layer video bitstream 14 can optionally include additional layers.
[0159] For example, with respect to the example shown in FIG. 11, OLS 181 can include a first time sublayer (to which access units 221 and 223 can belong), but the second time sublayer to which access unit 222 can belong is not included in the OLS in the example. According to this example, TL0 may be the largest time sublayer included in the OLS. Note that the time sublayers can have a hierarchical order. In the example of FIG. 11, among the time sublayers within access unit 22, at least one picture of a subset of the layers (layer 241 which is part of the OLS), for example picture 261 in access unit 221, is at most TL1. Therefore, the largest time sublayer included in the OLS is below the largest time sublayer TL1. Thus, a given time sublayer can be indicated by showing that it is below the largest time sublayer present in an access unit that is at least partially included in the OLS. In a further example, a given time sublayer can be indicated by showing an index that identifies the given time sublayer.
[0160] Therefore, in the example of OLS 181, the given time sublayer may be the first time sublayer. Syntax elements provided in the multi-layer video bitstream 14, such as max_tid_within_ols or vps_ptl_max_temporal_id, indicate a given time sublayer for the OLS. The syntax elements distinguish different states. According to one of the states, the given time sublayer is below the largest time sublayer within an access unit that includes at least one picture of a subset of the layers. For example, in FIG. 11, the above example of OLS 181 includes layers 241 and 242. The maximum of the time sublayers within the access units of the subset of the layers of OLS 181 is the second time sublayer to which access unit 222 belongs. In an example where the second time sublayer does not belong to OLS 181, the syntax element can indicate this state.
[0161] The extractor 30 according to the fifth aspect can derive syntax elements from the multi-layer video bitstream 14, and each picture belongs to one of the layers of the OLS 181. When the picture belongs to the access units 221, 223 that are equal to or below a predetermined temporal sublayer, the sub-bitstream 12 can be provided by selectively transferring the pictures of the multi-layer video bitstream 14 in the sub-bitstream 12.
[0162] That is, when the picture is equal to or belongs to a temporal sublayer below a predetermined temporal sublayer, the extractor 30 can provide the bitstream portion of each picture 26 in the sub-bitstream 12, and in other cases, the picture can be dropped, that is, not transferred.
[0163] That is, the extractor 30 can use syntax elements in constructing the sub-bitstream 12 to exclude pictures belonging to temporal sublayers that are not included in the OLS to be decoded but are included in any of the layers indicated by the OLS from being transferred in the sub-bitstream 12.
[0164] According to an embodiment, the multi-layer video bitstream 14 indicates, for each layer of the OLS, a syntax element indicating a predetermined temporal sublayer, for example, the maximum temporal sublayer, included in each OLS. The extractor 30 can identify whether the bitstream portion belonging to the temporal sublayer of the OLS belongs to a temporal sublayer that does not belong to the OLS based on the syntax element for the layer of the OLS, and consider the bitstream portion belonging to the temporal sublayer of the OLS as a transfer target to the sub-bitstream 12.
[0165] For example, the syntax element may be part of the OLS display 18. For example, the syntax element may be part of the OLS181 it references. For example, the syntax element may be part of the video parameter set for each OLS.
[0166] In one embodiment, signaling (e.g., of the maximum temporal sublayer) is provided in the bitstream to indicate that the OLS has a maximum sublayer different from vps_max_sublayers_minus1+1 or is the largest among all the layers present in the OLS. For this purpose, the existing syntax element vps_ptl_max_temporal_id[i][j] may be reused to indicate the maximum sublayer present in the OLS.
[0167] According to some embodiments, further, NumSublayerInLayer[i][j] representing the maximum sublayer included in the i-th OLS for layer j is changed to vps_ptl_max_temporal_id[i][j] when vps_max_tid_il_ref_pics_plus1[m][k] does not exist or when layer j is the output layer of the i-th OLS.
[0168] Alternatively, a new syntax element, e.g., max_tid_within_ols, indicating the maximum sublayer within the OLS can also be added.
[0169] According to an embodiment, the encoder 10 and / or the extractor 30 are configured to derive decoder function related parameters for each picture of the multi-layer video bitstream 14 for a substream, such as the sub-bitstream 12, obtained by selectively inheriting, when each picture belongs to one of the layers of the OLS 181, and when the picture is equal to or belongs to an access unit belonging to a predetermined time sublayer. That is, the encoder 10 and / or the extractor 30 can derive decoder function related parameters of a sub-bitstream that exclusively includes pictures belonging to the time sublayer belonging to the OLS that describes the sub-bitstream. The encoder 30 or the extractor 30 can signal the function related parameters within the sub-bitstream 12. Therefore, the encoder 10 can signal the function related parameters in the multi-layer video bitstream 14. For example, the parameters related to the decoder function may include parameters as described in Section 6.
[0170] 6. Treatment of Time Sublayers in Video Parameter Signaling In Section 6, embodiments according to the sixth aspect of the present invention will be described with reference to FIGS. 11 and 1. Therefore, FIGS. 1 and 11 can be optionally applied to the embodiments according to the sixth aspect. Further details regarding additional aspects can be optionally implemented in the embodiments described in this section.
[0171] Some examples according to the sixth aspect relate to the constraints on vps_ptl_max_temporal_id[i], vps_dpb_max_temporal_id[i], vps_hrd_max_tid[i] so as to be consistent for a given OLS.
[0172] The multi-layer video bitstream 14 and / or sub-bitstream 12 described with reference to FIG. 1 and / or FIG. 11 may optionally include a video parameter set 81. The video parameter set 81 can include one or more decoder requirement sets, such as profile-tier-level-sets (PTL sets), and / or one or more buffer requirement sets, such as DPB parameter sets, and / or one or more bitstream compliance sets, such as hypothetical reference decoder (HRD) parameter sets. For example, each of the OLSs shown by the OLS display 18 can be associated with a decoder requirement set, a buffer requirement set, and a bitstream compliance set for applying the bitstream described by each OLS. The video parameter set can indicate, for each of the decoder requirement set, the buffer requirement set, and the bitstream compliance set, the maximum temporal sublayer that each set refers to, i.e., the maximum temporal sublayer of the video bitstream or video sequence that each set refers to.
[0173] FIG. 12 shows an example of a video parameter set 81 that includes a first decoder requirement set 821, a first buffer requirement set 841, and a first bitstream compliance set 861 associated with the first OLS1 of the OLS display 18. Further, according to FIG. 12, the video parameter set 81 includes a second decoder requirement set 822, a second buffer requirement set 842, and a second bitstream compliance set 862 associated with the second output layer set OLS2.
[0174] For example, each of the OLSs described by the OLS representation 18 can be associated with one of the decoder requirement set 82, the buffer requirement set 84, and the bitstream compatibility set by associating each OLS with a respective index that refers to the decoder requirement set, the buffer requirement set, and the bitstream compatibility set. According to an embodiment of the sixth aspect, the multi-layer video bitstream includes access units, and each of the access units belongs to one of a set of temporal sub-layers of the encoded video sequence encoded in the multi-layer video bitstream 14. The multi-layer video bitstream 14 according to the sixth aspect further includes a video parameter set 81 and an OLS representation 18. For each of the bitstream compatibility set 86, the buffer requirement set 84, and the decoder requirement set 82, the temporal subset representation indicates constraints on the maximum temporal sub-layer, for example, the maximum temporal sub-layer referred to by each bitstream compatibility set / buffer requirement set / decoder requirement set. For example, each of the bitstream compatibility set 86, the buffer requirement set 84, and the decoder requirement set 82 signals a syntax element indicating a respective temporal subset representation (e.g., vps_ptl_max_temporal_id for the PTL set, vps_dpb_max_temporal_id for the DPB parameter set, and vps_hrd_max_tid for the bitstream compatibility set).
[0175] As shown in FIG. 12, each of the bitstream adaptation set 86, buffer requirement set 84, and decoder requirement set 82 may include a set of one or more parameters for each temporal sublayer in a set of layers that includes the bitstream portion of the video bitstream that the respective bitstream adaptation set 86, buffer requirement set 84, or decoder requirement set 82 refers to. For example, in FIG. 12, OLS1 includes layer L0 that includes the bitstream portion of temporal layer TL0, and OLS2 includes layers L0 and L1 that include the bitstream portions of temporal layers TL0 and TL1. The bitstream adaptation set 862 and decoder requirement set 822 associated with OLS2 include parameters for L0, and thus, according to this example, include parameters for temporal layer TL0, and further include parameters for L1, and thus parameters for temporal layer TL1. The buffer requirement set 842 includes a set of parameters DPB0 for temporal sublayer TL0 and DPB1 for temporal sublayer TL1.
[0176] Normally, there are three syntax structures in the VPS, which are generally defined and then mapped to specific OLSs. · Profile Tier Level (PTL), e.g., one or more decoder requirement sets · DPB parameters, e.g., one or more buffer requirement sets · HRD parameters, e.g., one or more bitstream adaptation sets The mapping from PTL to OLS is performed for the VPS of all OLSs (single-layer or multi-layer). However, the mapping of DPB and HRD parameters to OLS is performed only for the VPS of OLSs with multiple layers. As shown in FIG. 12, the parameters of PTL, DPB, and HRD are first described in the VPS and then mapped to indicate which parameters the OLS uses.
[0177] In the example shown in FIG. 12, there are two OLSs, and each of these has two parameters. However, since the definitions and mappings are specified such that multiple OLSs can share the same parameters, it is not necessary to repeat the same information multiple times, as shown in FIG. 13 for example.
[0178] FIG. 13 shows an example where OLS2 and OLS3 have the same PTL and the same DBP parameters, but different HRD parameters.
[0179] In the examples of FIGS. 12 and 13, the values of vps_ptl_max_temporal_id[i], vps_dpb_max_temporal_id[i], and vps_hrd_max_tid[i] for a given OLS are aligned (TL0 for OLS1, TL1 for OLS2, TL1 for OLS3), but this is not currently necessary. These three values associated with the same OLS are not currently restricted to having the same value. Currently, there are no constraints on any of these values to match the number of sublayers in the bitstream. For example, in the above example, the values are defined for two sublayers, but the bitstream can have a single sublayer for OLS2 and OLS3. Therefore, as the matching becomes more complex, the decoder may not be able to easily find out what the characteristics of the bitstream are.
[0180] In the first embodiment, the bitstream signals the maximum number of sublayers present in the OLS (although it may not necessarily be a bitstream as some may have been dropped), and can be understood at least as an upper limit, that is, there cannot be more sublayers in the bitstream for an OLS than the signal value, for example, vps_ptl_max_temporal_id[i]. Therefore, the DPB and HRD parameters are also used by the decoder.
[0181] If the values of vps_dpb_max_temporal_id[i] and vps_hrd_max_tid[i] are different from vps_ptl_max_temporal_id[i], the decoder needs to perform a more complex mapping. Therefore, in one embodiment, when the OLS indexes the PTL structure, DPB structure, and HRD parameter structure with vps_ptl_max_temporal_id[i], vps_dpb_max_temporal_id[i], and vps_hrd_max_tid[i], respectively, there is a bitstream constraint that vps_dpb_max_temporal_id[i] and vps_hrd_max_tid[i] must be equal to vps_ptl_max_temporal_id[i].
[0182] According to an embodiment, the encoder 10, such as the encoder 10 in FIG. 1 or FIG. 11, is configured to form an OLS representation 18 such that the maximum temporal sublayers indicated by the bitstream compliance set 86, buffer requirement set 84, and decoder requirement set 82 related to the OLS are equal to each other, and the parameters within the bitstream compliance set 86, buffer requirement set 84, and decoder requirement set 82 are completely valid for the OLS.
[0183] However, looking at the example in FIG. 12, the parameters of OLS1 with a single sublayer (TL0), that is, for level 0 (L0 of PTL0 (TL0) DPB parameter 0 (DPB0 of DPB0) and HRD parameter 0 (HRD0 of HRD0)), are also described by PTL1 822, DPB1 842, and HRD1 862. One option to avoid repeating many parameters is to exclude DPB0 841 and HRD0 861 and obtain the values of OLS1 from the DPB parameters and HRD parameters that include more sublayers. An example of this is shown in FIG. 14.
[0184] FIG. 14 shows an example of the definitions of PTL, DPB, and HRD and the sharing between different OLSs having sublayer information not related to some OLSs. Since PTL0 indicates that there is only one sublayer, only the TL0 parameters of DPB1 and HRD1 are used.
[0185] Therefore, in another embodiment, when the OLS indexes the PTL structure, DPB structure, and HRD parameter structure with vps_ptl_max_temporal_id[i], vps_dpb_max_temporal_id[i], and vps_hrd_max_tid[i] respectively, there is a bitstream constraint that vps_dpb_max_temporal_id[i] and vps_hrd_max_tid[i] are greater than or equal to vps_ptl_max_temporal_id[i], and large values corresponding to the upper sublayers of the DPB and HRD parameters are ignored in the OLS.
[0186] Therefore, according to a further embodiment, the encoder 10 is configured to form the OLS representation 18 and / or the video parameter set 81 (or generally, the multi-layer video bitstream 14) such that the maximum temporal sublayer indicated by the decoder requirement set 82 associated with the OLS is less than or equal to the maximum temporal sublayers indicated by the buffer requirement set 84 and the bitstream conformity set 86 associated with the OLS, and the parameters within the buffer requirement set 84 and the bitstream conformity set 86 are the same for the temporal layers equal to or lower than the maximum temporal sublayer indicated by the decoder requirement set 82 associated with the OLS, and are valid only for the OLS.
[0187] In other words, the encoder 10 can provide the OLS display 18 and / or the video parameter set 81 such that the maximum temporal sublayer indicated by the buffer requirement set 84 related to the OLS is greater than or equal to the maximum temporal sublayer indicated by the decoder requirement set 82 related to the OLS, and such that the maximum temporal sublayer indicated by the bitstream adaptation set 86 related to the OLS is greater than or equal to the maximum temporal sublayer indicated by the decoder requirement set 82 related to the OLS.
[0188] For example, FIG. 14 shows an example of a video parameter set 81 that includes a first decoder requirement set 821 for a first set of layers such as layer L0 that includes access units of a first time sublayer TL0. The video parameter set 81 further includes a second decoder requirement set 822 that refers to a second set of layers, the second set of layers includes layer L0 and layer L1, and the second set of layers includes access units of a first time sublayer and a second time sublayer, i.e., TL0 and TL1. Thus, the maximum time sublayer of the second set of layers is the second time sublayer TL1. The video parameter set 81 further includes a DPB parameter set 842 that refers to the second set of layers, a bitstream adaptation set 862 that refers to the second set of layers, and a bitstream adaptation set 863 that refers to a set of a third layer L2 that has access units of the first time sublayer and the second time sublayer. An OLS1, which is a first OLS, is associated with the first set of layers, and its decoder requirements are described by a first decoder requirement set 821 that indicates that the maximum time sublayer of the first set of layers is the first time sublayer. Since the first time sublayer is less than or equal to the maximum time sublayer indicated by the DPB parameter set 842 associated with OLS1 and the bitstream adaptation set 862 associated with OLS1 (note that time sublayers are hierarchically ordered), the DPB parameter set 842 and the bitstream adaptation set 862 include information regarding the first set of layers. Thus, the DPB parameter set 842 and the bitstream adaptation set 862 are valid for OLS1 as long as they are related to the first set of layers, and their access units belong to the first time sublayer. For example, as described with respect to FIG. 12 and also shown in FIGS. 13 and 14, the decoder requirement set 822, the buffer requirement set 842, and the bitstream adaptation set 862 include sets of parameters for each time sublayer included in the bitstream they refer to.According to this embodiment, the parameters related to the first time sublayer TL0 are valid for OLS1 because the first time sublayer is below the maximum time sublayer indicated by the decoder requirement set 822.
[0189] In other words, for the OLS to be decoded, the decoder 50 may (by way of example) use only those related to a time sublayer that is the same as or smaller than the maximum time sublayer related to the decoder requirement set 82 for the OLS among the parameters of the decoder requirement set 82, buffer requirement set 84, and bitstream conformity set 86 related to the OLS.
[0190] FIG. 15 shows another example of the video parameter set 81 and the OLS display 18. FIG. 15 shows alternatives that can occur when the parameters of TL1 and TL0 are the same, or when the value of TL1 is the maximum value allowed for a certain level. In such a case, instead of not including DPB0 and HRD0 in the VPS as shown previously, both can be included without including DPB1 and HRD1. And the value of the upper sublayer of OLS1 can be derived as equal to the value signaled at TL0 or the maximum value permitted at the level. Thus, FIG. 15 illustrates the sublayer information that needs to be estimated when there is no existence for a certain OLS regarding the sharing between different OLSs with different definitions of PTL, DPB, and HRD.
[0191] Therefore, in another embodiment, there are no bitstream constraints on the values vps_ptl_max_temporal_id[i], vps_dpb_max_temporal_id[i], vps_hrd_max_tid[i], but for values greater than vps_ptl_max_temporal_id[i] where i > vps_dpb_max_temporal_id[i], the DPB and HRD parameters of vps_hrd_max_tid[i] up to vps_ptl_max_temporal_id[i] are assumed to be the maximum values defined by the profile level or equal to the highest signal DPB and HRD parameters.
[0192] Thus, according to another embodiment, the encoder 10 is configured to form the OLS representation and / or the video parameter set 81 such that the maximum temporal sublayer indicated by the decoder requirement set 82 associated with the OLS is greater than or equal to the maximum temporal sublayer indicated by each of the buffer requirement set 84 and the bitstream adaptation set 86 associated with the OLS. According to these embodiments, parameters regarding temporal sublayers that are missing within the buffer requirement set 84 and the bitstream adaptation set 86 associated with the OLS, for example, the OLS2 of FIG. 15, and that are above the maximum temporal sublayer indicated by each of the buffer requirement set 84 and the bitstream adaptation set 86, are set equal to a fourth parameter or are set equal to the parameters within the buffer requirement set 84 and the bitstream adaptation set 86 associated with the OLS associated with the maximum temporal sublayer indicated by each of the buffer requirement set 84 and the bitstream adaptation set 86.
[0193] Accordingly, embodiments of a decoder that decodes a multi-layer video bitstream, such as decoder 50 in FIG. 1, are such that the maximum temporal sublayer indicated by the decoder requirement set 82 related to OLS is greater than or equal to the maximum temporal sublayer indicated by each of the buffer requirement set 84 and the bitstream adaptation set 86 related to OLS, i.e., the parameters related to OLS in the buffer requirement set 84 and the bitstream adaptation set 86 are set to be equal to default values such as the maximum value of each parameter indicated by the decoder requirement set 82 for each parameter of the buffer requirement set 84 and the bitstream adaptation set 86, or may be configured to estimate the value of each parameter within the buffer requirement set 84 or the bitstream adaptation set 86 associated with OLS related to the maximum temporal sublayer indicated by each of the buffer requirement set 84 and the bitstream adaptation set 86. For example, the selection of whether to use default values or the values of each parameter within the buffer requirement set or the bitstream adaptation set related to the maximum temporal sublayer indicated by each of the buffer requirement set and the bitstream adaptation set may be made differently for each parameter of the buffer requirement set 84 and the bitstream adaptation set 86.
[0194] 7. Selection of Output Layer in Region of Interest Application In section 7, with reference to FIG. 1, embodiments according to the seventh aspect are described. Accordingly, the description of FIG. 1 can be arbitrarily applied to embodiments according to the seventh aspect. Also, details described regarding further aspects can be optionally implemented in the embodiments described in this section.
[0195] Some embodiments according to the seventh aspect are related to the derivation of PicOutputFlag in RoI applications.
[0196] A multi-layer bitstream such as video bitstream 14 is used. If the picture of the specified output layer is not available on the decoder side (e.g., bitstream error or transmission loss), or if certain considerations are not followed, it may result in a sub-optimal user experience. Usually, when an access unit does not contain a picture in the output layer, to compensate for such errors or losses, it is possible to select and output a picture from a non-output layer, depending on the implementation, as is clear from the annotation of the derivation of the following variable PicOutputFlag.
[0197] - The variable PictureOutputFlag of the current picture is derived as follows. - If sps_video_parameter_set_id is greater than 0 and the current layer is not the output layer (i.e., nuh_layer_id is not equal to OutputLayerIdInOls[TargetOlsIdx][i] for any value i in the range from 0 to NumOutputLayersInOls[TargetOlsIdx]-1), or if any of the following conditions is true, PictureOutputFlag is set to 0: - The current picture is a RASL picture and the NoOutputBeforeRecoveryFlag of the associated IRAP picture is equal to 1. - The current picture is a GDR picture with NoOutputBeforeRecoveryFlag equal to 1, or a recovery picture of a GDR picture with NoOutputBeforeRecoveryFlag equal to 1. - Otherwise, PictureOutputFlag is set the same as ph_pic_output_flag.
[0198] Note - In the implementation, the decoder may output pictures that do not belong to the output layer. For example, while there is only one output layer, in the AU, if the picture of the output layer is not available due to, for example, loss or layer down-switching, the decoder can set the PictureOutputFlag equal to 1 for the picture with the highest nuh_layer_id value and ph_pic_output_flag equal to 1 among all the pictures of the AU available to the decoder, and set the PictureOutputFlag equal to 0 for all other pictures of the AU available to the decoder.
[0199] However, when the bitstream is created for a region of interest (RoI) application, that is, when the upper layer depicts only a subset of the lower layer pictures (by using a scaling window), since the switching between overview and detailed display becomes very fast, it is not desirable to switch between the layers of the decoder output in a short time frame. Therefore, as part of the present invention, in one embodiment, when a scaling window that does not cover the entire picture plane is used, the decoder implementation is not permitted to freely select the output layer.
[0200] In one implementation, the decoder can output pictures that do not belong to the output layer as long as the scaling window covers the entire picture plane. For example, while there is only one output layer, in the AU, if the picture of the output layer is not available due to, for example, loss or layer down-switching, the decoder can set the PictureOutputFlag equal to 1 for the picture with the highest nuh_layer_id value and ph_pic_output_flag equal to 1 among all the pictures of the AU available to the decoder, and set the PictureOutputFlag equal to 0 for all other pictures of the AU available to the decoder.
[0201] According to an embodiment of the seventh aspect, a decoder 50 for decoding a multi-layer video bitstream, such as multi-layer video bitstream 14 or sub-bitstream 12, uses vector-based inter-layer prediction from a reference picture 261 of a second layer 241 in which prediction vectors are scaled and offset according to the relative sizes and relative positions of the predicted picture and the reference picture defined in the multi-layer video bitstream 14 to a predicted picture 262 of a first layer 242. For example, the picture 262 of layer 242 in FIG. 1 can be encoded into the multi-layer video data stream 14 using inter-layer prediction from the picture 261 of layer 241, such as the picture 261 of the same access unit 221. According to the seventh aspect, the multi-layer video bitstream 14 can include an OLS display 18 of the OLS indicating a subset of the layers of the multi-layer video bitstream 14, and the OLS includes one or more output layers including the first layer 241 and one or more non-output layers including the second layer.
[0202] If a predetermined picture 262 of the first layer 242 of the OLS, such as picture 262, is lost, the decoder 50 according to the seventh aspect is configured such that the scaling window defined for the predetermined picture coincides with the picture boundary of the predetermined picture 262, and when the scaling window defined for a further predetermined picture coincides with the picture boundary of the further predetermined picture, the predetermined picture is replaced by a further predetermined picture of the second layer 241 of the OLS in the same access unit 22 as the predetermined picture. If at least one of the scaling pictures defined for the predetermined picture does not coincide with the picture boundary of the predetermined picture, and the scaling window defined for the predetermined picture does not coincide with the picture boundary of the further predetermined picture, the decoder 50 is configured to replace the predetermined picture by other means or not to replace it at all.
[0203] 8. Further Embodiments In the previous section, although several aspects have been described as features in the context of an apparatus, it is clear that such descriptions may also be regarded as descriptions of corresponding features of a method. Although several aspects have been described as features in the context of a method, it is clear that such descriptions can also be regarded as descriptions of corresponding features regarding the functions of an apparatus.
[0204] Some or all of the method steps may be performed by (or using) a hardware device such as, for example, a microprocessor, a programmable computer, or an electronic circuit. In some embodiments, one or more of the most important method steps may be performed by such a device.
[0205] The encoded image signal of the invention may be stored on a digital storage medium or transmitted on a transmission medium such as a wireless transmission medium like the Internet or a wired transmission medium.
[0206] Depending on certain implementation requirements, embodiments of the present invention may be implemented in hardware or software or at least partially in hardware or at least partially in software. This implementation may be carried out using a digital storage medium such as a floppy disk, a DVD, a Blu-Ray, a CD, a ROM, a PROM, an EPROM, an EEPROM, or a FLASH memory having electronically readable control signals, which can be stored thereon and cooperate (or be capable of cooperating) with a programmable computer system so that each method is executed. Thus, the digital storage medium may be computer-readable.
[0207] Some embodiments according to the present invention include a data carrier having electronically readable control signals, and by these control signals being able to cooperate with a programmable computer system, one of the methods described herein is executed.
[0208] Generally, embodiments of the present invention can be implemented as a computer program product having program code, which is operative to perform one of the methods when the computer program product is executed on a computer. The program code can be stored, for example, in a machine-readable carrier.
[0209] Other embodiments include a computer program that performs one of the methods described herein and is stored in a machine-readable carrier.
[0210] In other words, one embodiment of the method of the present invention is thus a computer program having program code for performing one of the methods described herein when the computer program is executed on a computer.
[0211] A further embodiment of the method of the present invention is thus a data carrier (or digital storage medium, or computer-readable medium), which includes a computer program for performing one of the methods described herein recorded thereon. The data carrier, digital storage medium, or recorded medium is typically tangible and / or non-transitory.
[0212] A further embodiment of the method of the present invention is thus a data stream or signal sequence representing a computer program for performing one of the methods described herein. The data stream or signal sequence can be configured to be transferred via a data communication connection, such as via the Internet.
[0213] Further embodiments include, for example, a computer or processing means, such as a programmable logic device, configured or adapted to perform one of the methods described herein.
[0214] A further embodiment includes a computer on which a computer program for performing one of the methods described herein is installed.
[0215] A further embodiment according to the present invention includes an apparatus or system configured to transfer (e.g., electronically or optically) a computer program for performing one of the methods described herein to a receiver. The receiver may be, for example, a computer, a mobile device, a memory device, etc. The apparatus or system may include, for example, a file server for transferring the computer program to the receiver.
[0216] In some embodiments, a programmable logic device (e.g., a field programmable gate array) may be used to perform some or all of the functions of the methods described herein. In some embodiments, a field programmable gate array may cooperate with a microprocessor to perform one of the methods described herein. Generally, the methods are preferably performed by any hardware device.
[0217] The apparatuses described herein may be implemented using a hardware device, or using a computer, or using a combination of a hardware device and a computer.
[0218] The methods described herein may be performed using a hardware device, or using a computer, or using a combination of a hardware device and a computer.
[0219] In the foregoing detailed description, it can be seen that various features are grouped in examples for the purpose of streamlining the disclosure. The present method of disclosure is not to be construed as reflecting an intention that the examples recited in the claims require more features than are explicitly recited in each claim. Rather, as the following claims reflect, the subject matter may lie in fewer features than all of the single disclosed example. Accordingly, the following claims are hereby incorporated by reference into the detailed description herein, and each claim can stand on its own as a separate example. While each claim can stand on its own as a separate example, dependent claims may in some cases refer to a particular combination with one or more other claims in the claim, but it should be noted that other examples may also include combinations of the subject matter of the dependent claim with each other dependent claim, or combinations of each feature with other dependent or independent claims. Such combinations are proposed herein unless it is stated that a particular combination is not intended. Further, even if a claim is not directly dependent on an independent claim, it is intended to include the features of the claim with respect to other independent claims.
[0220] The above embodiments are merely illustrative of the principles of the present disclosure. It will be understood that modifications and variations of the configurations and details described herein will be apparent to those skilled in the art. Accordingly, it is intended to be limited only by the scope of the pending claims, rather than by the specific details presented as descriptions and explanations of the embodiments herein.
Claims
1. A decoder (50) for decoding a video bit stream (14), comprising: a processor; a memory for storing instructions, which, when executed by the processor, cause the processor to: obtain a video bit stream (14) including access units (22), each of the access units (22) including one or more pictures (26), and each of the one or more pictures corresponding to one of a plurality of layers (24) of the video bit stream (14); derive one or more output layer sets (OLS) (181, 182) from the video bit stream (14), each indicating a set of layers from the plurality of layers of the video bit stream (14); identify a first access unit of the access units (22); when external information for identifying the OLS to be decoded from among the one or more OLSs is not available for the first access unit, identify the OLS to be decoded based on the number of layers of each of the one or more OLSs; decode the identified OLS; and the memory for causing the above operations; the decoder (50) including the above components.
2. To determine the OLS to be decoded, the instructions, when executed by the processor, cause the processor to identify the maximum or minimum value of the index of the OLS. The decoder (50) according to Claim 1.
3. To determine the OLS to be decoded, the instructions, when executed by the processor, cause the processor to identify the largest number of layers of the OLS. The decoder (50) according to Claim 1.
4. The external information exists outside the video bit stream. The decoder (50) according to Claim 1.
5. When the instructions are executed by the processor, the instructions further cause the processor to: obtain external information separate from the video bit stream; identify the OLS to be decoded based on the external information; and perform the above operations. The decoder (50) according to Claim 1.
6. When the instructions are executed by the processor, the instructions further cause the processor to: determining whether to obtain additional information indicating an operation point; identifying the operation point to be decoded in response to obtaining the additional information; selecting the OLS to be decoded in response to not obtaining the additional information; causing to perform; the decoder (50) according to claim 1.
7. The first access unit is at the head of one of the encoded video sequences (20) and represents at least one of the first in temporal order, the first in decoding order, or the first in output order. the decoder (50) according to claim 1.
8. The first access unit is of a sequence start access unit type. the decoder (50) according to claim 1.
9. The first access unit is at the head of the video bitstream. the decoder (50) according to claim 8.
10. A method for decoding a video bitstream, comprising: obtaining a video bitstream (14) including access units (22), each of the access units (22) including one or more pictures (26), and each of the one or more pictures corresponding to one of the plurality of layers (24) of the video bitstream (14); deriving one or more output layer sets (OLSs) (181, 182) from the video bitstream (14), each indicating a set of layers from the plurality of layers of the video bitstream (14); identifying a first access unit of the access units (22); when external information for identifying the OLS to be decoded from among the one or more OLSs is not available for the first access unit, identifying the OLS to be decoded based on the number of layers of each of the one or more OLSs; decoding the identified OLS; the method including.
11. Determining the identified OLS includes identifying the highest or lowest value of the index of the OLS. The method according to claim 10.
12. Determining the identified OLS includes identifying the largest number of layers of the OLS. The method according to claim 10.
13. Determining the identified OLSs is based on at least one of the index of each of the OLSs, the number of output layers of each of the OLSs, or the number of layers of each of the OLSs. The method according to claim 10.
14. Obtaining external information separate from the video bitstream, Identifying the OLSs to be decoded based on the external information, further comprising The method according to claim 10.
15. Determining whether to obtain additional information indicating an operation point, Identifying the operation point to be decoded in response to obtaining the additional information, Selecting the OLSs to be decoded in response to not obtaining the additional information, further comprising The method according to claim 10.
16. The first access unit is at the beginning of one of the encoded video sequences (20) and represents at least one of the first in temporal order, the first in decoding order, or the first in output order. The method according to claim 10.
17. The first access unit is of a sequence start access unit type. The method according to claim 10.
18. The first access unit is at the beginning of the video bitstream. The method according to claim 17.
19. A non-transitory computer-readable medium storing a computer program, wherein when the computer program is executed by a processor of a decoder, the decoder is caused to obtain a video bitstream (14) including access units (22), each of the access units (22) including one or more pictures (26), and each of the one or more pictures corresponding to one of the plurality of layers (24) of the video bitstream (14), the obtaining; derive one or more output layer sets (OLSs) (181, 182) from the video bitstream (14), each indicating a set of layers from the plurality of layers of the video bitstream (14), the deriving; identifying a first access unit of the access units (22), For the first access unit, when external information for identifying the OLS to be decoded from among the one or more OLSs is not available, identifying the OLS to be decoded based on the number of layers of each of the one or more OLSs; decoding the identified OLS; The non-transitory computer-readable medium that causes the above to be performed.