Processing of output layer sets of coded video

The method addresses the challenge of handling multi-layered video bitstreams by deriving syntax elements for OLSs and using interlayer prediction to maintain decoding consistency and quality, particularly in error-prone scenarios.

AU2026205066A1Pending Publication Date: 2026-07-16FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
AU · AU
Patent Type
Applications
Current Assignee / Owner
FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
Filing Date
2026-06-29
Publication Date
2026-07-16

AI Technical Summary

Technical Problem

Existing video coding standards fail to efficiently handle multi-layered video bitstreams, particularly in scenarios where there is a requirement for precise extraction of Output Layer Sets (OLS) and handling of temporal sublayers, leading to suboptimal decoding and user experience when errors or losses occur.

Method used

Implementing a method to derive and signal syntax elements in the video bitstream to indicate the maximum temporal sublayer within an OLS, ensuring consistent decoder capability parameters, and using vector-based interlayer prediction for picture substitution in case of layer losses, while maintaining consistent scaling and offsetting predictions.

Benefits of technology

Ensures accurate extraction of OLSs and improved decoding performance by aligning decoder requirements with bitstream constraints, enhancing user experience by minimizing rapid layer switching and maintaining video quality during errors or losses.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000063_0000
    Figure 00000063_0000
  • Figure 00000064_0000
    Figure 00000064_0000
  • Figure 00000065_0000
    Figure 00000065_0000
Patent Text Reader

Abstract

Aspects of output layer sets in video coding Abstract An output layer set may describe a video bitstream, or video sequence, which is extractable 5 from a multi-layered video bitstream. Concepts related to output layer sets are described, including a concept for a random-accessible sub-bitstream indication, a concept for handling reference picture alignment, a concept for a bitstream-based OLS determination, a concept for sequence start access units of an extracted video bitstream, a concept for handling temporal sublayers in the extraction process of an output layer set, a concept for 10 handling temporal sublayers in video parameter signaling and a concept for output layer selection in region of interest applications. (Figure 1) Abstract 5 10 20 26 20 50 66 29 J un 2 02 6 2 0 2 6 2 0 5 0 6 6 2 9 J u n 2 0 2 6 A b s t r a c t
Need to check novelty before this filing date? Find Prior Art

Description

2026205066   29 Jun 2026 There can be at most one EOB NAL unit in an AU, and when vps_max_layers_minus1 is greater than 0, there shall be one and only one AUD NAL unit in each AU that contains only IRAP or GDR NAL units in all layers of at least one multi-layer OLS. 5 In other words, using the term IRAP or GDR pictures: There can be at most one EOB NAL unit in an AU, and when vps_max_layers_minus1 is greater than 0, there shall be one and only one AUD NAL unit in each AU that contains only IRAP or GDR pictures in all layers of at least one multi-layer OLS. 10 In this embodiment, the related OLS extraction process would be extended through the following step: [...] The output sub-bitstream OutBitstream is derived as follows: 15       - The bitstream outBitstream is set to be identical to the bitstream inBitstream. - Remove from outBitstream all NAL units with TemporalId greater than tIdTarget. - Remove from outBitstream all NAL units with nal_unit_type not equal to any of VPS_NUT, DCI_NUT, and EOB_NUT and with nuh_layer_id not included in the list LayerIdInOls[ targetOlsIdx ]. 20      - When an AU contains only NAL units with nal_unit_type equal to a single type of IDR_NUT, CRA_NUT or GDR_NUT in two or more layers, re-write the flag aud_irap_or_gdr_au_flag to be equal to 1 for the AUD of the AU. Accordingly, the above described extractor 30 according to the fourth aspect may, provide 25 the sequence start indication by setting a value of a sequence start indicator for the respective access unit 22*. For example, the sequence start indicator may be syntax element signaled in the access unit of the multi-layered video bitstream 14, and extractor 30 may amend or keep the value of the sequence start indicator when forwarding the access unit in the sub-bitstream 12. 30 In other words, according to embodiments, the above described extractor 30 according to the fourth aspect may, for each of the access units 22 of the sub-bitstream, if all bitstream portions of the respective access unit are bitstream portions of the same out of a set of predetermined bitstream portion types, e.g. access units 22*, provide within the sub-35 bitstream 12 a sequence start indication indicating the respective access unit to be a starting 2026205066   29 Jun 2026 access unit of a subsequence of the coded video sequence by, in forwarding the respective access unit 22* of the mulit-layered video bitstream 14 in the sub-bitstream 12, set a value of a sequence start indicator, e.g. the aud_irap_or_gdr_flag, present in the respective access unit of the multi-layered video bitstream, e.g. present in an AUD NAL unit of the 5 respective access unit 22*, to a predetermined value, e.g. 1, the predetermined value indicating that the respective access unit to be a starting access unit of a subsequence of the coded video sequence. For example, extractor 30 may change the value of the sequence start indicator to be the predetermined value, if it does not have the predetermined value in the multi-layered video bitstream 14. 10 Accordingly, embodiments of the encoder 10 for providing the multi-layered video bitstream 14 according to the fourth aspect, provides, in the multi-layered video bitstream 14, an OLS indication 18 indicating an OLS including layer of the video bitstream 14, i.e. at least two layers. For each access unit which comprises pictures of one of the predetermined picture 15 types (e.g. the same type, or not necessarily the same type) for the layers of the OLS, the encoder 10 may provide the sequence start indicator in the mulit-layered video bitstream 14, the sequence start indicator indicating, whether all picture of the access unit, i.e. also those which are not part of the OLS, are of one of the predetermined picture types, e.g. by means of a value of the sequence start indicator. 20 In other words, encoder 10 may signal the sequence start indicator for access units, e.g. access units 22*, the pictures of which access units, which pictures belong to one of the layers of the OLS, are of one of the predetermined type. 25 5. Handling of temporal sublayers in the extraction process of an output layer set Section 5 describes embodiments in accordance with the fifth aspect of the invention. Embodiments according to the fifth aspect may optionally be in accordance with embodiments of the encoder 10 and the extractor 30 as described with respect to Fig. 1. 30 Also details described with respect to the further aspects may optionally be implemented in embodiments described in this section. Some embodiments according to the fifth aspect are related to an extraction process of an OLS and vps_ptl_max_temporal_id[ i ][ j ]. Some embodiments according to the fifth spect 35 may relate to the derivation of NumSublayerInLayer[ i ][ j ]. 2026205066   29 Jun 2026 In order to extract an Output Layer Set (OLS), it is necessary to drop or remove the layers that do not belong to the OLS from the bitstream. However, note that layers belonging to an OLS might have different amount of sublayers (temporal layer TLx in Fig. 10). 5 Fig. 10 illustrates an example of a two-layer bitstream wherein each of the two layers has a different frame rate. The bitstream of Fig. 10 comprises access unit 221 associated with a first temporal layer TL0 and access units 222 associated with a second temporal layer TL1. The first layer 241 includes pictures of both temporal sublayers TL0, TL1, whereas the second layer 242 includes pictures of TL0 only. Thus, the first layer 241 has the double frame 10 rate or picture rate as the second layer 242. The bitstream of Fig. 10 could have two OLS: one consisting of L0 only, which would have two operating points (e.g., 30fps TL0 and 60fps TL0+TL1). The other OLS could consist of L0 and L1 but only of TL0. 15 The current specification allows to signal the profile and level of the OLS with TL0 and TL1 for L0 but the extraction process misses to generate a bitstream that only has TL0. Currently, NumSublayerInLayer[ i ] [ j ] representing the maximum sublayer included in the 20 i-th OLS for layer j is set to vps_max_sublayers_minus1 +   1 when vps_max_tid_il_ref_pics_plus1[ m ][ k ]is not present or layer j is an output layer in the i-th OLS. Fig. 11 illustrates an encoder 10 and an extractor 30 according to embodiments of the fifth 25 aspect. Encoder 10 and extractor 30 according to Fig. 11 may optionally correspond to the encoder 10 and the extractor 30 as described with respect to Fig. 1. Accordingly, the description of Fig. 1 may optionally also apply to the elements shown in Fig. 11. Encoder 10 according to Fig. 11 is configured for encoding a coded video sequence 20 into a multilayered video bitstream 14. The multi-layered video bitstream 14 comprises access units 30 22, eg. access units 221, 222, 223 in Fig. 11 (or also Fig. 1), each of which comprises one or more pictures 26 of the coded video sequence. For example, each of the access units 22 comprises one or more pictures related to one common temporal instant or frame of the coded video sequence as described with respect to Fig. 1. Each of the pictures 26 belongs to one of layers 24 of the multi-layered video bitstream. For example, in the illustrative 35 example of Fig. 11, the multi-layered video bitstream 14 comprises a first layer 241 and a second layer 242, the first layer comprising pictures 261, and the second layer comprising 2026205066   29 Jun 2026 pictures 262. According to the fifth aspect, each of the access units 22 belongs to a temporal sublayer of a set of temporal sublayers TL0, TL1 of the coded video sequence 20. For example, access units 221 and 223 of Fig. 11 may belong to a first temporal sublayer TL0, and access unit 222 may belong to a second temporal sublayer TL1, e.g. as described with 5 respect to Fig. 10. Temporal sublayers may also be referred to as temporal subsets or temporal layers. For example, each of the temporal sublayers is indicated or indexed with a temporal identifier and may be characterized, for example, by a frame rate and / or a temporal relation with respect to the other temporal subsets. For example, each of the access units may comprise or may have associated therewith a temporal identifier which 10 associates respective access units with one of the temporal sublayers. According to the fifth aspect, encoder 10 is configured for providing the multi-layered video bitstream 14 with a syntax element, e.g. max_tid_within_ols as described above, the syntax element indicating a predetermined temporal sublayer for an OLS, the OLS comprising or 15 indicating a (not necessarily proper) subset of layers of the multi-layered video bitstream 14. The syntax element indicates the predetermined temporal sublayer of the OLS in a manner discriminating between different states including a state according to which the predetermined temporal sublayer is beneath a maximum of temporal sublayers within access units of which a picture of at least one of the subset of layers is. For example, the 20 predetermined temporal sublayer is the maximum temporal sublayer included in the OLS. For example, the encoder 10 may provide within the multi-layered video data stream 14 the OLS indication 18, for example as described with respect to Fig. 1. The OLS indication 18 may comprise a description or an indication for one or more OLSs, e.g. OLS 181 as shown 25 in Fig. 11. Each of the OLSs may be associated with a set of layers of the multi-layered video data stream 14. In the example illustrated in Fig. 11, OLS 181 is associated with the first layer 241 and the second layer 242. Note that the multi-layered video bitstream 14 may optionally comprise further layers. 30 For example, with respect to the illustrated example of Fig. 11, OLS 181 may include the first temporal sublayer (to which access units 221 and 223 may belong), but the second temporal sublayer, to which the access units 222 may belong, may, in examples, not be included in the OLS, so that according to this example TL0 may be the maximum temporal sublayer included in the OLS. It is noted that the temporal sublayers may have a hierarchical 35 order. In the example of Fig. 11, the maximum of temporal sublayers within access units 22 of which a picture, e.g. picture 261 if access unit 221, of at least one of the subset of layers 2026205066   29 Jun 2026 (layer 241 which is part of the OLS) is, is TL1. Accordingly, the maximum temporal sublayer included in the OLS is beneath the maximum temporal sublayer TL1. Thus, the predetermined temporal sublayer may, for example, be indicated by indicated that the predetermined temporal sublayer is beneath the maximum temporal sublayer present in 5 access units included at least partially in the OLS. In further example, the predetermined temporal sublayer may be indicated by indicating an index which identifies the predetermined temporal sublayer. Thus, in the example of the OLS 181, the predetermined temporal sublayer may be the first 10 temporal sublayer. The syntax element provided in the multi-layered video bitstream 14, e.g. max_tid_within_ols or vps_ptl_max_temporal_id, indicates the predetermined temporal sublayer for an OLS. The syntax element discriminates between different states. According to one of the states, the predetermined temporal sublayer is beneath a maximum of temporal sublayers within access units of which a picture of at least one of the subset of 15 layers is. For example, in Fig. 11, the example of the OLS 181 as described above, includes layers 241 and 242. The maximum of temporal sublayers within access units of the subset of layers of OLS 181 is the second temporal sublayer, to which access units 222 belong. In examples, in which the second temporal sublayer does not belong to OLS 181, the syntax element may be indicative of this state. 20 Extractor 30 according to the fifth aspect may derive the syntax element from the multilayered video bitstream 14 and may provide the sub-bitstream 12 by selectively forwarding the pictures of the multi-layered video bitstream 14 in the sub-bitstream 12 if the respective picture belongs to one of the layers of the OLS 181, and if the picture belongs to an access 25  unit 221, 223 that belongs to a temporal sublayer equal to, or beneath, the predetermined temporal sublayer. That is, extractor 30 may provide the bitstream portions of the respective picture 26 in the sub-bitstream 12 if the picture belongs to a temporal sublayer equal to, or beneath, the 30 predetermined temporal sublayer, and may drop, i.e. not forward, the picture otherwise. In other words, extractor 30 may use the syntax element in the construction of the subbitstream 12 for excluding pictures which belong to temporal sublayers which are not part of the OLSs to be decoded, but which are part of one of the layers indicated by the OLS, 35 from being forwarded in the sub-bitstream 12. 2026205066   29 Jun 2026 According to embodiments, the multi-layered video bitstream 14 indicates, for each of the layers of the OLS, a syntax element which indicates the predetermined temporal sublayer, e.g. the maximum temporal sublayer, included in the respective OLS. The extractor 30 may, based on the syntax elements for the layers of the OLS, discriminate between bitstream 5 portions belonging to a temporal sublayer of the OLS belonging to a temporal sublayer which is not part of the OLS, and consider those bitstream portions for forwarding into the sub-bitstream 12 which belong to a temporal sublayer of the OLS. For example, the syntax element may be part of the OLS indication 18, for example, the 10 syntax element may be part of the OLS 181 to which it refers. For example, the syntax element may be part of a video parameter set for the respective OLS. In one embodiment the signaling (e.g. of the maximum temporal sublayer) is provided into the bitstream to indicate that an OLS has a maximum sublayer different to 15 vps_max_sublayers_minus1 + 1 or to the maximum among all layers present in the OLS. For this purpose, the existing syntax element vps_ptl_max_temporal_id[ i ][ j ] may be repurposed to indicate also the maximum sublayer present in an OLS. According to some embodiments, in addition NumSublayerInLayer[ i ][ j ] which represents 20 the maximum sublayer included in the i-th OLS for layer j is changed to vps_ptl_max_temporal_id[ i ][ j ] when vps_max_tid_il_ref_pics_plus1[ m ][ k ] is not present or layer j is an output layer in the i-th OLS. Alternatively, a new syntax element could be added that indicates the maximum sublayer 25 within an OLS, e.g. max_tid_within_ols. According to embodiments, encoder 10 and / or extractor 30 are configured for deriving, for a substream, e.g. the sub-bitstream 12, which is obtained by selectively taking over, for each of the pictures of the multi-layered video bitstream 14, the respective picture, if the 30 picture belongs to one of the layers of the OLS 181, and if the picture belongs to an access unit that belongs to a temporal sublayer equal to, or beneath, the predetermined temporal sublayer, decoder capability-related parameters. In other words, encoder 10 and / or extractor 30 may derive the decoder capability-related parameters for a sub-bitstream which exclusively comprises pictures which belong to temporal sublayers belonging to the OLS 35 describing the sub-bitstream. Encoder 30 or extractor 30 may signal the capability-related parameters in the sub-bitstream 12. Accordingly, encoder 10 may signal the capability- 2026205066   29 Jun 2026 related parameters in the multi-layered video bitstream 14. For example, the decoder capability-related parameters may include parameters as described in section 6. 6. Handling of temporal sublayers in video parameter signaling 5 Section 6 describes embodiments in accordance with the sixth aspect of the invention, making reference to Fig. 11 and Fig. 1. Thus, the description of Figs. 1 and 11 may optionally apply to the embodiments in accordance with the sixth aspect. Also details described with respect to the further aspects may optionally be implemented in 10 embodiments described in this section. Some examples in accordance with the sixth aspect relate to a constraint on vps_ptl_max_temporal_id[ i ], vps_dpb_max_temporal_id[ i ], vps_hrd_max_tid[ i ] to be consistent for a given OLS. 15 The multi-layered video bitstream 14 and / or the sub-bitstream 12 as described with respect to Fig. 1 and / or Fig. 11 may optionally include a video parameter set 81. The video parameter set 81 may include one or more decoder requirement sets, e.g. profile-tier-level-sets (PTL sets), and / or one or more buffer requirement sets, e.g. DPB parameter sets, 20 and / or one or more bitstream conformance sets, e.g. hypothetical reference decoder (HRD) parameter sets. For example, each of OLSs indicated in the OLS indication 18 may be associated with each one of the decoder requirement sets, buffer requirement sets, and bitstream conformance sets applying for the bitstream described by the respective OLS. The video parameter set may indicate, for each of the decoder requirement sets, buffer 25 requirement sets, and bitstream conformance sets a maximum temporal sublayer to which the respective set refers, i.e. a maximum temporal sublayer of a video bitstream or video sequence to which the respective set refers. Fig. 12 illustrates an example of a video parameter set 81 comprising a first decoder 30   requirement set 821, a first buffer requirement set 841 and a first bitstream conformance set 861 which are associated with a first OLS1 of the OLS indication 18. Further, according to Fig. 12, the video parameter set 81 comprises a second decoder requirement set 822, a second buffer requirement set 842 and a second bitstream conformance set 862 which are associated with a second output layer set OLS2. 35 2026205066   29 Jun 2026 For example, each of the OLSs described by the OLS indication 18 may be associated with one of the decoder requirement sets 82, the buffer requirement sets 84 and the bitstream conformance sets 86 by having associated to the respective OLS respective indices pointing to the decoder requirement set, the buffer requirement set and the bitstream 5 conformance set. According to embodiments of the sixth aspect, the multi-layered video bitstream comprises access units, each of which belongs to one of a temporal sublayer of a set of temporal sublayers of a coded video sequence coded into the multi-layered video bitstream 14. The multi-layered video bitstream 14 according to a sixth aspect further comprises the video parameter set 81 and the OLS indication 18. For each of the bitstream 10 conformance sets 86, the buffer requirement sets 84, and the decoder requirement sets 82, a temporal subset indication is indicative of a constraint on a maximum temporal sublayer, e.g. a maximum temporal sublayer to which the respective bitstream conformance set / buffer requirement set / decoder requirement set refers. For example, each of the bitstream conformance sets 86, the buffer requirement sets 84, and the decoder requirement set 82 15 signal a syntax element indicating the respective temporal subset indication (e.g. vps_ptl_max_temporal_id for the PTL sets, vps_dpb_max_temporal_id for the DPB parameter sets, and vps_hrd_max_tid for the bitstream conformance sets). As illustrated in Fig 12, each of the bitstream conformance sets 86, the buffer requirement 20 sets 84, and the decoder requirement set 82 may comprise a set of one or more parameters for each temporal sublayer present in the layer set of layers which comprise bitstream portions of the video bitstream to which the respective bitstream conformance sets 86, buffer requirement sets 84, or decoder requirement set 82 refers. E.g., in Fig. 12, OLS1 includes layer L0 which comprises bitstream portions of the temporal layer TL0, and OLS2 25 includes layers L0 and L1 which comprise bitstream portions of temporal layers TL0 and TL1. The bitstream conformance set 862 and the decoder requirement set 822 which are associated with OLS2 include parameters for L0, and thus, according to this example, for the temporal layer TL0, and further include parameters for L1, and thus for the temporal layer TL1. The buffer requirement set 842, includes sets of parameters DPB0 for temporal 30 sublayer TL0 and DPB1 for temporal sublayer TL1. Conventionally, there are three syntax structures in the VPS that are defined generally and subsequently mapped to a specific OLS: • Profile-tier-level (PTL), e.g. one or more decoder requirement sets 35      • DPB parameters, e.g. one or more buffer requirement sets • HRD parameters, e.g. one or more bitstream conformance sets 2026205066   29 Jun 2026 The mapping of PTL to OLSs is done in the VPS for all OLS (with single layer or with multilayer). However, the mapping for the DPB and HRD parameters to OLS is only done in the VPS for OLS with more than one layer. As illustrated in Fig. 12, the parameters for PTL, DPB and HRD are described in the VPS first and then OLSs are mapped to indicate 5 which parameter they use. In the example shown in Fig. 12 there are 2 OLS and 2 of each of these parameters. The definition and mapping has been however specified to allow more than one OLS to share the same parameters and thus not require repeating the same information multiple times, 10 as for instance illustrated in Fig. 13. Fig. 13 illustrates an example where OLS2 and OLS3 have the same PTL and DBP parameters but different HRD parameters. 15 In the examples of Fig. 12 and Fig. 13, the values of vps_ptl_max_temporal_id[ i ], vps_dpb_max_temporal_id[ i ], vps_hrd_max_tid[ i ] for a given OLS are aligned (TL0 for OLS1, TL1 for OLS2, TL1 for OLS2), but this is currently not necessary. These three values that are associated to the same OLS are not currently restricted to have the same value. Currently, none of these values is constraint in any manner to be consistent or match the 20 number of sublayers in a bitstream. For instance, in the example above, the bitstream could have a single sublayer for OLS 2 and 3 although values are defined for two sublayers. Therefore, a decoder would not find it easily what are the characteristics of the bitstream as the matching becomes more complicated. 25 In a first embodiment, the bitstream signals the maximum number of sublayers that are present in an OLS (not necessarily the bitstream as some might have been dropped), but at least it can be understood as an upper bound, i.e. no more sublayers can be present for a OLS in the bitstream than the signal value, e.g. vps_ptl_max_temporal_id [ i ]. Thus, also DPB and HRD parameters are used by the decoder. 30 If the values of vps_dpb_max_temporal_id[ i ], vps_hrd_max_tid[ i ] are different to vps_ptl_max_temporal_id [ i ] the decoders would need to carry out a more complicated mapping. Therefore, in one embodiment there is a bitstream constraint that if a OLS indexes a PTL structure, DPB structure and HRD parameters structure with 35 vps_ptl_max_temporal_id[ i ], vps_dpb_max_temporal_id[ i ], vps_hrd_max_tid[ i ] 2026205066   29 Jun 2026 respectively, vps_dpb_max_temporal_id[ i ], vps_hrd_max_tid[ i ] shall be equal to vps_ptl_max_temporal_id[ i ]. According to an embodiment, the encoder 10, e.g. the encoder 10 of Fig. 1 or Fig. 11, is 5 configured for forming the OLS indication 18 such that the maximum temporal sublayers indicated by the bitstream conformance set 86, the buffer requirement set 84, and the decoder requirement set 82 associated with the OLS are equal to each other, and parameters within the bitstream conformance set 86, the buffer requirement set 84, and the decoder requirement set 82 are valid for the OLS completely. 10 However, looking at the example in Fig. 12, the parameters for OLS1 having a single sublayer (TL0), i.e. the level 0 (L0 (TL0) in PTL0) DPB parameters 0 (DPB0 in DPB 0) and HRD parameters 0 (HRD0 in HRD 0) are also described in PTL 1 822, DPB 1 842 and HRD 1 862. In order to not repeat so many parameters, one option would be to not include DPB 15 0 841 and HRD 0 861, and take the values for OLS1 from DPB parameters and HRD parameters that include more sublayers. An example is illustrated in Fig. 14. Fig. 14 illustrates an example of PTL, DPB and HRD definition and sharing among different OLS with sublayer information irrelevant for some OLSs. Since PTL0 indicates that there is 20 only one sublayer only the parameters for TL0 of DPB1 and HRD1 would be used. Therefore, in another embodiment there is a bitstream constraint that if a OLS indexes a PTL structure, DPB structure and HRD parameters structure with vps_ptl_max_temporal_id[ i ], vps_dpb_max_temporal_id[ i ], vps_hrd_max_tid[ i ] 25 respectively, vps_dpb_max_temporal_id[ i ], vps_hrd_max_tid[ i ] shall be greater than or equal to vps_ptl_max_temporal_id[ i ] and greater values corresponding to higher sublayers for DPB and HRD parameters are ignored for the OLS. Accordingly, according to a further embodiment, the encoder 10 is configured for forming 30 the OLS indication 18 and / or the video parameter set 81 (or, in general, the multi-layered video bitstream 14) such that the maximum temporal sublayer indicated by the decoder requirement set 82 associated with the OLS is smaller than or equal to the maximum temporal sublayer indicated by each of the buffer requirement sets 84 and the bitstream conformance set 86 associated with the OLS, and the parameters within the buffer 35 requirement set 84 and the bitstream conformance set 86 are valid for the OLS only as far 2026205066   29 Jun 2026 as same relate to temporal layers equal to and beneath the maximum temporal sublayer indicated by the decoder requirement set 82 associated with the OLS. In other words, the encoder 10 may provide the OLS indication 18 and / or the video 5 parameter set 81 so that the maximum temporal sublayer indicated by the buffer requirement set 84 associated with the OLS is greater than or equal to the maximum temporal sublayer indicated by the decoder requirement set 82 associated with the OLS and so that the maximum temporal sublayer indicated by the bitstream conformance set 86 associated with the OLS is greater than or equal to the maximum temporal sublayer 10 indicated by the decoder requirement set 82 associated with the OLS. For example, Fig. 14 illustrates an example of a video parameter set 81 comprising a first decoder requirement set 821 for a first set of layers, e.g. layer L0, comprising access units of a first temporal sublayer, TL0. The video parameter set 81 further comprises a second 15 decoder requirement set 822 referring to a second set of layers, the second set of layers comprising layer L0 and layer L1, the second set of layers including access units of the first temporal sublayer and a second temporal sublayer, i.e. TL0 and TL1. Thus, the maximum temporal sublayer of the second set of layers is the second temporal sublayer TL1. The video parameter set 81 further comprises a DPB parameter set 842 referring to the second 20 set of layers, a bitstream conformance set 862 referring to the second set of layers, and a bitstream conformance set 863 which refers to a third set of layers comprising access units of the first temporal sublayer and a third layer, L2, having access units of the second temporal sublayer. A first OLS, OLS1, is associated with the first set of layers, the decoder requirements of which are described by the first decoder requirement set 821 indicating a 25 maximum temporal sublayer of the first set of layers being the first temporal sublayer. As the first temporal sublayer is smaller than or equal to (note that the temporal sublayers are hierarchically ordered) than the maximum temporal sublayers indicated by the DPB parameter set 842 which is associated to OLS1, and the bitstream conformance set 862 which is associated with OLS1, the DPB parameter set 842 and the bitstream conformance 30 set 862 comprise information about the first set of layers. Thus, the DPB parameter set 842 and the bitstream conformance set 862 are valid for OLS1 as far as they relate to the first set of layers, the access units of which belong to the first temporal sublayer. For example, as described with respect to Fig 12., and also illustrated in Fig. 13 and Fig. 14, the decoder requirement set 822, the buffer requirement set 842, and the bitstream conformance set 862 35 include sets of parameters for each of the temporal sublayers included in the bitstream to which they refer. According to this embodiment, the parameters relating to the first temporal 2026205066   29 Jun 2026 sublayer TL0 are valid for OLS1, as the first temporal sublayer is equal to or smaller than the maximum temporal sublayer indicated by the decoder requirement set 822. In other words, decoder 50 may use, for the OLS to be decoded, those (and in examples 5 only those) parameters of the decoder requirement set 82, the buffer requirement set 84, and the bitstream conformance set 86 associated with the OLS which relate to a temporal sublayer which is equal to or smaller than the maximum temporal sublayer associated with the decoder requirement set 82 for the OLS. 10 Fig. 15 illustrates another example of a video parameter set 81 and an OLS indication 18. Fig. 15 illustrates an alternative which may occur when the parameters for TL1 and TL0 are the same or when the values for TL1 are the maximum ones that are allowed for a level. In such a case, instead of not including DPB0 and HRD0 in the VPS as shown before, both could be included without including DPB1 and HRD1. Then the values for higher sublayers 15 for OLS1 could be derived as being equal to those signaled for TL0 or to the maximum allowed by the level. Thus, Fig. 15 illustrates an example of PTL, DPB and HRD definition and sharing among different OLS with sublayer information needed to be inferred when not present for some OLSs. 20 Therefore, in another embodiment there is no bitstream constraint on the values vps_ptl_max_temporal_id[ i ], vps_dpb_max_temporal_id[ i ], vps_hrd_max_tid[ i ], but for values of vps_ptl_max_temporal_id[ i ] greater than vps_dpb_max_temporal_id[ i ], vps_hrd_max_tid[ i ] DPB and HRD parameters for i > vps_dpb_max_temporal_id[ i ], vps_hrd_max_tid[ i ] up to vps_ptl_max_temporal_id[ i ] shall be inferred to be a maximum 25 value specified by the profile lever or equal to the highest signaled DPB and HRD parameters. Accordingly, according to another embodiment, encoder 10 is configured for forming the OLS indication and / or the video parameter set 81 such that the maximum temporal sublayer 30 indicated by the decoder requirement set 82 associated with the OLS is greater than or equal to the maximum temporal sublayer indicated by each of the buffer requirement set 84 and the bitstream conformance set 86 associated with the OLS. According to these embodiments, parameters missing within the buffer requirement set 84 and the bitstream conformance set 86 associated with the OLS, e.g. OLS 2 of Fig. 15, and relating to temporal 35 sublayers above the maximum temporal sublayer indicated by each of the buffer requirement set 84 and the bitstream conformance set 86, are to be set equal to the fourth 2026205066   29 Jun 2026 parameters or equal to parameters within the buffer requirement set 84 and the bitstream conformance set 86 associated with the OLS which relate to the maximum temporal sublayer indicated by each of the buffer requirement set 84 and the bitstream conformance set 86. 5 Accordingly, an embodiment of a decoder for decoding a multi-layered video bitstream, such as decoder 50 of Fig. 1 may be configured for inferring, if the maximum temporal sublayer indicated by the decoder requirement set 82 associated with the OLS is greater than or equal to the maximum temporal sublayer indicated by each of the buffer requirement 10 set 84 and the bitstream conformance set 86 associated with the OLS, i.e. the OLS to be decoded, inferring that parameters associated with the OLS for the buffer requirement set 84 and the bitstream conformance set 86 related to temporal sublayers above the maximum temporal sublayer indicated by each of the buffer requirement set and the bitstream conformance set are to be set equal to, for each of the parameters for the buffer requirement 15 set 84 and the bitstream conformance set 86 a default value such as a maximum value for the respective parameter indicated in the decoder requirement set 82, or a value for the respective parameter within the buffer requirement set 84 or the bitstream conformance set 86 associated with the OLS which relate to the maximum temporal sublayer indicated by each of the buffer requirement set 84 and the bitstream conformance set 86. For example, 20 the choice whether the default value is to be used or whether the value for the respective parameter within the buffer requirement set or the bitstream conformance set relating to the maximum temporal sublayer indicated by each of the buffer requirement set and the bitstream conformance set may be made differently for each of the parameters of the buffer requirement set 84 and the bitstream conformance set 86. 25 7. Output layer selection in region of interest applications Section 7 describes embodiments according to the seventh aspect making reference to Fig. 1. Thus, the description of Figs. 1 may optionally apply to the embodiments in accordance 30 with the seventh aspect. Also details described with respect to the further aspects may optionally be implemented in embodiments described in this section. Some embodiments according to the seventh aspect relate to PicOutputFlag derivation in RoI applciations. 35 2026205066   29 Jun 2026 When a multi layer bitstream, such as video bitstream 14, is used and pictures of the designated output layers are not available on decoder side (e.g. bitstream error or transmission loss), it may result in suboptimal user experience when certain considerations are not obeyed. Usually, when an access unit does not contain pictures in the output layer, 5 it is up to the implementation to select pictures from non-output layers for output as to compensate for the error / loss as evident from the following note below the derivation of the PicOutputFlag variable: - The variable PictureOutputFlag of the current picture is derived as follows: 10          - If sps_video_parameter_set_id is greater than 0 and the current layer is not an output layer (i.e., nuh_layer_id is not equal to OutputLayerIdInOls[ TargetOlsIdx ][ i ] for any value of i in the range of 0 to NumOutputLayersInOls[ TargetOlsIdx ] - 1, inclusive), or one of the following conditions is true, PictureOutputFlag is set equal to 0: - The current picture is a RASL picture and NoOutputBeforeRecoveryFlag of the 15               associated IRAP picture is equal to 1. - The current picture is a GDR picture with NoOutputBeforeRecoveryFlag equal to 1 or is a recovering picture of a GDR picture with NoOutputBeforeRecoveryFlag equal to 1. - Otherwise, PictureOutputFlag is set equal to ph_pic_output_flag. 20 NOTE - In an implementation, the decoder could output a picture not belonging to an output layer. For example, when there is only one output layer while in an AU the picture of the output layer is not available, e.g., due to a loss or layer down-switching, the decoder could set PictureOutputFlag set equal to 1 for the picture that has the highest value of nuh_layer_id among all pictures of the AU available to the decoder and having 25               ph_pic_output_flag equal to 1, and set PictureOutputFlag equal to 0 for all other pictures of the AU available to the decoder. It is, however, undesirable to change between layers in the decoder output on short time frames when the bitstream is made for a region of interest (RoI) application, i.e. higher 30 layers depict only a subset of the lower layer pictures (via the use of scaling window) as this would result in a very fast switching between overview and detail view. Therefore, as part of the invention, in one embodiment a decoder implementation is not permitted to freely select output layer when scaling windows are in use that do not cover the whole picture plane as follows: 2026205066   29 Jun 2026 In an implementation, the decoder could output a picture not belonging to an output layer as long as scaling windows cover the complete picture plane. For example, when there is only one output layer while in an AU the picture of the output layer is not available, e.g., due to a loss or layer down-switching, the decoder could set PictureOutputFlag set equal to 1 for the 5           picture that has the highest value of nuh_layer_id among all pictures of the AU available to the decoder and having ph_pic_output_flag equal to 1, and set PictureOutputFlag equal to 0 for all other pictures of the AU available to the decoder. According to an embodiment of the seventh aspect, the decoder 50 for decoding a multi-10 layered video bitstream, for example the multi-layered video bitstream 14 or the subbitstream 12, is configured for using vector-based interlayer prediction of predicted pictures 262 of a first layer 242 from reference pictures 261 of a second layer 241 with scaling and offsetting prediction vectors according to relative sizes and relative positions of scaling windows of the predicted pictures and the reference pictures which are defined in the multi-15 layered video bitstream 14. For example, picture 262 of layer 242 of Fig. 1 may be encoded into the multi-layered video data stream 14 using interlayer prediction from picture 261 of layer 241, e.g. the picture 261 of the same access unit 221. According to the seventh aspect, the multi-layered video bitstream 14 may comprise an OLS indication 18 of an OLS indicating a subset of layers of the multi-layered video bitstream 14, the OLS comprising 20 one or more output layers including the first layer 241 and one or more non-output layers including the second layer. In case of a loss of a predetermined picture of the first layer 242 of the OLS, such as picture 262, decoder 50 according to the seventh aspect is configured for substituting the 25 predetermined picture 262 by a further predetermined picture of the second layer 241 of the OLS which is in the same access unit 22 as the predetermined picture, in case of the scaling window defined for the predetermined picture 262 coinciding with the picture boundary of the predetermined picture and the scaling window defined for the further predetermined picture coinciding with the picture boundary of the further predetermined picture. In case of 30 at least one of the scaling pictures defined for the predetermined picture not coinciding with the picture boundary of the predetermined picture and the scaling window defined for the predetermined picture not coinciding with the picture boundary of the further predetermined picture, decoder 50 is configured for substituting the predetermined picture by other means or not at all. 35 8. Further embodiments 2026205066   29 Jun 2026 In the previous sections, although some aspects have been described as features in the context of an apparatus it is clear that such a description may also be regarded as a description of corresponding features of a method. Although some aspects have been described as features in the context of a method, it is clear that such a description may also 5 be regarded as a description of corresponding features concerning the functionality of an apparatus. Some or all of the method steps may be executed by (or using) a hardware apparatus, like for example, a microprocessor, a programmable computer or an electronic circuit. In some 10 embodiments, one or more of the most important method steps may be executed by such an apparatus. The inventive encoded image signal can be stored on a digital storage medium or can be transmitted on a transmission medium such as a wireless transmission medium or a wired 15 transmission medium such as the Internet. Depending on certain implementation requirements, embodiments of the invention can be implemented in hardware or in software or at least partially in hardware or at least partially in software. The implementation can be performed using a digital storage medium, for 20 example a floppy disk, a DVD, a Blu-Ray, a CD, a ROM, a PROM, an EPROM, an EEPROM or a FLASH memory, having electronically readable control signals stored thereon, which cooperate (or are capable of cooperating) with a programmable computer system such that the respective method is performed. Therefore, the digital storage medium may be computer readable. 25 Some embodiments according to the invention comprise a data carrier having electronically readable control signals, which are capable of cooperating with a programmable computer system, such that one of the methods described herein is performed. 30 Generally, embodiments of the present invention can be implemented as a computer program product with a program code, the program code being operative for performing one of the methods when the computer program product runs on a computer. The program code may for example be stored on a machine readable carrier. 35 Other embodiments comprise the computer program for performing one of the methods described herein, stored on a machine readable carrier. 2026205066   29 Jun 2026 In other words, an embodiment of the inventive method is, therefore, a computer program having a program code for performing one of the methods described herein, when the computer program runs on a computer. 5 A further embodiment of the inventive methods is, therefore, a data carrier (or a digital storage medium, or a computer-readable medium) comprising, recorded thereon, the computer program for performing one of the methods described herein. The data carrier, the digital storage medium or the recorded medium are typically tangible and / or non-transitory. 10 A further embodiment of the inventive method is, therefore, a data stream or a sequence of signals representing the computer program for performing one of the methods described herein. The data stream or the sequence of signals may for example be configured to be transferred via a data communication connection, for example via the Internet. 15 A further embodiment comprises a processing means, for example a computer, or a programmable logic device, configured to or adapted to perform one of the methods described herein. 20 A further embodiment comprises a computer having installed thereon the computer program for performing one of the methods described herein. A further embodiment according to the invention comprises an apparatus or a system configured to transfer (for example, electronically or optically) a computer program for 25 performing one of the methods described herein to a receiver. The receiver may, for example, be a computer, a mobile device, a memory device or the like. The apparatus or system may, for example, comprise a file server for transferring the computer program to the receiver. 30 In some embodiments, a programmable logic device (for example a field programmable gate array) may be used to perform some or all of the functionalities of the methods described herein. In some embodiments, a field programmable gate array may cooperate with a microprocessor in order to perform one of the methods described herein. Generally, the methods are preferably performed by any hardware apparatus. 2026205066   29 Jun 2026 The apparatus described herein may be implemented using a hardware apparatus, or using a computer, or using a combination of a hardware apparatus and a computer. The methods described herein may be performed using a hardware apparatus, or using a 5 computer, or using a combination of a hardware apparatus and a computer. In the foregoing Detailed Description, it can be seen that various features are grouped together in examples for the purpose of streamlining the disclosure. This method of disclosure is not to be interpreted as reflecting an intention that the claimed examples 10 require more features than are expressly recited in each claim. Rather, as the following claims reflect, subject matter may lie in less than all features of a single disclosed example. Thus the following claims are hereby incorporated into the Detailed Description, where each claim may stand on its own as a separate example. While each claim may stand on its own as a separate example, it is to be noted that, although a dependent claim may refer in the 15 claims to a specific combination with one or more other claims, other examples may also include a combination of the dependent claim with the subject matter of each other dependent claim or a combination of each feature with other dependent or independent claims. Such combinations are proposed herein unless it is stated that a specific combination is not intended. Furthermore, it is intended to include also features of a claim 20 to any other independent claim even if this claim is not directly made dependent to the independent claim. The above described embodiments are merely illustrative for the principles of the present disclosure. It is understood that modifications and variations of the arrangements and the 25 details described herein will be apparent to others skilled in the art. It is the intent, therefore, to be limited only by the scope of the pending patent claims and not by the specific details presented by way of description and explanation of the embodiments herein. It is to be understood that, if any prior art publication is referred to herein, such reference 30 does not constitute an admission that the publication forms a part of the common general knowledge in the art, in Australia or any other country. In the claims which follow and in the preceding description of the invention, except where the context requires otherwise due to express language or necessary implication, the word 35 “comprise” or variations such as “comprises” or “comprising” is used in an inclusive sense, 2026205066   29 Jun 2026 i.e. to specify the presence of the stated features but not to preclude the presence or addition of further features in various embodiments of the invention.

Claims

1. A video decoder comprising a processor configured to:obtain a bitstream that includes a plurality of layers and access units;5          decode an indication that the bitstream includes a plurality of output layer sets(OLSs), each OLS of the plurality of OLSs includes a respective number of the plurality of layers;determine that external information for identifying an OLS to be decoded among the plurality of OLSs is not available; and10          in response to the determination, for a first access unit of the access units, select,from the plurality of OLSs, an OLS with a highest number of the plurality of layers to be decoded.

2. The video decoder of Claim 1, wherein the processor is further configured to: 15         decode the selected OLS.

3. The video decoder of Claim 1, wherein to select the OLS, the processor is further configured to:select the OLS to be decoded from the plurality of OLSs with:20                  the highest number of the plurality of layers, anda lowest index value.

4. The video decoder of Claim 3, wherein to select the OLS to be decoded, the processor is further configured to:25          identify a set of the plurality of OLSs with the highest number of the plurality of layers;identify, from the set of the plurality of OLSs with the highest number of the pluralityof layers, an OLS that has the lowest index value; andselect the identified OLS that has the lowest index value as the OLS to be decoded.30          5. The video decoder of claim 1, wherein the external information is external tothe bitstream.2026205066   29 Jun 20266. A video decoding method comprising:obtaining a bitstream that includes a plurality of layers and access units;decoding an indication that the bitstream includes a plurality of output layer sets (OLSs), each OLS of the plurality of OLSs includes a respective number of the plurality of 5 layers;determining that external information for identifying an OLS to be decoded among the plurality of OLSs is not available; andin response to the determination, for a first access unit of the access units, selecting, from the plurality of OLSs, an OLS with a highest number of the plurality of layers to be 10 decoded.

7. The video decoding method of Claim 6, further comprising:decoding the selected OLS.15         8. The video decoding method of Claim 6, wherein selecting the OLS, themethod further comprises:selecting the OLS to be decoded from the plurality of OLSs with:the highest number of the plurality of layers, and a lowest index value.

209. The video decoding method of Claim 8, wherein selecting the OLS to be decoded further comprises:identifying a set of the plurality of OLSs with the highest number of the plurality of layers;25          identifying, from the set of the plurality of OLSs with the highest number of theplurality of layers, an OLS that has the lowest index value; andselecting the identified OLS that has the lowest index value as the OLS to be decoded.30          10. The video decoding method of claim 6, wherein the external information isexternal to the bitstream.2026205066   29 Jun 202611. A non-transitory computer readable storage medium containing instructions, that when executed by at least one processor of an electronic device, cause the at least one processor to:obtain a bitstream that includes a plurality of layers and access units;5          decode an indication that the bitstream includes a plurality of output layer sets(OLSs), each OLS of the plurality of OLSs includes a respective number of the plurality of layers;determine that external information for identifying an OLS to be decoded among the plurality of OLSs is not available; and10          in response to the determination, for a first access unit of the access units, select,from the plurality of OLSs, an OLS with a highest number of the plurality of layers to be decoded.

12. The non-transitory computer readable storage medium of Claim 11, further15 containing instructions that when executed cause the at least one processor to: decode the selected OLS.

13. The non-transitory computer readable storage medium of Claim 11, wherein the instructions that when executed cause the at least one processor to select the OLS, 20 comprise instructions that when executed cause the at least one processor to:select the OLS to be decoded from the plurality of OLSs with: the highest number of the plurality of layers, and a lowest index value.25         14. The non-transitory computer readable storage medium of Claim 13, whereinthe instructions that when executed cause the at least one processor to select the OLS to be decoded, comprise instructions that when executed cause the at least one processor to: identify a set of the plurality of OLSs with the highest number of the plurality of layers; identify, from the set of the plurality of OLSs with the highest number of the plurality30 of layers, an OLS that has the lowest index value; andselect the identified OLS that has the lowest index value as the OLS to be decoded.

15. The non-transitory computer readable storage medium of claim 11, wherein the external information is external to the bitstream.