Multi-layered video bitstream and wide-ranging signaling concept for leading output timing

By encoding RASL pictures in the next access unit and signaling output timing, the issue of incorrect decoding in multi-layer video bitstreams is resolved, ensuring accurate picture output and flexible frame rates.

JP2025186515APending Publication Date: 2025-12-23FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2025163602
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2020-06-10
Filing Date
2025-09-30
Publication Date
2025-12-23

AI Technical Summary

Technical Problem

In multi-layer video bitstreams, independently coded reference pictures (RASL pictures) may not be correctly decoded due to missing reference pictures, leading to incorrect decoding of dependent pictures.

Method used

Prevent pictures with inter-layer references to RASL pictures from being shown for output by encoding them in the next access unit following an end-of-sequence identifier, using decode refresh, and signaling picture output timing through supplemental extension information.

Benefits of technology

Ensures correct decoding of pictures by preventing inter-layer references to undecodable RASL pictures, maintaining consistent output timing, and allowing for different frame rates across layers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025186515000001_ABST
    Figure 2025186515000001_ABST
Patent Text Reader

Abstract

To provide a method for processing a coding layer video sequence boundary in a multi-layered video bitstream, a non-transitory computer readable medium, and a coding device.SOLUTION: An encoder 10 encodes an access unit 22 into a bit stream portion 16 of a video bitstream 14 that could be a single layer or multi-layered video bitstream including one or more layers. A bit stream portion 16 where a picture 26 is encoded is called a video coding layer (VCL) NAL unit. The video bitstream 14 further includes a non-VCL NAL unit where descriptive data is coded, and bit stream portions 23, 29. A video 12 is coded in a sequence of a coding video sequence 20.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] FIELD OF THE DISCLOSURE Embodiments of the present disclosure relate to a video encoder, a video decoder, a method for encoding a video sequence into a video bitstream, and a method for decoding a video sequence from a video bitstream. Further embodiments relate to video bitstreams.

[0002] Video may be coded into a video bitstream in units of one or more coded video sequences, each including a sequence of access units containing one or more pictures of a common time frame of the video. In the case of a multi-layer video bitstream, video data is encoded into multiple layers of the video bitstream, and a layer may include individual coded layer video sequences, with coded layer video sequences of different layers not necessarily starting / stopping at the same access unit. A coded layer video sequence may start with an independently coded picture, e.g., an IRAP picture, which may be independently decoded from a picture of an access unit different from that of a dependently coded picture. There may be an independently coded first type picture, e.g., a CRA picture, to which pictures of the same layer may be associated. This picture follows the first type picture in coding order but is scheduled for presentation before the first type picture. Such a picture may be referred to as a RASL picture. Such a RASL picture may have a reference to a picture preceding the first type picture with which the RASL picture is associated in decoding order. In other words, an RASL picture may have references to pictures in a previous coding layer video sequence. Therefore, the absence of a reference picture of an RASL picture in the video bitstream may prevent the RASL picture from being decoded correctly. In this case, there may be an indication that pictures preceding the first type of picture in presentation order should be excluded from the output. On the other hand, an RASL picture can serve as an inter-layer reference picture for a picture of another layer. If an RASL picture cannot be decoded, the pictures dependent on the RASL picture cannot be decoded correctly either. Summary of the Invention [Means for solving the problem]

[0003] A first aspect of the present disclosure provides a concept for handling coded layer video sequence boundaries in a multi-layered video bitstream. Embodiments according to the first aspect can prevent pictures having inter-layer references to RASL pictures from being shown for output, where the RASL pictures are not decodable or excluded from output. To this end, such pictures having inter-layer references to RASL pictures that cannot be correctly decodable are not present in the video bitstream or are excluded from output.

[0004] According to an embodiment of the first aspect, when encoding a first layer and a second layer into a multi-layer video bitstream such that the first layer depends on the second layer, the encoder encodes, using decode refresh, a picture to be encoded in the first layer in a next access unit among subsequent access units following in coding order an access unit including an end-of-sequence identifier for the second layer, the next access unit being the closest access unit to the access unit including the end-of-sequence identifier for the second layer in which the picture is encoded in the first layer, without outputting the first picture. Because the first picture of the next access unit is encoded using decode refresh and the first picture is not output, output can be prevented, for example, by indicating that a picture having an inter-layer reference to a RASL picture in the second layer must also be a RASL picture, and therefore the first picture should not be output. By indicating that the first picture will not be output, it is possible to prevent pictures that are indirectly dependent by inter-layer reference on a picture that is part of a coding layer video sequence of the second layer that is indicated to end with the sequence end identifier of the second layer from being output.

[0005] A second aspect of the present disclosure relates to output timing of decoded pictures, i.e., the output time at which a decoded picture is output from an output buffer of a decoder. The derivation of picture output timing may be signaled at the access unit level, for example, by picture timing supplemental extension information (PT SEI). Additionally or alternatively, picture output timing may be signaled at the output layer level, i.e., with reference to individual output layers of a video bitstream. For example, information regarding picture output timing of an access unit and / or output layer may include information regarding the number of times a picture is output, i.e., repeated.

[0006] According to a first sub-aspect of the second aspect, a gating flag is provided in a video bitstream to signal whether a PT SEI message included in the video bitstream includes a picture output multiplication syntax element. The picture output multiplication syntax element indicates whether a picture of the access unit it references is subject to multiplied picture output, and if so, how many output pictures are generated from the picture of the access unit. The gating flag provides a means for distinguishing whether information about the multiplied picture output is obtained from a PT SEI message or from other means, such as a frame field SEI message referencing individual output layers. Thus, signaling the gating flag enables signaling of different frame rates for different output layers of an output layer set. In other words, the gating flag enables signaling of individually multiplied picture outputs for different pictures within one access unit.

[0007] A second sub-aspect of the second aspect provides a concept for using multiplied picture output of pictures of an access unit in conjunction with a frame-field syntax element that indicates where a picture of a picture sequence represents a field or a frame, e.g., an interlaced or progressive picture. Thus, an embodiment of the second sub-aspect enables signaling of picture output times when frames or fields are used.

[0008] A third sub-aspect of the second aspect provides a concept for signaling the number of picture outputs by a picture output multiplication syntax element of a PT SEI that references an access unit and a further picture output multiplication of a video bitstream, and a further picture output multiplication syntax element of a frame field SEI that references an output layer of the video bitstream. According to an embodiment, the picture output multiplication syntax element is equal to or less than the further picture output multiplication syntax element, e.g., the further picture output multiplication syntax element is an integer multiple of the picture output multiplication syntax element. The picture output multiplication syntax element included in the PT SEI message may be access unit-specific and may therefore enable determining a picture refresh interval for the output achieved as a result of the repetition or multiplication indicated by the picture output multiplication syntax element. Thus, the picture output multiplication syntax element allows determining timing information of the picture output associated with the access unit and, in the case of multiplication, the interval between the pictures presented therebetween. The further picture output multiplication syntax element signaled in the frame field SEI message can provide layer-specific information regarding how often a picture needs to be repeated to enable content to be presented at the picture refresh interval, i.e., the interval determined from the picture output multiplication syntax element. Requiring the further picture output multiplication syntax element to be, for example, an integer multiple or greater than the picture output multiplication syntax element can ensure that the picture refresh interval signaled by the picture output multiplication syntax element is achieved by the numeric value of the picture output time signaled by the further picture output multiplication syntax element. For example, for the first and second fields, the further picture output multiplication syntax element can signal a multiplication value corresponding to twice the multiplication value signaled by the picture output multiplication syntax element.

[0009] A fourth sub-aspect of the second aspect provides a concept for deriving from a video bitstream whether the output frame rate is constant across boundaries between subsequent coded video sequences, for example, without explicitly signaling this information in the video bitstream. By inferring this information instead of explicitly signaling it, there is an advantage that when splicing video bitstreams, it is not necessary to correct each piece of information or check whether it is still constant.

[0010] A fifth sub-aspect of the second aspect provides a concept for deriving an element picture output time (e.g., an output time of an access unit) of a coded video sequence based on an element output picture duration syntax element, e.g., elemental_duration_in_tc_minus1, which may be part of a parameter set, e.g., a video parameter set or a sequence parameter set having HRD and timing information. This concept relies on the idea that if one or more syntax elements encoded in a video bitstream have an initial state, it is possible to infer that the access unit is not subject to multiply-output. Thus, if information about whether pictures of an access unit are subject to multiply-output can be inferred, this concept may enable determining the element picture output time without requiring a PT SEI message signaling this information. For example, in this case, the element output picture time may be derived based on information about the output time of each picture, as may be provided by, e.g., an element output picture duration syntax element. Therefore, this concept may allow for deriving element picture output times in the absence of a PT SEI message and / or may allow for omitting signaling of a PT SEI message.

[0011] A sixth sub-aspect of the second aspect provides a concept for processing when there are no output pictures in a video bitstream for which a constant picture rate is signaled. According to the sixth sub-aspect, when a constant picture rate is indicated for a video bitstream, pictures preceding pictures that will not be output, i.e., pictures preceding pictures that are indicated to be omitted from output, are repeated. Thus, a constant picture rate can be maintained even when there are no output pictures. [Brief explanation of the drawings]

[0012] Embodiments and advantageous implementations of the present disclosure are explained in more detail below with reference to the figures. [Figure 1] 1 illustrates an encoder, a decoder, and a video bitstream according to an embodiment. [Figure 2] An example of two layers of a bitstream is shown, with the two layers having different IRAP periods. [Figure 3] An example of random access to a two-layer video bitstream with no end-of-sequence indication is given. [Figure 4] 3 shows an example of an encoded video sequence with aligned end-of-sequence indications according to an embodiment of the first aspect; [Figure 5] 3 shows an example of a dependency layer according to an embodiment of the first aspect; [Figure 6] An example of a temporal sublayer is shown below. [Figure 7] An example of bitstream splicing is shown below. [Figure 8] An example of frame repetition is shown below. [Figure 9] Here is an example of a bitstream with two layers with different frame rates: [Figure 10] Here is an example of a two layer bitstream with one layer repeating the output frame: [Figure 11] 10 shows an encoder and a video bitstream according to an embodiment of sub-aspects 2 and 3. [Figure 12] Examples regarding GOP size, DPB parameters and reordering are given below. [Figure 13] 10 shows examples of an encoder, a decoder and a video bitstream according to an embodiment of sub-aspects 2 and 5. [Figure 14] An example of a bitstream containing pictures that are not output is shown below. DETAILED DESCRIPTION OF THE INVENTION

[0013] Although the following describes embodiments in detail, it should be understood that the embodiments provide many applicable concepts that can be embodied in a wide variety of video coding concepts. The specific embodiments described are merely illustrative of specific ways to implement and use the concepts and do not limit the scope of the embodiments. In the following description, numerous details are set forth to provide a more thorough description of embodiments of the present invention. However, it will be apparent to one skilled in the art that other embodiments can be practiced without these specific details. In other instances, well-known structures and devices are shown in block diagram form, rather than in detail, in order to avoid obscuring the examples described herein. Furthermore, features of different embodiments described herein can be combined with each other, unless otherwise noted.

[0014] In the following description of the embodiments, identical or similar elements or elements having the same functions are given the same reference numerals or identified by the same names, and repeated descriptions of elements given the same reference numerals or identified by the same names are generally omitted. Therefore, the descriptions provided for elements having the same or similar reference numerals or identified by the same names can be mutually interchangeable or applied to each other in different embodiments.

[0015] A detailed description of embodiments of the disclosed concepts begins with a description of example encoders, decoders, and video bitstreams, which provide a framework into which embodiments of the present invention can be incorporated. Below, a description of embodiments of the present concepts is presented along with an explanation of how such concepts can be incorporated into the encoder and decoder of FIG. 1. However, the embodiments described with respect to FIG. 2 and subsequent figures may be used to form encoders and decoders that do not operate according to the framework described with respect to FIG. 1. Furthermore, it should be noted that the encoder and decoder, although described together for illustrative purposes in FIG. 1, may be implemented separately from one another. It should also be noted that the encoder and decoder may be combined within a single device, or one of the two may be implemented as part of the other. Additionally, some embodiments of the present invention will be described with reference to FIG. 1.

[0016] 0. Encoder 10, decoder 50 and video bitstream 14 according to FIG. FIG. 1 shows examples of an encoder 10 and a decoder 50. The encoder 10 (which may also be referred to as an apparatus for encoding) encodes a video sequence 12 into a video bitstream 14 (which may also be referred to as a bitstream, a data stream, a video data stream, or a stream). The video sequence 12 includes a sequence of pictures 13, which are arranged in a presentation order or picture order 17. In other words, each of the pictures 13 may represent a frame of the video sequence 12 and may be associated with a time instant in the presentation order of the video sequence 12. Based on the video sequence 12, the encoder 10 may encode a coded video sequence 20 into the video bitstream 14. The encoder 10 may form the coded video sequence 20 in the form of access units 22, each of which encodes video data belonging to a common time instant. In other words, each of the access units 22 may encode one of the frames of the video sequence 12. Encoder 10 encodes coded video sequence 20 into video bitstream 14 according to a coding order 19 that may differ from picture order 17 of video sequence 12 .

[0017] Encoder 10 can encode coded video sequence 20 into one or more layers. That is, video bitstream 14 can be a single-layer or multi-layer video bitstream containing one or more layers. Each of access units 22 contains one or more coded pictures 26 (e.g., pictures 260, 261 in FIG. 1 , where apostrophes and stars are used to refer to specific ones and subscript indexes indicate the layer to which the picture belongs). Note that hereinafter, coded pictures may be simply referred to as pictures. Each of pictures 26 belongs to one of layers 24 of the coded video sequence, e.g., layers 240, 241 in FIG. 1 . An exemplary number of two layers, namely, first layer 241 and second layer 240, are shown in FIG. 1 . In embodiments consistent with the disclosed concepts, coded video sequence 20 and video bitstream 14 do not necessarily contain multiple layers, but may contain one, two, or more layers. 1, each of the access units 22 includes a coded picture 261 of a first layer 241 and a coded picture 260 of a second layer 240. However, it should be noted that each of the access units 22 may, but need not, include coded pictures for each of the layers of the coded video sequence 20. For example, the layers 240, 241 may have different frame rates (or picture rates) and / or may include pictures for complementary subsets of the access units of the access units 22.

[0018] As mentioned above, pictures 260, 261 of one access unit represent picture content at the same time. For example, pictures 260, 261 of the same access unit 22 can represent the same picture content at different qualities, e.g., resolution or fidelity. In other words, layer 240 can represent a first version of coded video sequence 20, and layer 241 can represent a second version of coded video sequence 20. Thus, a decoder, such as decoder 50, or extractor, can select between different versions of coded video sequence 20 to decode or extract from video bitstream 14. For example, layer 240 can be decoded independently from further layers of the coded video sequence to provide a decoded video sequence of a first quality, while joint decoding of first layer 241 and second layer 240 can provide a decoded video sequence of a second quality that is higher than the first quality. For example, first layer 241 can be encoded based on second layer 240. In other words, the second layer 240 may be a reference layer for the first layer 241. For example, in this scenario, the first layer 241 may be referred to as an enhancement layer, and the second layer 240 may be referred to as a base layer. The picture 260 may have a smaller, equal, or larger picture size than the picture 261. For example, the picture size may refer to the number of samples in a two-dimensional array of the picture. Note that the pictures 260 and 261 do not necessarily represent the same picture content, but, for example, the picture 261 may represent an excerpt of the picture content of the picture 260. For example, in some scenarios, different layers of the video bitstream 14 may include different sub-pictures of a picture encoded in the video bitstream, which may be encoded independently of each other. Thus, in a further example, the layers 240 and 241 may be encoded independently of each other in the video bitstream 14.

[0019] Encoder 10 encodes access units 22 into bitstream portions 16 of video bitstream 14. For example, each of access units 22 may be encoded into one or more bitstream portions 16. For example, pictures 26 may be subdivided into tiles of slices, and each of the slices may be encoded into one bitstream portion 16. The bitstream portions 16 into which pictures 26 are encoded are sometimes referred to as video coding layer (VCL) NAL units. Video bitstream 14 may further include non-VCL NAL units, e.g., bitstream portions 23, 29, into which description data is encoded. The description data may provide information for decoding or information about coded video sequence 20. The bitstream portions into which description data is encoded may be associated with individual bitstream portions; for example, they may reference individual slices, or may be associated with one of pictures 26 or one of access units 22, or may be associated with a sequence of access units, i.e., related to coded video sequence 20. Note that video 12 may be encoded into a sequence of coded video sequence 20.

[0020] Decoder 50 (which may also be referred to as an apparatus for decoding) decodes video bitstream 14 to obtain decoded video sequence 51. It should be noted that video bitstream 14 provided to decoder 50 does not necessarily correspond to video bitstream 14 provided by an encoder, but may be derived from the video bitstream provided by the encoder, such that the video bitstream decoded by decoder 50 may be a sub-bitstream of a video bitstream encoded by an encoder, such as encoder 10. As mentioned above, decoder 50 may decode the entire coded video sequence 20 encoded in video data stream 14, or may decode a portion thereof, e.g., a subset of layers of coded video sequence 20 (i.e., a video sequence having a frame rate lower than the maximum frame rate provided by coded video sequence 20) and / or a temporal subset of coded video sequence 20. Thus, decoded video sequence 51 does not necessarily correspond to video sequence 12 encoded by encoder 10. It should also be noted that decoded video sequence 51 may further differ from video sequence 12 due to coding losses, such as quantization losses. Decoded video sequence 51 includes, for each frame of the decoded video sequence, one or more decoded pictures 53 that have been decoded from respective layers of coded video sequence 20. In other words, in the example, decoded video sequence 51 may include one or more layers, similar to coded video sequence 20. Decoded pictures 53 may be output according to an output order 18, which in the example may correspond to picture order 17. However, decoded video sequence 51 does not necessarily include all frames of video sequence 12, and may also include multiple instances, i.e., repetitions, of a single picture, as detailed in Section 2.

[0021] Picture 26 may be encoded using a predictive tool to predict signals or coefficients representing a picture in video bitstream 14 from previously encoded pictures. * For example, a prediction tool can be used to encode a currently to be encoded picture using previously encoded pictures. Correspondingly, decoder 50 can derive a currently to be decoded picture 26 from a previously decoded picture. * In the following description, a given picture or block, e.g., the picture or block currently being coded, is referred to as ( * ) are referenced using the image. For example, picture 261 in Figure 1 * is considered to be the currently coded picture, and the currently coded picture 26 * may equally refer to the currently encoded picture being encoded by encoder 10 and the currently decoded picture in the decoding process performed by decoder 50.

[0022] Predicting a picture from another picture in coded video sequence 20 is sometimes called inter-prediction. * Picture 261 * , may be encoded using temporal inter prediction from a picture 261′ belonging to one of the access units 22 different from * belongs to the same layer but is part of picture 261 * Additionally or alternatively, the picture 261 may include an intra-layer reference 32 to a picture 261' that belongs to a different access unit than the picture 261. * 261 may optionally be predicted using inter-layer (inter) prediction from a picture of another layer, e.g., a lower layer (the lower layer depending on the layer index that may be associated with each of the layers 24). *may contain inter-layer references 34 to pictures 260' that belong to the same access unit but to different layers. In other words, in FIG. 1, pictures 261', 260' are currently coded pictures 261 * , ...

[0023] The embodiments described herein may be implemented in the context of Versatile Video Coding (VVC) or other video codecs.

[0024] In the following, some concepts and embodiments will be described with reference to Fig. 1, and features will be described with respect to Fig. 1. It is pointed out that a feature described with respect to an encoder, a video bitstream, or a decoder shall be understood to be a description with respect to other of these entities as well. For example, a feature described as being present in a video data stream shall be understood as a description of an encoder configured to encode this feature into a video bitstream, and a decoder or extractor configured to read the feature from the video bitstream. It is further pointed out that the estimation of information based on the representation encoded in the video bitstream can be performed equally on the encoder side and on the decoder side. It is further noted that the aspects described in the following sections can be combined with each other.

[0025] 1. Impact of End of Sequence (EOS) on Multi-layer Bitstreams In this section, an embodiment according to the first aspect is described with reference to Figure 1. The details described in Section 0 may optionally be applied to the embodiment according to the first aspect.

[0026] According to an embodiment of the first aspect, the video bitstream 40 is a multi-layered video bitstream, for example, as shown in FIG. 1. As described in Section 0, a picture 26 may be coded with inter-layer prediction, such that decoding requires information from another picture in the same layer. Alternatively or additionally, a picture may be coded using inter-layer prediction, such that decoding requires information from a picture in another layer. In contrast, an independently coded picture, or a randomly accessible picture, may be a picture that is not dependent on a picture that belongs to an access unit 22 different from its own access unit. In other words, an independently coded picture is encoded without using temporal inter-prediction. For example, an intra-random access point (IRAP) picture may be an independently coded picture. Examples of IRAP pictures include instantaneous decode refresh (IDR) pictures and clean random access (CRA) pictures. As noted in Section 0, the coding order 19 does not necessarily correspond to the picture order and presentation order (also called output order). A dependently coded picture that precedes a previous picture in coding order and presentation order may be called a trailing picture. Another example of a dependently coded picture is a picture that depends on a picture of a previously coded access unit, but the dependently coded picture precedes the dependent picture in presentation order 19. An example of such a picture may be a random access skip ahead (RASL) picture. An RASL picture may be associated with and dependent on an independently coded picture, such as a CRA picture, that precedes the RASL picture in coding order but follows the RASL picture in presentation order. Furthermore, an RASL picture may depend on (i.e., contain references to) one or more further pictures, including one or more pictures that precede the associated independently coded (e.g., CRA) picture in coding order.If an independently coded picture is the start picture of a coded video (layer) sequence (e.g., because it is the first picture in the bitstream or the first picture after an end-of-sequence indication), this may mean that pictures preceding the independently coded picture in coding order are cleared from the buffer, and therefore RASL pictures may be omitted from the output since they may be incorrectly decoded due to having no references.

[0027] The coded video sequence 21 may include one or more coded layer video sequences in each of the layers 24. A coded layer video sequence may begin with a coded layer video sequence start picture, e.g., an independently coded picture, and may include all pictures of the respective layer from the coded layer video sequence start picture to the next coded picture in coding order 19, exclusively, or to the end of the coded video layer sequence. Note that each of the layers 24 may have a different number and / or a different arrangement of coded layer video sequences. In other words, the coded layer video sequence start pictures of different layers are not necessarily aligned within the same access unit.

[0028] FIG. 2 shows an example of two layers 240, 241 with different durations of IRAP pictures.

[0029] If a bitstream includes multiple layers, the IRAP pictures of each layer do not need to be aligned; for example, a lower layer L0, e.g., layer 240 in Figure 1, may have more frequent IRAP pictures than a higher, dependent layer L1, e.g., layer 241 in Figure 1. The CLVS 21' of the lower layer stops at each such IRAP picture, but the CLVS of the higher layer can continue, for example, by having an IDR AU in the lower layer but not in the higher layer, or by an end-of-sequence (EOS) NAL unit as shown in Figure 2.

[0030] The bitstream contains a CLVSS picture 260 (CRA in Figure 2). * There are cases where it is necessary to include a so-called end-of-sequence (EOS) NAL unit 41 before the start of a new CLVS 21' that stops the first CLVS 21' (in the example). Subsequently, if a CLVSS picture (CRA) has NoOutputBeforeRecoveryFlag equal to 1, the CLVS 21' will be the first CLVS picture 260'. * 2, the L1 Trail, e.g., picture 261′ in FIG. 2, will also be incorrectly reconstructed and subsequently output.

[0031] According to a first embodiment of the first aspect, the encoder 10 is configured to encode a non-RASL picture, e.g., picture 261′ in the picture 261 of the first layer 241 of FIG. 2, which is then encoded into a RASL picture of the second layer 240, e.g., picture 260′ in FIG. 2, which is then encoded into a RASL picture of the second layer 240, e.g., picture 260′ in FIG. ’2. Furthermore, the encoder 10 encodes the RASL picture of the first layer 241 using the RASL picture of the second layer 240 as an inter-layer prediction reference for the RASL picture of the first layer 241, such that the RASL picture is temporally aligned to the RASL picture of the second layer 240. For example, referring to FIG. 2, assuming that picture 261′ is an RASL picture, the encoder 10 encodes picture 261′, e.g., referred to as the first RASL picture, using RASL picture 260′ of the second layer 240 as an inter-layer prediction reference. RASL picture 260′ of the second layer 240 can be referred to as the second RASL picture. In the case shown in FIG. 2, the picture 261' is a non-RASL picture, and the encoder 10 according to the first embodiment encodes the first RASL picture 261' without prediction from the third RASL picture 260'.

[0032] Using an inter-layer prediction reference of a picture may mean considering the same picture when employing previously encoded pictures to form a reference picture list for inter-predicting the currently encoded picture with respect to vector-based inter-prediction and / or motion vector prediction. Optionally, using a picture as an inter-layer prediction reference may further mean considering the same picture as referenced in a reference picture list via a reference index for an inter-predicted block of the currently encoded picture.

[0033] Encoding a picture without predicting from a particular picture may mean, with respect to vector-based inter prediction and / or motion vector prediction, avoiding a particular picture when employing previously encoded pictures to form a reference picture list for inter predicting the currently encoded picture, and / or avoiding a particular picture that is referenced in the reference picture list of the currently encoded picture via the reference index of an inter-predicted block of the currently encoded picture.

[0034] That is, according to the first embodiment of the first aspect, if the reference picture is not an RASL picture, the RASL picture is prohibited as an inter-layer reference picture, so that the L1 trailing picture in the same access unit as the RASL picture can be correctly reconstructed. The constraint described here means that the RASL picture is not used as a reference by not being present in the RPL (Reference Picture List) or not being selected from the RPL. That is, they may be inactive or unused references in the RPL.

[0035] However, this constraint is unnecessarily strict, and there are cases where this is not an issue, for example, when there is no EOS NUT, as shown in Figure 3.

[0036] FIG. 3 shows an example of a random access picture without a preceding EOS indication.

[0037] In such cases, the above reference from L1 Trail to L0 RASL is CRA position 22 * This only matters when tuning (random access) in , but the decoder will anyway skip all pictures in enhancement layer L1 until it encounters an IRAP picture in the respective enhancement layer.

[0038] In the example of FIG. 3, the picture 260 in the second layer 240 *is a CRA picture, e.g., an independently coded picture to which one or more RASL pictures may depend. In the example of Figure 3, in contrast to Figure 2, CRA picture 260 is coded in coding order 19. * The second layer picture 240 preceding the CRA picture 260 does not have an end of sequence indication 41. * may be the starting picture of a coding layer video sequence, e.g., a CRA picture 260, such as a RASL picture 260' in this case. * For example, CRA picture 260′ may be a coding layer video sequence start picture if pre-recovery no-output slack, described below, is set to 1 and HandleCraAsClvsStartFlag, described below, is set to 1.

[0039] According to a second embodiment of the first aspect, the encoder 10 may use the RASL picture of the second layer 240 if the picture of the first layer is an RASL picture and the CRA picture 260' with which the RASL picture 260' of the second layer 240 is associated is the start picture of the coded layer video sequence, and proposes the RASL picture 260' as an inter-layer predicted reference picture for a picture of the first layer 241 that is temporally aligned with the RASL picture of the second layer 240, e.g., part of the same access unit. In other words, in the example of FIG. 3 , the picture 261' is temporally aligned with the picture 260', i.e., the two pictures are in the same access unit 22'. The picture 260' of the second layer 240 is temporally aligned with the picture 260' of the CRA picture 260'. * According to the second embodiment, if the picture 261' is a RASL picture, the encoder 10 uses the picture 260' as an inter-layer reference picture for the picture 261'. If the picture 261' is a non-RASL picture, the encoder 10 uses the picture 260' as an inter-layer reference picture for the picture 261'. *If picture 260' does not form the start of coded layer video sequence 21'', then picture 260' is used as an inter-layer reference picture for picture 261'. Otherwise, i.e., picture 261' is a non-RASL picture and CRA picture 260 * form the start of coded layer video sequence 21'', encoder 10 encodes picture 261' as an inter-layer reference picture and does not encode picture 260'. * forms the start of a coding layer video sequence 21'', the RASL picture 260' may not be decoded correctly because the previous picture in coding order 19 is not available.

[0040] In other words, according to an example of the second embodiment of the first aspect, the above first embodiment (wherein if the reference picture is not a RASL picture, the RASL picture is prohibited as an inter-layer reference picture) is subject to the following conditions: RASL picture associated with the CRA following the EOS NAL unit (see Figure 2) · RASL pictures associated with a CRA that has NoOutputBeforeRecoveryFlag set to 1 by external means of setting HandleCraAsClvsStartFlag to 1 (see Figure 3).

[0041] The latter case is when the decoder is informed by the API that any CRA is to be treated as the start of a new CLVS, and therefore the processing is the same as when there is an EOS NAL unit where no such NAL unit exists.

[0042] That is, the constraint (do not allow RASL as ILRP reference picture) is conditioned on the relevant CRA setting NoOutputBeforeRecoveryFlag equal to 1 (either by way of an EOS NAL or the presence of external means):

[0043] The following constraints apply to the pictures referenced by each ILRP entry, if any, in RefPicList[0] or RefPicList[1] of a slice of the current picture: The picture is assumed to be in the same AU as the current picture. *The picture must exist in the DPB. o The picture shall have a nuh_layer_id ref PicLayerId that is smaller than the nuh_layer_id of the current picture. If the associated CRA's NoOutputBeforeRecoveryFlag is set to 1 and the current picture is not a RASL, the picture must not be a RASL picture. One of the following constraints applies: The picture is assumed to be an IRAP picture. The TemporalId of a picture shall be less than or equal to Max(0,vps_max_tid_il_ref_pics_plus1[currLayerIdx][refLayerIdx]-1), where currLayerIdx and refLayerIdx are equal to GeneralLayerIdx[nuh_layer_id] and GeneralLayerIdx[refpicLayerId], respectively.

[0044] 4 shows an example of how coded video sequence 20 may be encoded into video data stream 14. Coded video sequence 20 includes first layer 240 and second layer 241, where the second layer is a reference layer for the first layer. In access unit 22′, second layer 240 has sequence end indicator 41 indicating that access unit 22′ is the last access unit in coding order 19 of coded layer video sequence 210′ ​​of second layer 240 and that access unit 22″ following access unit 22′ in coding order 19 indicates the start of a new coded layer video sequence 210″ of second layer 240.

[0045] According to a third embodiment of the first aspect, the encoder 10 is configured to insert in each such access unit 22′ having a sequence end indication 41 (also called a sequence end identifier) ​​into the first layer 241 and the sequence end indication 41, as shown in FIG.

[0046] As a result, because the first layer 241 has an end-of-sequence indication 41 at access unit 22′, the end of the coding layer video sequence 211′ of the first layer 24′ ends at access unit 22′, and therefore ends at the same access unit as the coding layer video sequence 210′ ​​of the second layer 240. Thus, in layer 241, a new coding layer video sequence 211″ starts at access unit 22″, synchronized with the start of the next coding layer video sequence 210″ in the second layer 240. Due to the end-of-sequence indication 41 in the first layer 241, a RASL picture should not occur in the next coding layer video sequence 211″, avoiding the above-mentioned problems with asynchronous coding layer video sequence boundaries.

[0047] An alternative to the third embodiment is described with reference to FIG. 5, which shows a video bitstream according to the scenario described with reference to FIG. 4. However, according to this embodiment, the encoder 10 does not necessarily insert an end-of-sequence indication 41 in the first layer 241 in the access unit 22′ (although it may do so in the example). According to the alternative third embodiment, the encoder 10 is configured such that in the next access unit 22″, 22″ following the access unit 22′ having the end-of-sequence indication 41 in the second layer 240, a picture 261″ is encoded in that access unit 22″, 22″, and to encode that picture 261″ to be encoded in the first layer 241 using decoder refresh and without outputting a preceding picture. In other words, the next access unit 22″, 22′″ is the access unit following the access unit 22′ in coding order 19 in which a picture of the first layer is encoded, and of these access units, the next access unit 22″, 22′″ is the one that is closest to the access unit 22′. FIG. 5 shows two examples of first layers, referred to using reference numerals 241 and layer 242. Note that these two examples of first layers are shown together in FIG. 5 for illustrative purposes, but may represent independent examples. Thus, in an example, one or more such first layers may be present in coded video sequence 20. In layer 242, the next access unit is access unit 22''', in which picture 262''' is coded (because access unit 22'' has no pictures in the first layer). In layer 241, the next access unit is 22'', in which picture 261'' is coded.

[0048] Encoding pictures 261'', 261''' of the first layer 241, 242 that are coded dependently on the second layer 240 using decode refresh without outputting the preceding picture results in pictures of the first layer 241, 242 that are dependent on pictures of the second layer 240 that are RASL pictures, as may be the case for example for picture 260'', being not presented because they may also be RASL pictures.

[0049] Encoding pictures 261", 261'" using decode refresh may mean that pictures 261", 261'" are encoded without reference to pictures of access units other than the access unit to which pictures 261", 261'" belong, i.e., access units 22", 22'", respectively; for example, pictures 261", 261'" may be IDR or CRA pictures. The term "leading picture" of pictures 261", 261'" may mean a picture that follows pictures 261", 261'" in coding order 19 but precedes pictures 261", 262'" in presentation order 18, and that is dependent on (i.e., includes a reference to) the picture preceding pictures 261", 262'" in coding order 19. An example of a leading picture may be a RASL picture. Thus, encoding pictures 261'', 261''' without outputting leading pictures may mean, for example, that there are no pictures in the first layer 241, 242 that follow pictures 261'', 262''' in coding order 19 but precede pictures 261'', 262''' in presentation order (e.g., there are no RASL pictures), or that no pictures are indicated for output. In other words, according to the alternative third example embodiment, the encoder 10 is configured to encode the first layer such that picture 261'', encoded in the first layer 241 in the next access unit 22'', is a coded layer video sequence start picture.

[0050] For example, encoding pictures 261'', 261''' as IDR may imply that there are no leading pictures, e.g., RASL pictures, in the first layers 241, 242 that follow pictures 261'', 261''' in coding order. In other words, for example, picture 261'' being an IDR picture may prohibit picture 261'' from being a leading picture and therefore may also prohibit picture 261'''' from having inter-layer references to RASL pictures, such as the reference between pictures 261' and 260' in FIG. 3. Therefore, one example of encoding pictures 261'', 261''' using decode refresh and without outputting leading pictures is to encode pictures 261'', 261''' as IDR.

[0051] Alternatively, pictures 261", 261'" can be encoded as CRAs, and an end-of-sequence indication 41 may be used in access unit 22' of first layer 241, 242 to prevent output of leading pictures. In this case, leading pictures following pictures 261", 261'" in coding order may be present in first layer 241, 242, but may be excluded from output to prevent erroneous decoding. Thus, even if picture 261'" in layer 241 depends on picture 260'" in layer 240 and picture 260'" is a RASL picture, picture 260'" is not output.

[0052] Compared to the constraints of the first alternative of the third embodiment described with respect to FIG. 4, the second alternative with respect to FIG. 5 has the advantage that the encoder 10 does not necessarily need to insert an end-of-sequence indication 41 into the first layer 241.

[0053] In other words, according to the example of the third embodiment, the bitstream constraint is that if layer k is dependent on layer l and layer l contains an EOS NAL unit, then layer k must also contain an EOS NAL at the same position, or the next AU must contain a CLVSS picture (e.g., an IDR) of layer k.

[0054] 2. Picture output timing This section describes an embodiment according to a second aspect, which includes sub-aspects 1 to 6. The embodiment according to the second aspect is described with reference to FIG. 1 , and details and features thereof may optionally be included in the embodiment according to the second aspect. Also, details described in Section 1, such as details regarding the coding scheme of pictures, dependency relationships between pictures, coded video layer sequences, etc., may optionally be applied to the embodiment according to the second aspect.

[0055] As described with respect to FIG. 1 , decoder 50 decodes video bitstream 14, or portions thereof, to provide decoded video sequence 51. Decoder 50 may provide decoded pictures 53 of decoded video sequence 51 to a decoded picture buffer (DPB). Some embodiments according to a second aspect relate to the output timing of decoded pictures 53 from the decoded picture buffer, from which decoded pictures 53 may be provided, for example, to a display for presentation. Decoder 50 may provide decoded pictures 53 to the decoded picture buffer according to presentation order 18. As described with respect to FIG. 1 , decoder 50 need not necessarily decode and / or output all pictures of video bitstream 14, but may decode and / or output a subset of pictures encoded in video bitstream 14, where this subset of pictures may be defined by the layer with which the picture is associated and / or by a definition of a temporal subset of pictures. The temporal subsets may be defined by temporal layers as described with respect to FIG.

[0056] Figure 6 illustrates example temporal layers of coded video sequence 20. Figure 6 illustrates layer 240, which includes pictures 26 of a first temporal sublayer 250. A further layer 241 includes pictures of a second temporal layer 251. As illustrated in Figure 6, the first temporal sublayer 250 and the second temporal sublayer 251 may have equal frame rates, but the pictures of the first temporal sublayer 250 may belong to different access units 22 than the pictures of the second temporal sublayer. That is, the pictures of the first temporal sublayer 250 may belong to different times than the pictures of the second temporal sublayer 251. Note that in Figure 6, pictures 26 are arranged according to presentation order 18, not coding order 19. Figure 6 illustrates a further example layer, namely layer 242, which includes pictures of the first temporal sublayer 250 and the second temporal sublayer 251, respectively. Thus, layer 242 has a higher frame rate than layers 240 and 241. 6 is an illustrative example, and a coded video sequence may include any combination of layers, each of which may include one or more temporal sub-layers. As a result, each of access units 22 may include one or more pictures of layers 24.

[0057] The temporal sublayers may have a hierarchical order that may be defined, for example, by an index associated with the temporal sublayer. For example, the second temporal sublayer 251 may be higher in the hierarchical order than the first temporal sublayer 250.

[0058] Decoder 50 may select one or more of the temporal sub-layers included in video bitstream 14 for decoding, for example, by selecting the largest temporal sub-layer for decoding, i.e., the decoder may decode all temporal sub-layers down in hierarchical order up to the largest temporal sub-layer.

[0059] For example, decoder 50 may receive instructions indicating up to which temporal sub-layer video bitstream 14 should be decoded. In other examples, decoder 50 may independently determine the maximum temporal sub-layer to decode. In other words, the temporal subset of pictures 26 to be decoded by decoder 50 may be defined by the selection of the maximum temporal sub-layer to decode.

[0060] As mentioned above, the subset of pictures 26 to decode may be further defined by selecting a subset of layers 24 included in the video bitstream 14 for decoding. The video bitstream 14 may provide several options for a decodable bitstream. For example, a single layer of the video bitstream 14 may represent a decodable bitstream that may be decodable by the decoder 50 independently of additional layers. Alternatively, a combination of layers, or all layers, of the video bitstream 14 may represent a decodable bitstream and may be selected for decoding. For example, the video bitstream 14 may include an OLS indication, e.g., description data 23 of the video bitstream 14. The OLS indication may indicate one or more output layer sets (OLS). Each OLS may indicate one or more of the layers 24 of the video bitstream 14 as belonging to the OLS. In other words, an OLS may include one or more or all of the layers 24. In an example, one or more or all of the layers of an OLS may be indicated as output layers of the OLS. The OLS may optionally further include non-output layers. For example, in the example of a quality scalable bitstream, a reference layer for an output layer of the OLS may be included in the OLS because an output layer that references a reference layer may need the reference layer to be decoded, but the reference layer itself does not necessarily need to be an output layer.

[0061] Decoder 50 may select, for example, based on instructions provided by external means, one OLS from among multiple OLSs indicated in an OLS indication of video bitstream 14 for decoding. In other examples, decoder 50 may itself select an OLS to decode. Thus, the bitstream to be decoded may be defined by selecting an OLS and a maximum temporal sublayer for decoding.

[0062] For example, the decoder 50 may provide a decoded picture 53 to the decoded picture buffer for each picture that is part of an output layer of the OLS to be decoded and that is included in a temporal subset of pictures, for example defined by a maximum temporal sublayer.

[0063] For example, the maximum temporal sub-layer to decode may be represented by a variable Htid, described below, which may be provided to or derived by the decoder 50 .

[0064] The decoder 50 can output the decoded pictures 53 from the output buffer, i.e., the decoded picture buffer, at the output time. In other words, the decoder 50 can determine, for each of the decoded pictures 53, the output time at which the respective picture should be output. For example, the output time of the picture can be provided to the decoder 50 in description data of the video bitstream 14. For example, the output time of the picture can be provided by a Picture Timing (PT) Supplemental Enhancement Information (SEI) message in the video bitstream 14. For example, a PT SEI message can be provided for each of the access units 22. However, the video bitstream 14 does not necessarily need to provide such output timing information. Rather, the decoder 50 can determine the output timing of the picture. For example, the decoder 50 can derive the output timing of the picture by itself if the bitstream decoded by the decoder 50 has a constant output picture rate.

[0065] The current (e.g., VVC) specification includes the following text, which indicates that the output picture rate of the bitstream is constant:

[0066] For a CVS containing picture n, if Htid is equal to i, fixed_pic_rate_general_flag[i] is equal to 1, and picture n is an output picture that is not the last picture (in output order) in the output bitstream, the calculated value for DpbOutputElementalInterval[n] is equal to ClockTick * (elemental_duration_in_tc_minus1[i]+1), where ClockTick is the output order nextPicInOutputOrder specified for use in equation C.16, as specified in equation C.1 (using the value of ClockTick of the CVS containing picture n) when any of the following conditions are true:

[0067] - Picture nextPicInOutputOrder is in the same CVS as picture n.

[0068] - The picture nextPicInOutputOrder is in another CVS, and fixed_pic_rate_general_flag[i] is equal to 1 in the CVS that contains the picture nextPicInOutputOrder, the value of ClockTick is the same in both CVSs, and the value of elemental_duration_in_tc_minus1[i] is the same in both CVSs.

[0069] For the CVS containing picture n, if Htid is equal to i, fixed_pic_rate_within_cvs_flag[i] is equal to 1, and picture n is an output picture that is not the last picture (in output order) in the output CVS, the calculated value for DpbOutputElementalInterval[n] is ClockTick *(elemental_duration_in_tc_minus1[i]+1), where ClockTick is as specified in equation C.1 (using the value of ClockTick in the CVS containing picture n) when the next picture in the output order nextPicInOutputOrder specified for use in equation C.16 is in the same CVS as picture n.

[0070] In summary, there are two control flags in the SPS (sequence parameter set, e.g., description data, associated with each coded video sequence, e.g., coded video sequence (CVS) 20) as part of the hypothetical reference decoder (HRD) parameters. One of the control flags is fixed_pic_rate_within_cvs_flag, which indicates that within a CVS (or CVS), all output pictures have equidistant output times. The other control flag is fixed_pic_rate_general_flag, which indicates that a CVS starting from the first AU that references an SPS containing such a flag will also satisfy the equidistant output times between output pictures at the boundary with the previous CVS, as long as the value of ClockTick is the same in both CVSs and the value of elemental_duration_in_tc_minus1 is the same.

[0071] Note that signaling of whether the output rate is constant (fixed_pic_rate_within_cvs_flag) is given for each sub-layer, e.g., temporal sub-layer 25. That is, when a bitstream that allows temporal scalability is generated, this (constant picture rate) property is signaled for each possible frame rate that can be achieved when different numbers of sub-layers are received. HTid refers to the highest temporal ID present in the bitstream. For example, if the original bitstream had four sub-layers with temporal IDs from 0 to 3 and the highest one was dropped, HTid would become 2, and the parameter fixed_pic_rate_within_cvs_flag for HTid=2 would be taken into account by the decoder to evaluate whether the output rate is constant.

[0072] The problem is that this solution requires modifying the SPS (e.g., fixed_pic_rate_general_flag) when splicing or editing involving CVS concatenation is performed, as shown in Figure 7.

[0073] Figure 7 shows examples of SPS modifications that may be made to a coded video sequence relative to a previously coded video sequence. The upper panel of Figure 7 shows a first case in which splicing a first video sequence 201 and a second video sequence 202 results in a bitstream whose picture rate, defined by a timing interval 71 between successive pictures, is constant at the boundary between the first video sequence 201 and the second video sequence 202. The lower panel of Figure 7 shows a second case in which the timing interval 71 and the boundary between the first video sequence 201 and the second video sequence 202 are not constant at the boundary.

[0074] In FIG. 7 and the following FIGS. 8-10 and 12, pictures belonging to a common temporal sub-layer 25 are shown at a common level with respect to their vertical position.

[0075] One advantage of indicating this information (e.g., fixed_pic_rate_general_flag) in the bitstream beyond signaling that the bitstream has a constant output frame rate is that this information also allows the constant output frame rate characteristics to be used to derive output times without using (or ignoring) even more complex HRD parameters, such as buffering period SEI messages or picture timing SEI messages, to derive output times of decoded pictures. That is, for example, the PT SEI message and / or the BP SEI message may be omitted, i.e., not present in video bitstream 14, or may be ignored by decoder 50 when deriving output times of decoded pictures.

[0076] One problem that requires specifications to derive output times in this way (e.g., without using PT and / or BP SEI messages) is that there is no way to determine the output time of the first AU in a CVS. If the output time of the first AU in a CVS is known, it is possible to add that signaled delta (ClockTick) to the previous output picture when the CVS output time is flagged as constant. * (elemental_duration_in_tc_minus1[i]+1)) to easily determine the output time of the subsequent picture.

[0077] Also, note that the current specification states:

[0078] When htid is equal to i, elemental_duration_in_tc_minus1[i] plus 1 (if present) specifies the temporal distance in clock ticks between elemental units that specify the HRD output times of successive pictures in the output order specified below. The value of elemental_duration_in_tc_minus1[i] shall be in the range from 0 to 2047 inclusive.

[0079] For the CVS containing picture n, if Htid is equal to i, fixed_pic_rate_within_cvs_flag[i] is equal to 1, and picture n is an output picture that is not the last picture (in output order) of the CVS being output, then the value of the variable DpbOutputElementalInterval[n] is defined by:

[0080] DpbOutputElementalInterval[n]=DpbOutputInterval[n]×elementalOutputPeriods(113) where DpbOutputInterval[n] is defined in equation C.16, and elementalOutputPeriods is defined as follows:

[0081] - If a PT SEI message is present for picture n, elementalOutputPeriods is equal to the value of pt_display_elemental_periods_minus1+1.

[0082] - Otherwise, elementalOutputPeriods is 1.

[0083] This means that the constant output rate does not necessarily apply to the pictures being decoded, but to the pictures being displayed / output. That is, it does not apply to DpbOutputInterval[n], but to DpbOutputElementalInterval[n]. This means that the constant output rate includes frame repetition, i.e., elementalOutputPeriods not equal to 1 means that a particular picture is repeated. Figure 8 shows an example.

[0084] FIG. 8 shows an example video bitstream, such as the example in the top panel of FIG. 7, with access unit 22 * In such a situation, which may occur, for example, when a picture is lost during transmission, or is incorrectly decoded, or is omitted from the output, or is not present, the picture 26 of the previous access unit is used to achieve a constant frame rate. * can be repeated. For example, picture 26 * The access unit to which the picture belongs may have an indication that the picture of the access unit is repeated. For example, the indication may indicate the number of repetitions.

[0085] For example, DpbOutputInterval[n] may represent the duration of a time interval for output of an access unit 22, i.e., the output of content belonging to a common time frame, such as repeated output of pictures of an access unit, e.g., access unit output interval 61 in Figure 8. In contrast, DpbOutputElementalInterval[n] may represent the output of a single element, i.e., a time interval for a picture or a repetition of a picture, e.g., picture output interval 63 in Figure 8.

[0086] The syntax element (pt_display_elemental_periods_minus1) used below for repetition does not always mean a repetition of frames, and may also be used for interlaced content when frames that are encoded and decoded as frames are displayed as fields in the display step.

[0087] See the specification text below.

[0088] When sps_field_seq_flag is equal to 0 and fixed_pic_rate_within_cvs_flag[TemporalId] is equal to 1, a value of pt_display_elemental_periods_minus1 greater than 0 may be used to indicate the frame repetition period for displays using the same fixed frame refresh interval as DpbOutputElementalInterval[n], as given in Equation 113.

[0089] The following syntax examples and their corresponding semantics are provided for illustrative purposes and ease of understanding. [Table 1-1] [Table 1-2] The PT SEI message provides information on the CPB removal delay and DPB output delay of the AU associated with the SEI message.

[0090] If bp_nal_hrd_params_present_flag or bp_vcl_hrd_params_present_flag of the BP SEI message applicable to the current AU is equal to 1, then the variable CpbDpbDelaysPresentFlag is set equal to 1. Otherwise, CpbDpbDelaysPresentFlag is set equal to 0.

[0091] The presence of a PT SEI message is defined as follows:

[0092] - If CpbDpbDelaysPresentFlag is 1, the PT SEI message shall be associated with the current AU.

[0093] - Otherwise (CpbDpbDelaysPresentFlag is equal to 0), there shall be no PT SEI message associated with the current AU.

[0094] The TemporalId of the PT SEI message syntax is the TemporalId of the SEI NAL unit that contains the PT SEI message.

[0095] pt_cpb_removal_delay_minus1[i] plus 1 is used to calculate the nominal CPB removal time between the AU associated with the PT SEI message and the preceding AU in decoding order that contains the BP SEI message when Htid is i. This value is also used to calculate the earliest time that AU data can arrive at the CPB in the HSS. The length of pt_cpb_removal_delay_minus1[i] is bp_cpb_removal_delay_length_minus1+1 bits.

[0096] Specifies that when pt_cpb_alt_timing_info_present_flag is equal to 1, the following syntax elements may be present in the PT SEI message: pt_nal_cpb_alt_initial_removal_delay_delta[i][j], pt_nal_cpb_alt_initial_removal_offset_delta[i][j], pt_nal_cpb_delay_offset[i], pt_nal_dpb_delay_offset[i], pt_vcl_cpb_alt_initial_removal_delay_delta[i][j], pt_vcl_cpb_alt_initial_removal_offset_delta[i][j], pt_vcl_cpb_delay_offset[i], and pt_vcl_dpb_delay_offset[i]. It specifies that these syntax elements are not present in the PT SEI message if pt_cpb_alt_timing_info_present_flag is equal to 0. If the associated picture is a RASL picture, the value of pt_cpb_alt_timing_info_present_flag shall be equal to 0.

[0097] NOTE 1 - The value of pt_cpb_alt_timing_info_present_flag may be equal to 1 for multiple AUs that follow the IRAP picture in decoding order, but the alternative timing applies only to the first AU that has pt_cpb_alt_timing_info_present_flag equal to 1 and that follows the IRAP picture in decoding order.

[0098] pt_nal_cpb_alt_initial_removal_delay_delta[i][j] specifies the alternative initial CPB removal delay delta of the ith sublayer of the jth CPB of the NAL HRD in units of 90 kHz clock. The length of pt_nal_cpb_alt_initial_removal_delay_delta[i][j] is bp_cpb_initial_removal_delay_length_minus1+1 bits.

[0099] If pt_cpb_alt_timing_info_present_flag is equal to 1 and pt_nal_cpb_alt_initial_removal_delay_delta[i][j] is not present for values ​​of i less than bp_max_sublayers_minus1, its value is inferred to be equal to 0.

[0100] pt_nal_cpb_alt_initial_removal_offset_delta[i][j] specifies the alternative initial CPB removal offset delta of the ith sublayer of the jth CPB of the NAL HRD in units of 90 kHz clock. The length of pt_nal_cpb_alt_initial_removal_offset_delta[i][j] is bp_cpb_initial_removal_delay_length_minus1+1 bits.

[0101] If pt_cpb_alt_timing_info_present_flag is equal to 1 and pt_nal_cpb_alt_initial_removal_offset_delta[i][j] is not present for values ​​of i less than bp_max_sublayers_minus1, its value is inferred to be equal to 0.

[0102] pt_nal_cpb_delay_offset[i] specifies, for the i-th sublayer of the NAL HRD, the offset to be used in deriving the nominal CPB removal time of the AU associated with a PT SEI message and the AU that follows in decoding order, if the AU associated with the PT SEI message directly follows the AU associated with a BP SEI message in decoding order. The length of pt_nal_cpb_delay_offset[i] is bp_cpb_removal_delay_length_minus1+1 bits. If not present, the value of pt_nal_cpb_delay_offset[i] is inferred to be equal to 0.

[0103] pt_nal_dpb_delay_offset[i] specifies, for the i-th sublayer of the NAL HRD, the offset to be used in deriving the DPB output time of an IRAP AU associated with a BP SEI message when the AU associated with a PT SEI message directly follows the IRAP AU associated with the BP SEI message in decoding order. The length of pt_nal_dpb_delay_offset[i] is bp_dpb_output_delay_length_minus1+1 bits. If not present, the value of pt_nal_dpb_delay_offset[i] is inferred to be equal to 0.

[0104] pt_vcl_cpb_alt_initial_removal_delay_delta[i][j] specifies the alternative initial CPB removal delay delta of the ith sublayer of the jth CPB of the VCL HRD in units of 90 kHz clock. The length of pt_vcl_cpb_alt_initial_removal_delay_delta[i][j] is bp_cpb_initial_removal_delay_length_minus1+1 bits.

[0105] If pt_cpb_alt_timing_info_present_flag is equal to 1 and pt_vcl_cpb_alt_initial_removal_delay_delta[i][j] is not present for values ​​of i less than bp_max_sublayers_minus1, its value is inferred to be equal to 0.

[0106] pt_vcl_cpb_alt_initial_removal_offset_delta[i][j] specifies the alternative initial CPB removal offset delta of the ith sublayer of the jth CPB of the VCL HRD in units of 90 kHz clock. The length of pt_vcl_cpb_alt_initial_removal_offset_delta[i][j] is bp_cpb_initial_removal_delay_length_minus1+1 bits.

[0107] If pt_cpb_alt_timing_info_present_flag is equal to 1 and pt_vcl_cpb_alt_initial_removal_offset_delta[i][j] is not present for values ​​of i less than bp_max_sublayers_minus1, its value is inferred to be equal to 0.

[0108] pt_vcl_cpb_delay_offset[i] specifies, for the i-th sublayer of the VCL HRD, the offset to be used in deriving the nominal CPB removal time for the AU associated with a PT SEI message and the AU that follows in decoding order, if the AU associated with a PT SEI message directly follows the AU associated with a BP SEI message in decoding order. The length of pt_vcl_cpb_delay_offset[i] is bp_cpb_removal_delay_length_minus1+1 bits. If not present, the value of pt_vcl_cpb_delay_offset[i] is inferred to be equal to 0.

[0109] pt_vcl_dpb_delay_offset[i] specifies, for the i-th sublayer of the VCL HRD, the offset to be used in deriving the DPB output time of an IRAP AU associated with a BP SEI message when the AU associated with the PT SEI message directly follows the IRAP AU associated with the BP SEI message in decoding order. The length of pt_vcl_dpb_delay_offset[i] is bp_dpb_output_delay_length_minus1+1 bits. If not present, the value of pt_vcl_dpb_delay_offset[i] is inferred to be equal to 0.

[0110] The variable BpResetFlag for the current picture is derived as follows:

[0111] - BpResetFlag is set to 1 if the current picture is associated with a BP SEI message.

[0112] Otherwise, BpResetFlag is set to 0.

[0113] pt_sublayer_delays_present_flag[i] equal to 1 specifies that pt_cpb_removal_delay_delta_idx[i] or pt_cpb_removal_delay_minus1[i] and pt_du_common_cpb_removal_delay_increment_minus1[i] or pt_du_cpb_removal_delay_increment_minus1[][] are present in the sublayer with TemporalId equal to i. sublayer_delays_present_flag[i] equal to 0 specifies that for the sublayer with TemporalId equal to i, neither pt_cpb_removal_delay_delta_idx[i] nor pt_cpb_removal_delay_minus1[i] is present, nor pt_du_common_cpb_removal_delay_increment_minus1[i] nor pt_du_cpb_removal_delay_increment_minus1[][] is present. The value of pt_sublayer_delays_present_flag[bp_max_sublayers_minus1] is inferred to be equal to 1. If not present, the value of pt_sublayer_delays_present_flag[i] for any i in the range 0 to bp_max_sublayers_minus1-1 (inclusive) is inferred to be equal to 0.

[0114] A pt_cpb_removal_delay_delta_enabled_flag[i] equal to 1 specifies that pt_cpb_removal_delay_delta_idx[i] is present in the PT SEI message. A pt_cpb_removal_delay_delta_enabled_flag[i] equal to 0 specifies that pt_cpb_removal_delay_delta_idx[i] is not present in the PT SEI message. If not present, the value of pt_cpb_removal_delay_delta_enabled_flag[i] is inferred to be equal to 0.

[0115] pt_cpb_removal_delay_delta_idx[i] specifies the index of the CPB removal delta to apply to Htid equal to i in the list in bp_cpb_removal_delay_delta_val[j], for j in the range 0 to bp_num_cpb_removal_delay_deltas_minus1 inclusive. The length of pt_cpb_removal_delay_delta_idx[i] is Ceil(Log2(bp_num_cpb_removal_delay_deltas_minus1+1)) bits. If pt_cpb_removal_delay_delta_idx[i] is not present and pt_cpb_removal_delay_delta_enabled_flag[i] is equal to 1, the value of pt_cpb_removal_delay_delta_idx[i] is inferred to be equal to 0.

[0116] The variables CpbRemovalDelayMsb[i] and CpbRemovalDelayVal[i] for the current picture are derived as follows:

[0117] If the current AU is the AU that initializes the HRD, then CpbRemovalDelayMsb[i] and CpbRemovalDelayVal[i] are both set equal to 0, and the value of cpbRemovalDelayValTmp[i] is set equal to pt_cpb_removal_delay_minus1[i]+1.

[0118] - Otherwise, let picture prevNonDiscardablePic be the previous picture in decoding order that is not RASL or RADL and has TemporalId equal to 0, and set prevCpbRemovalDelayMinus1[i], prevCpbRemovalDelayMsb[i], and prevBpResetFlag equal to the values ​​of cpbRemovalDelayValTmp[i]-1, CpbRemovalDelayMsb[i], and BpResetFlag of picture prevNonDiscardablePic, respectively, and then the following apply:

[0119] -CpbRemovalDelayMsb[i] is derived as follows:

[0120] cpbRemovalDelayValTmp[i]=pt_cpb_removal_delay_delta_enabled_flag[i]? pt_cpb_removal_delay_minus1[bp_max_sublayers_minus1]+1+ bp_cpb_removal_delay_delta_val[pt_cpb_removal_delay_delta_idx[i]]: pt_cpb_removal_delay_minus1[i]+1 if(prevBpResetFlag) CpbRemovalDelayMsb[i]=0 elseif(cpbRemovalDelayValTmp[i] <prevCpbRemovalDelayMinus1[i]) CpbRemovalDelayMsb[i]=prevCpbRemovalDelayMsb[i]+2 bp_cpb_removal_delay_length_minus1+1 (D.1) else CpbRemovalDelayMsb[i]=prevCpbRemovalDelayMsb[i] -CpbRemovalDelayVal is derived as follows:

[0121] if(pt_sublayer_delays_present_flag[i]) CpbRemovalDelayVal[i]=CpbRemovalDelayMsb[i]+cpbRemovalDelayValTmp[i](D.2) else CpbRemovalDelayVal[i]=CpbRemovalDelayVal[i+1] The value of CpbRemovalDelayVal[i] is between 1 and 2. 32 The range is assumed to be within the range (inclusive).

[0122] The variable AuDpbOutputDelta[i] is derived as follows:

[0123] AuDpbOutputDelta[i]=CpbRemovalDelayVal[i]- (pt_cpb_removal_delay_minus1[bp_max_sublayers_minus1]+1)-(D.3) (i==bp_max_sublayers_minus1?0:bp_dpb_output_tid_offset[i]) Here, the value of bp_dpb_output_tid_offset[i] is found in the associated BP SEI message.

[0124] pt_dpb_output_delay is used to calculate the DPB output time of a picture. It specifies the number of clock ticks to wait before the decoded picture is output from the DPB after removing the AU from the CPB.

[0125] NOTE 2 – Decoded pictures are not removed from the DPB on output if they remain marked as "used for short-term reference" or "used for long-term reference".

[0126] The length of pt_dpb_output_delay is bp_dpb_output_delay_length_minus1+1 bits. If max_dec_pic_buffering_minus1[Htid] is equal to 0, the value of pt_dpb_output_delay shall be equal to 0.

[0127] The output time derived from the pt_dpb_output_delay of any picture output from an output timing compliant decoder shall precede the output times derived from the pt_dpb_output_delay of all pictures in any subsequent CVS in decoding order.

[0128] The picture output order established by the value of this syntax element shall be the same as the order established by the value of PicOrderCntVal.

[0129] For pictures that are not output by the "bump" process because they precede in decoding order a CLVSS picture whose ph_no_output_of_prior_pics_flag is 1 or is presumed to be 1, the output time derived from pt_dpb_output_delay shall increase as the value of PicOrderCntVal increases for all pictures in the same CVS.

[0130] pt_dpb_output_du_delay is used to calculate the DPB output time of a picture when DecodingUnitHrdFlag is equal to 1. It specifies the sub-clock ticks to wait before the decoded picture is output from the DPB after removing the last DU of an AU from the CPB.

[0131] The length of the syntax element pt_dpb_output_du_delay is given in bits by bp_dpb_output_delay_du_length_minus1+1.

[0132] The output time derived from pt_dpb_output_du_delay of any picture output from an output timing compliant decoder shall precede the output times derived from pt_dpb_output_du_delay of all pictures in any subsequent CVS in decoding order.

[0133] The picture output order established by the value of this syntax element shall be the same as the order established by the value of PicOrderCntVal.

[0134] For pictures that are not output by the "bump" process because they precede in decoding order a CLVSS picture whose ph_no_output_of_prior_pics_flag is 1 or is presumed to be 1, the output time derived from pt_dpb_output_du_delay shall increase as the value of PicOrderCntVal increases for all pictures in the same CVS.

[0135] For any two pictures in the CVS, the difference in output time between the two pictures when DecodingUnitHrdFlag is equal to 1 shall be the same as the same difference when DecodingUnitHrdFlag is equal to 0.

[0136] Pt_num_decoding_units_minus1 plus 1 specifies the number of DUs in the AU with which the PT SEI message is associated. The value of pt_num_decoding_units_minus1 shall be in the range from 0 to PicSizeInCtbsY-1, inclusive.

[0137] pt_du_common_cpb_removal_delay_flag equal to 1 specifies that the syntax element pt_du_common_cpb_removal_delay_increment_minus1[i] is present. pt_du_common_cpb_removal_delay_flag equal to 0 specifies that the syntax element pt_du_common_cpb_removal_delay_increment_minus1[i] is not present. If not present, pt_du_common_cpb_removal_delay_flag is inferred to be equal to 0.

[0138] pt_du_common_cpb_removal_delay_increment_minus1[i] plus 1 specifies the period, in clock subticks (see Section C.1), between the nominal CPB removal times of any two consecutive DUs, in decoding order, in the AU associated with the PT SEI message, when Htid is equal to i. This value is also used to calculate the earliest possible arrival time of DU data at the CPB in the HSS, as specified in Annex C. The length of this syntax element is bp_du_cpb_removal_delay_increment_length_minus1+1 bits.

[0139] If pt_du_common_cpb_removal_delay_increment_minus1[i] does not exist for values ​​of i less than bp_max_sublayers_minus1, its value is inferred to be equal to pt_du_common_cpb_removal_delay_increment_minus1[bp_max_sublayers_minus1].

[0140] pt_num_nalus_in_du_minus1[i] plus 1 specifies the number of NAL units in the i-th DU of the AU with which the PT SEI message is associated. The value of pt_num_nalus_in_du_minus1[i] shall be in the range from 0 to PicSizeInCtbsY-1, inclusive.

[0141] The first DU of an AU consists of the first pt_num_nalus_in_du_minus1[0]+1 consecutive NAL units in the AU's decoding order. The i-th DU (i greater than 0) of an AU consists of the pt_num_nalus_in_du_minus1[i]+1 consecutive NAL units that immediately follow the last NAL unit of the DU preceding the AU in decoding order. Each DU shall have at least one VCL NAL unit. All non-VCL NAL units associated with a VCL NAL unit shall be contained in the same DU as the VCL NAL unit.

[0142] pt_du_cpb_removal_delay_increment_minus1[i][j] plus 1 specifies the duration, in clock subticks, between the CPB removal time of the (i+1)th DU and the ith DU, in decoding order, in the AU associated with the PT SEI message, when Htid is equal to j. This value is also used to calculate the earliest possible time that DU data can arrive at the CPB in the HSS, as specified in Annex C. The length of this syntax element is bp_du_cpb_removal_delay_increment_length_minus1+1 bits.

[0143] If pt_du_cpb_removal_delay_increment_minus1[i][j] does not exist for values ​​of j less than bp_max_sublayers_minus1, its value is inferred to be equal to pt_du_cpb_removal_delay_increment_minus1[j][bp_max_sublayers_minus1].

[0144] pt_delay_for_concatenation_ensured_flag equal to 1 specifies that the nominal removal time of the next AU from the CPB calculated by bp_cpb_removal_delay_delta_minus1 should be applied if the difference between the final arrival time of the AU associated with the PT SEI message and the CPB removal time is followed by an AU with a BP SEI message for which bp_concatenation_flag is equal to 1 and InitCpbRemovalDelay[][][Ed.(YK): where "InitCpbRemovalDelay[Htid][ScIdx]" is less than or equal to the value of bp_max_initial_removal_delay_for_concatenation. pt_delay_for_concatenation_ensured_flag equal to 0 specifies that the difference between the final arrival time of the AU associated with the PT SEI message and the CPB removal time may or may not exceed the value of max_val_initial_removal_delay_for_splicing.

[0145] If sps_field_seq_flag is equal to 0 and fixed_pic_rate_within_cvs_flag[TemporalId] is equal to 1, the value of pt_display_elemental_periods_minus1 plus 1 indicates the number of elementary picture periodicity intervals that the current coded picture occupies for the display model.

[0146] If fixed_pic_rate_within_cvs_flag[TemporalId] is equal to 0 or sps_field_seq_flag is equal to 1, the value of pt_display_elemental_periods_minus1 shall be equal to 0.

[0147] When sps_field_seq_flag is equal to 0 and fixed_pic_rate_within_cvs_flag[TemporalId] is equal to 1, a value of pt_display_elemental_periods_minus1 greater than 0 may be used to indicate the frame repetition period for displays using the same fixed frame refresh interval as DpbOutputElementalInterval[n], as given in Equation 112.

[0148] We will resume the discussion of the issues raised here.

[0149] A further problem that needs to be solved is that a similar result may need to be obtained even if the PT SEI message is not present, i.e., since the PT SEI message is optional, repetition needs to be allowed even if the PT SEI message is not present.

[0150] Also note that there is an interaction between the Frame Field Information SEI message, which is required if sps_field_seq_flag is 1 and optional if it is 0, and the information in the PT SEI message (pt_display_elemental_periods_minus1). Such an SEI (Frame Field Information SEI) also has syntax elements with the same values ​​as the PT SEI message, namely: If present, the value of display_elemental_periods_minus1 plus 1 (which may only be coded if field_pic_flag is off, or may not be coded if field_pic_flag is on) and FixedPicRateWithinCvsFlag is equal to 1, indicates the number of elementary picture periodicity intervals that the current coded picture occupies for the display model. The value of display_elemental_periods_minus1 shall be equal to DisplayElementalPeriods-1, subject to the following constraints:

[0151] - If display_fields_from_frame_flag is 1, then display_elemental_periods_minus1 shall be 1 or 2.

[0152] Otherwise, if FixedPicRateWithinCvsFlag is equal to 0, then display_elemental_periods_minus1 shall be equal to 0.

[0153] The interpretation of combinations of field_pic_flag (in the frame field SEI; assumed to be equal to sps_field_seq_flag), FixedPicRateWithinCvsFlag, bottom_field_flag, display_fields_from_frame_flag, top_field_first_flag, and display_elemental_periods_minus1 (DisplayElementalPeriods) is specified in Table 14. Note that in the table, absent syntax elements are marked with "-". Combinations of syntax elements not listed in Table 14 are reserved for future use by ITU-T|ISO / IEC and shall not be present in bitstreams conforming to this version of the specification.

[0154] NOTE 1 - When FixedPicRateWithinCvsFlag is equal to 1, the indicated display time is constrained to take into account the duration for display for a display model that follows the display pattern indicated by the values ​​of the syntax elements in the Frame Field Information SEI message (however, the display process is outside the scope of this specification). Although a video decoder model may be specified to output only the entire cropped decoded picture, the modeled display behavior may include other steps, such as repeatedly displaying a frame over multiple time intervals when display_fields_from_frame_flag is equal to 0, or sequentially displaying individual fields of a frame when display_fields_from_frame_flag is 1.

[0155] NOTE 2 - Frame doubling can be used, for example, to facilitate the display of 25 Hz progressive scan video on a 50 Hz progressive scan display, or 30 Hz progressive scan video on a 60 Hz progressive scan display. Alternating combinations of frame doubling and tripling, used every other frame, can easily display 24 Hz progressive scan video on a 60 Hz progressive scan display.

[0156] Table 14 - Interpretation of Frame Field Information Syntax Elements [Table 2] Note that in multi-layer, if some of the output layers contain fields (e.g., in interlaced video, a picture may be split into a first field and a second field, and the first field and the second field may be output at consecutive time instances. Thus, the first field picture may be considered to belong to the first temporal sub-layer, e.g., 250, and the second field picture may be considered to belong to the second temporal sub-layer, e.g., 251, and the video sequence of the layer with fields will therefore have a higher, e.g., twice the frame rate), and if some layers do not (i.e., for example, a bitstream including both layers with and without fields), both sets will have different output frame rates, so in the more general case, this means that all multi-layer bitstreams with different output frame rates will have problems, as shown in Figure 9 below.

[0157] 9 shows a video bitstream having a first layer 241 and a second layer 240. The first layer 241 includes three temporal sublayers: sublayer 250, sublayer 251, and sublayer 252. In contrast, the second layer 240 includes pictures of the first temporal sublayer 250 and the second temporal sublayer 251, but does not include pictures of the third temporal sublayer 252. Thus, the second layer 240 has a lower frame rate than the first layer, e.g., half the frame rate.

[0158] In other words, Figure 9 shows an example of a bitstream with two layers having different frame rates: for example, upper layer 241 can have fields and lower layer 240 can have progressive frames (e.g., no interlacing).

[0159] In the example of Figure 9, the highest layer 241 has a frame rate twice that of the lower layer 240 and no repetition. However, if a fixed frame rate is signaled for OLS, repetition may actually be desired in the lowest layer 240, as shown in Figure 10 below.

[0160] Figure 10 shows the example video bitstream of Figure 9, where output frames in lower layer 240, which has a lower coding frame rate than upper layer 241, are repeated. The output frame repetition may result in the output frame rates of both layers being equal. Similar to Figure 8, access unit output interval 63 of layer 240 includes multiple picture output intervals 63.

[0161] In summary, the following problems are solved by this aspect of the invention, where the numbering indicates sub-aspects, embodiments of which are described in the respective sections:

[0162] 1. Different frame rates of output layers in the output layer set (e.g. multiview where enhancement layers have fields and base layers have frames) 2. Interaction between PT SEI message and Frame Field Information SEI message 3. Frame / field repetition without PT SEI message 4. A constant output frame rate between CVSs is derived, otherwise it is not signaled that SPSs would need to be rewritten after splicing.

[0163] 5. Derivation of output time when PT SEI message is not present or used.

[0164] 6. Handling of Pictures Not Outputted and the Impact on a Constant Output Frame Rate: In practice, although not mentioned above, if some pictures are not outputted, are not present in the PT SEI message, or are ignored, it may be complicated to derive the output time.

[0165] Before describing the sub-aspects of the second aspect, a brief summary of the above-described means for determining the output time of a decoded picture will be provided. For example, the video bitstream 14 may include a PT SEI message conveying information about picture output timing. The PT SEI message may include a picture output multiplication syntax element, such as pt_display_elemental_periods_minus1. For example, the PT SEI message may signal information about the access unit level. In other words, the PT SEI message may be related to an access unit 22. That is, the PT SEI message may be valid for all pictures in one access unit 22. The picture output multiplication syntax element signaled in the PT SEI message may reveal information about whether the access unit referenced by the respective PT SEI message is subject to multiplied picture output, as shown for the lower layer 240 in FIG. 10, for example. For example, pt_display_elemental_periods_minus1 equal to zero may indicate that the picture output is not multiplied, and pt_display_elemental_periods_minus1 > 0 may indicate that the picture output is repeated. If the picture output is multiplied, the picture output multiplication syntax element may indicate, for example, how many output pictures are generated from one picture of each access unit, depending on the value of pt_display_elemental_periods_minus1.

[0166] Note that the decoder can derive a variable called elementalOutputPeriods with pt_display_elemental_periods_minus1.

[0167] For example, decoder 50 may set the elemental output periods equal to pt_display_elemental_periods_minus1+1.

[0168] In other words, encoder 10 may encode the PT SEI message, video bitstream 14 may include the PT SEI message, and decoder 50 may decode the PT SEI message.

[0169] Additionally, video bitstream 14 may include a frame-field supplemental extension information (frame-field SEI) message, also referred to as an FFI SEI message, encoded by encoder 10 and decoded by decoder 50, which conveys information on the frame-field structure of a given access unit, such as whether the pictures of that access unit are frame- or field-coded, and if frame-coded, whether they are converted to fields or not, and the order in which the pictures are output: bottom-field or top-field (first field vs. second field). For example, the FFI SEI message may include a further picture output multiplication syntax element, such as FFI_display_elemental_periods_minus1. For example, the further picture output multiplication syntax element may indicate, for example, for a frame-field use case, whether a coded picture is repeated, i.e., whether the picture referenced by the FFI SEI message is the target of multiple-picture output. Note that, in contrast to a PT SEI message, the FFI SEI may reference a single picture rather than an entire access unit.

[0170] In the following, the above-mentioned picture output multiplication syntax element of the PT SEI message may be referred to as a PT multiplication indicator, and the further picture output multiplication syntax element of the FFI SEI message may be referred to as an FF multiplication indicator.

[0171] For example, when the decoder 50 derives information that a picture is subject to multiplied picture output, the decoder 50 can set the number of picture outputs, e.g., the repetition or generation of fields from the decoded frame, for each access unit (in the case of a PT multiplication representation) or for each picture (e.g., in the case of an FF multiplication representation), according to the number indicated by the picture output multiplication syntax element or the further picture output multiplication syntax element, respectively. In other words, the decoder 50 can provide one or more repetitions of each picture to the output buffer. Note that, according to embodiments of the present disclosure, this may only be true in certain circumstances.

[0172] 2.1. Different frame rates for different output layers in an output layer set As explained, if there are different frame rates for different output layers of an output layer set, this can be problematic, since it is not possible to indicate different repetition patterns or different values ​​of elementalOutputPeriods with a single value of pt_display_elemental_periods_minus1, since all HRD SEI messages (BP, PT, DUI) apply globally to the respective AU, i.e., to each picture within the AU without distinguishing between layers. Note that this also applies to the embodiments of other sub-aspects, e.g., sub-aspects 2.3 and 2.5.

[0173] Thus, in one embodiment, there is a gating flag (eg, pt_display_elemental_periods_present_flag) in the PT SEI message that indicates whether elementalOutputPeriods is set in the PT SEI message.

[0174] [Table 3] Thus, according to one embodiment, decoder 50 may decode the PT SEI message for access unit 22. Decoder 50 may decode, from the picture timing supplemental enhancement information message, a gating flag (e.g., pt_display_elemental_periods_present_flag) and a picture output multiplication syntax element (e.g., pt_display_elemental_periods_minus1>0) that reveals information about whether a given access unit is eligible for multiplied picture output if the gating flag is in a first state, and if so (e.g., pt_display_elemental_periods_minus1 is greater than or equal to 0), how many output pictures will be generated from the given access unit (e.g., pt_display_elemental_periods_minus1>0).

[0175] Semantics of when pt_display_elemental_periods_minus1 is inferred: If not present, the value of pt_display_elemental_periods_minus1 is unused and therefore not inferred. Instead, the information is obtained by other means as per 2.3.

[0176] This case applies unless the aspects of 2.5 are taken into account, in which case the constraint is signaled that the value of elementalOutputPeriods is constrained to 1 (see 2.5).

[0177] In another embodiment, there is a bitstream constraint that requires pt_display_elemental_periods_present_flag to be equal to 0 if any of the following is true:

[0178] - The frame rate of the OLS output layer is different ·The values ​​of sps_field_seq_flag of all SPSs referenced by the output layer of the OLS corresponding to the bitstream are not the same.

[0179] 2.2. Interaction between PT SEI Message and Frame Field Information SEI Message As explained, there may be a Frame Field Information SEI message that provides detailed information about how to output a frame. For that purpose, in one embodiment, the use of information in the Picture Timing SEI message in addition to the Frame Field Information SEI message is used in a constrained manner. (Note that according to the above embodiment, this only applies if a PT SEI message is present and the syntax element pt_display_elemental_periods_present_flag is equal to 1.) If sps_field_seq_flag of the SPS referenced by the VCL NAL units of a layer is equal to 1 (the bitstream contains fields), pt_display_elemental_periods_minus1 shall be equal to 0 regardless of the value of the Frame Field Information SEI message. Otherwise, if sps_field_seq_flag of the SPS referenced by the VCL NAL units of a layer is equal to 0 (the bitstream contains frames), display_fields_from_frame_flag is equal to 0 (frames are not displayed as fields), and fixed_pic_rate_within_cvs_flag[TemporalId] is equal to 0, the value of pt_display_elemental_periods_minus1 shall be 0, i.e., in such cases there is no constant output frame rate and no fields are output from frames.

[0180] According to an embodiment of sub-aspect 2.2, the encoder 10 is configured to, for a given access unit 22 of the video data stream, encode in the video data stream 14 a PT SEI message conveying information about picture output timing for the given access unit. Furthermore, the encoder 10 is configured to encode, for a picture sequence including pictures in the given access unit, a sequence parameter set (e.g., SPS) including a frame field syntax element (e.g., sps_field_seq_flag) indicating whether the pictures of the picture sequence represent fields or frames (e.g., progressive frames), and a fixed picture rate flag (e.g., in the SPS or VPS) indicating whether the output of the video data stream involves a fixed picture rate. The encoder 10 is configured to set a picture output multiplication syntax element (e.g., pt_display_elemental_periods_minus1) as indicating that there is no multiplied picture output for the given access unit, For a frame-field syntax element indicating that a picture in a picture sequence represents a field, and / or For a frame-field syntax element indicating that a picture in a picture sequence represents a frame, including setting a frame-to-field syntax element in the video data stream indicating that the frame is not displayed as a field (e.g., indicated in a frame-field supplemental extension information message or inferred in its absence), and a fixed picture rate flag indicating picture output without a fixed picture rate.

[0181] According to an embodiment, the video data stream 14 is a multi-layer video data stream including an output layer set (OLS) of one or more output layers (e.g., those layers for which pictures are output, possibly one or more reference layers that are not output but serve as reference layers). According to this embodiment, the picture timing supplemental enhancement information message conveys information regarding picture output timing for all output layers of a multi-layer video data stream having pictures coded in a given access unit. According to this embodiment, the picture sequence is of a given output layer and includes pictures of the given output layer in the given access unit, and a frame field syntax element (e.g., sps_field_seq_flag) indicates whether pictures of the picture sequence of the given output layer represent fields or frames. According to this embodiment, a fixed picture rate flag (e.g., in the SPS or VPS) indicates whether the picture output for the output layer set includes a fixed picture rate for the output layer set. According to this embodiment, the encoder is configured to set a picture output multiplication syntax element (e.g., pt_display_elemental_periods_minus1) to indicate that there is no multiplied picture output for a given access unit; For a frame-field syntax element indicating that a picture in a picture sequence for a given output layer represents a field, and / or For a frame field syntax element indicating that a picture of a picture sequence of a given output layer represents a frame, it is configured to set a frame-to-field syntax element (e.g., display_fields_from_frame_flag) indicating that the frame is not displayed as a field (e.g., indicated in a frame field supplemental extension information message or inferred in its absence), and a fixed picture rate flag indicating picture output of the output layer set without a fixed picture rate.

[0182] According to an embodiment, the encoder 10 is configured to encode, for a given access unit of the video data stream, an FFI SEI message into the video data stream 14 that conveys information about the frame-to-field structure for the given access unit, including a frame-to-field syntax element (such as display_fields_from_frame_flag).

[0183] According to an embodiment, the video data stream 14 is a multi-layer video data stream including an output layer set (OLS) of one or more output layers (e.g., those layers for which pictures are output, possibly one or more reference layers that are not output but serve as reference layers). According to this embodiment, the picture timing supplemental enhancement information message conveys information about picture output timing for all output layers of the multi-layer video data stream having pictures coded in a given access unit. According to this embodiment, the frame field supplemental enhancement information message is specific to a given output layer of the multi-layer video data stream and conveys information about a frame field structure associated with the given output layer in a given access unit. According to this embodiment, the picture sequence is of a given output layer and includes pictures of the given output layer in the given access unit, and a frame field syntax element (e.g., sps_field_seq_flag) indicates whether pictures of the picture sequence of the given output layer represent fields or frames. According to this embodiment, a fixed picture rate flag (e.g., in the SPS or VPS) indicates whether picture output for the output layer set includes a fixed picture rate for the output layer set. According to this embodiment, the encoder is configured to set a picture output multiplication syntax element (e.g., pt_display_elemental_periods_minus1) to indicate that there is no multiplied picture output for a given access unit; For a frame-field syntax element indicating that a picture in a picture sequence for a given output layer represents a field, and / or For a frame-field syntax element indicating that a picture in a picture sequence of a given output layer represents a frame, it is configured to set a frame-to-field syntax element indicating that the frame is not displayed as a field (e.g., indicated in a frame-field supplemental extension information message or inferred in its absence), and a fixed picture rate flag indicating picture output of the output layer set without a fixed picture rate.

[0184] 2.3.Repetition of a frame or field without a PT SEI message In another embodiment, the third problem listed above (frame / field repetition without a PT SEI message) is solved by external means (e.g., an API) for elementalOutputPeriods to be not just 1 when a PT is not present or when a frame field information SEI message is present. This also solves the problem indicated in 1) when the output frame rates of different layers are different, since there is no information in the PT SEI message related to elementalOutputPeriods and the frame field information SEI message provides this information as a per-layer SEI message.

[0185] When htid is equal to i, elemental_duration_in_tc_minus1[i] plus 1 (if present) specifies the temporal distance in clock ticks between elemental units that specify the HRD output times of successive pictures in the output order specified below. The value of elemental_duration_in_tc_minus1[i] shall be in the range from 0 to 2047 inclusive.

[0186] For the CVS containing picture n, if Htid is equal to i, fixed_pic_rate_within_cvs_flag[i] is equal to 1, and picture n is an output picture that is not the last picture (in output order) of the CVS being output, then the value of the variable DpbOutputElementalInterval[n] is defined by:

[0187] DpbOutputElementalInterval[n]=DpbOutputInterval[n]×elementalOutputPeriods(113) where DpbOutputInterval[n] is defined in equation C.16, and elementalOutputPeriods is defined as follows:

[0188] - If a PT SEI message is present for picture n and pt_display_elemental_periods_present_flag is equal to 1, then elementalOutputPeriods is equal to the value of pt_display_elemental_periods_minus1+1.

[0189] If an external means is provided, elementalOutputPeriods is set equal to the value of elementalOutputPeriods provided by the external means.

[0190] Otherwise (no external means are provided for setting the value of elementalOutputPeriods), if a Frame Field Information SEI message is provided for a layer with a predefined index, the value of elementalOutputPeriods is set to display_elemental_periods_minus1+1 (note that when the value of elementalOutputPeriods is set to display_elemental_periods_minus1+1, in this description display_elemental_periods_minus1 is an alias for the ffi_display_elemental_periods_minus1 syntax element described above).

[0191] - Otherwise, elementalOutputPeriods is 1.

[0192] Here, the layer with the predefined index to be used, e.g., the layer with the highest frame rate, is identified in one of the following ways:

[0193] indicated by additional signaling in the HRD SEI, VPS / SPS, or other means (e.g., fixed_pic_rate_layer_index), or The layer with the largest sublayer identifier value (temporal_id), or A frame field SEI message with a display_elemental_periods_minus1 value that is different from the pt_display_elemental_periods_minus1 of the corresponding PT SEI message Instead of using the layer index to identify which Frame File Info SEI message to use to determine the elementalOuputPeriods, one of the following is considered:

[0194] 1) Otherwise, if a Frame Field Information SEI message is provided for a layer for which there is an output picture in both AUn and the next AU in output order, i.e., the AU containing nextPicInOutputOrder, the value of elementalOutputPeriods is set to display_elemental_periods_minus1+1.

[0195] 2) Otherwise, if a Frame Field Information SEI message is provided for an output picture present in an AU, the value of elementalOutputPeriods is set to the minimum value display_elemental_periods_minus1+1 among all output layers.

[0196] Thus, according to an embodiment of sub-aspect 2.3, the decoder 50 is configured to derive the picture output number for a given access unit of the video data stream, e.g., the currently decoded access unit, according to one or more of the following criteria: if the picture output number for the given access unit is provided via the decoder's API, adopt the picture output number provided via the API, and / or if a frame field supplemental enhancement information message conveying information about the frame field structure of the given access unit and including a further picture output multiplication syntax element (e.g., display_elemental_periods_minus1) is present in the video data stream, decode a further picture output multiplication syntax element from the frame field supplemental enhancement information message and set the picture output number for the given access unit according to the further picture output multiplication syntax element.

[0197] Note that in general, a picture output multiplication syntax element can be expressed as pt_display_elemental_periods_minus1+1 or as pt_display_elemental_periods_minus1. That is, the value of a picture output multiplication syntax element can correspond to the picture output number or the picture output number minus 1. Similarly for further picture output multiplication syntax elements.

[0198] According to an embodiment, the video data stream 14 is a multi-layered video data stream and, as described above, includes a PTSEI message conveying information regarding picture output timing for all output layers of the multi-layered video data stream having pictures encoded in a given access unit.

[0199] According to an embodiment, the above set of criteria further includes the following: if a PT SEI message is present in the video data stream that includes a picture output multiplication syntax element (e.g., pt_display_elemental_periods_minus1), decode the picture output multiplication syntax element from the PT SEI message, and set the picture output count for a given access unit according to the picture output multiplication syntax element.

[0200] As mentioned above, the PT SEI message can relate to all pictures of a given access unit, and the FFI SEI message can convey information about the frame field structure of pictures of a given output layer that are encoded in a given access unit.

[0201] For example, decoder 50, when setting the number of picture outputs for a given access unit according to the further picture output multiplication syntax element, can set the picture output number of a given output layer in response to the further picture output multiplication syntax element and can use the further picture output multiplication syntax element to determine the inter-output picture interval (e.g., the variable DpbOutputElementalInterval introduced above) for the given access unit. According to an embodiment, decoder 50 performs this number setting selection if the given output layer has pictures coded in the given access unit and the access unit immediately following it in output order.

[0202] According to an embodiment, decoder 50 can perform setting of the number of picture outputs for a given access unit in accordance with the further picture output multiplication syntax element in accordance with one or more of the following criteria: According to a first criterion, if a given output layer has pictures coded in the given access unit and the access unit immediately following it in output order, when setting the number of picture outputs for the given access unit in accordance with the further picture output multiplication syntax element, the decoder 50 sets the picture output number for the given output layer in accordance with the further picture output multiplication syntax element and uses the further picture output multiplication syntax element to determine the inter-output picture interval (DpbOutputElementalInterval) for the given access unit (e.g., alternative 1 of the above alternatives that uses a layer's index to identify which frame file information SEI message to use to determine elementalOuputPeriods). According to the second criterion, if there is a picture coded in a given output layer in another output layer and this other output layer has a frame field supplemental enhancement information message that further includes an additional picture output multiplication syntax element, when setting the number of picture outputs for a given access unit according to the further picture output multiplication syntax element, the picture output number for the given output layer is set according to the further picture output multiplication syntax element, and the further picture output multiplication syntax element is used to determine the inter output picture interval (DpbOutputElementalInterval) for the given access unit (e.g., alternative 2 of the above alternative in which a layer index is used to identify which frame file information SEI message to use to determine the elementalOuputPeriod).

[0203] According to an embodiment, the plurality of output layers have coded pictures in a given output layer, including a frame field supplemental enhancement information message, and the number of further picture output multiplication syntax elements in the given output layer is minimal. According to this embodiment, when setting the number of picture outputs for a given access unit according to the further picture output multiplication syntax element, the decoder 50 sets the number of picture outputs for the given output layer according to the further picture output multiplication syntax element, and determines the inter-output picture interval (DpbOutputElementalInterval) for the given access unit using the further picture output multiplication syntax element.

[0204] According to this embodiment, when decoder 50 sets the number of picture outputs for a given access unit according to a picture output multiplication syntax element, it sets the picture output number equally for all output layers and uses the picture output multiplication syntax element to determine the inter-output picture interval (DpbOutputElementalInterval) for the given access unit. When decoder 50 sets the number of picture outputs for a given access unit according to a further picture output multiplication syntax element, it sets the picture output number for a given output layer according to the further picture output multiplication syntax element and uses the further picture output multiplication syntax element to determine the inter-output picture interval (DpbOutputElementalInterval) for the given access unit.

[0205] According to one embodiment, the decoder 50 determines a given output layer according to one of the following:

[0206] That is, based on the signaling of the multi-layer video data stream (e.g., specifying a given output layer), As the output layer with the highest temporal sublayer (e.g., the decoder determines which sublayers belong to each output layer and designates the highest output layer (hierarchically, i.e., the one on which no other temporal layers of the output layer depend, as the topmost temporal layer)), or The further picture output multiply syntax element determines as an output layer of an output layer set different from the picture output multiply syntax element.

[0207] According to an embodiment, encoder 10 can provide signaling for a predetermined output layer in a multi-layer video data stream (e.g., specifically indicating the predetermined output layer). Alternatively, encoder 10 can select the output layer with the highest temporal sublayer as the predetermined output layer (e.g., the decoder determines the sublayers belonging to each output layer and designates the highest output layer (hierarchically, i.e., the one on which other temporal layers of the output layer do not depend, as the topmost temporal layer)). Alternatively, encoder 10 can select the predetermined output layer as an output layer of an output layer set in which a further picture output multiplication syntax element differs from the picture output multiplication syntax element.

[0208] As an alternative to the above embodiment of sub-aspect 2.3, a further embodiment solves the problem by bitstream constraints, as will be explained, for example, with reference to FIG.

[0209] Figure 11 shows an encoder 10 according to an embodiment of sub-aspect 2.3. The encoder 10 of Figure 11 may optionally correspond to the encoder 10 of Figure 1. The video bitstream 14 according to this embodiment may be, for example, a single-layer video bitstream or a multi-layer video bitstream, as described with respect to Figure 1. As described in Section 0, the video bitstream 14 encodes a sequence of access units 22. Certain of the access units 22 are denoted by reference numeral 22 in Figure 11. *For example, the given access unit is the access unit currently being encoded. The encoder 10 according to this embodiment uses the given access unit 22 * , a PT SEI message, such as the PT SEI messages described earlier in this section, is encoded into the video bitstream 14. The PT SEI message is * The PT SEI message 73 conveys information about the picture output timing of a given access unit 22. The PT SEI message 73 includes a picture output multiplication syntax element 74, also referred to hereinafter as a PT multiplication indicator 73. For example, the PT multiplication indicator 73 may be used to indicate the picture output timing of a given access unit 22, as previously described. * It may also indicate the number of times the picture is output.

[0210] According to this embodiment, the encoder 10 further encodes an FFI SEI message 83 into the video data stream 14, which is included in a given access unit 22. * The FFI SEI message 83 conveys information about the frame field structure of a given access unit 22, for example, the FFI SEI message described earlier in this section. The FFI SEI message 83 includes a further picture output multiplication syntax element 84, also referred to hereinafter as FF multiplication indicator 84. As previously mentioned, the FF multiplication factor indicator 83 indicates the frame field structure of a given access unit 22. * For example, FFI SEI message 83 may reference one of the layers, e.g., one of the output layers of video bitstream 14.

[0211] The PT multiplication indicator 84 and the FF multiplication indicator 84 may be encoded into the video bitstream using a symbolization scheme, and for purposes of encoding the perspective syntax elements, the actual value of each syntax element may be derived by subtracting 1 from the number of picture outputs represented by the respective syntax element. In other words, in examples, the actual values ​​of the PT multiplication indicator 74 and the FF multiplication indicator 84 written into the video bitstream 14 may differ from the values ​​represented by the PT multiplication indicator 74 and the FF multiplication indicator 84 by, for example, a value of 1. Nevertheless, the values ​​of the PT multiplication indicator 74 and the FF multiplication indicator 84 shall be understood as values ​​representing the actual picture output times.

[0212] According to the embodiment of FIG. 11, the PT multiplication indicator is less than or equal to the FF multiplication indicator.

[0213] According to an embodiment, the information in the FFI SEI 83 is specific to a layer of the video bitstream 14. For example, the video bitstream 14 may include an FFI SEI message 84 for each output layer of the video bitstream 14. Alternatively, the FFI SEI messages 83 may be provided at the access unit level G, with one FFI SEI message 83 for a given access unit 22. * , and the FFI SEI message 83 may be provided to a given access unit 22 * The PT multiplication indicator 84 includes a respective multiplication indicator 84 for each of one or more output layers having the coded picture. According to this embodiment, the PT multiplication indicator 84 is less than or equal to all of the FF multiplication indicators 84 of the one or more output layers. For example, an output layer may represent a layer, the picture of which is considered to be output by the decoder 50, e.g., as described in the introduction to Section 2.

[0214] For example, video bitstream 14 may be a multi-layer video bitstream, and encoder 10 may provide one FFI SEI message 83 for each of one or more output layers of video bitstream 14, and thus provide one or more FFI SEI messages 83, each including a respective FF multiplication indicator 84. Each of the one FFI SEI messages 83 may reference one of the layers, one or more of which may be indicated as being output layers of the OLS represented in video bitstream 14. According to this embodiment, all of the FF multiplication indicators 84 signaled in each FFI SEI message 83 signaled for an output layer are equal to or greater than the PT multiplication indicator 74. Thus, the smallest value exceeding the FF multiplication indicator 84 is equal to or greater than the PT multiplication indicator 74.

[0215] According to one embodiment, a given access unit 22 * The PT multiplication indicator 74 of a given access unit 22 * is equal to the smallest value that exceeds the value of the FF multiplication indicator 84 of all FFI SEI messages of the output layer in

[0216] In other words, alternatively, pt_display_elemental_periods_minus1 in the PT SEI message applied to the AU is equal to the minimum value of display_elemental_periods_minus1 in all frame field information SEI messages of the output layer in the AU.

[0217] According to one embodiment, the FF multiplication indicator 84 is an integer multiple of the PT multiplication indicator 74. Note that this constraint is particularly valid for the actual values ​​of the picture output times represented by the respective syntax elements.

[0218] For example, the decoder 50 may derive the above-mentioned ElementalOuputPeriods and DisplayElementalPeriods by setting the element output period to PT_display_elemental_periods_minus1+1 and the display element period to display_elemental_periods_minus1+1. In this case, the above constraints may apply to the variables element output period and display element period. That is, the display element period may be an integer multiple of the element output period.

[0219] According to another embodiment, the video bitstream 14 is a multi-layer video data stream 14, and accordingly the PT SEI messages 83 refer to all output layers of the multi-layer video bitstream 14. The FFI SEI messages relate to a given output layer, and therefore to pictures of a given output layer, and to a given access unit 22. * The encoder 10 encodes the pictures of a given output layer into the encoded access units 22. * The PT multiplication syntax element 74 and the FF multiplication indicator 84 are configured to be encoded such that the FF multiplication indicator 84 is x times the PT multiplication indicator 74, where x is the distance between the access unit 22 and the preceding or succeeding access unit.

[0220] That is, alternatively, the pt_display_elemental_periods_minus1 in the PT SEI message applied to the AU and the display_elemental_periods_minus1 in the frame field information SEI message applied to each picture of the layer present in the AU may not be the same, but there are bitstream constraints as follows:

[0221] For each layer, let picA and picB be two consecutive output pictures, and let AuA and AuB be the nth and mth output AUs in the output order. Then, the value of display_elemental_periods_minus1 = ((mn) * (pt_display_elemental_periods_minus1+1))-1.

[0222] 2.4. Derivation of frame rate periodicity According to an embodiment of this sub-aspect, a constant output frame rate across the CVS is derived and not signaled, as this would otherwise require rewriting the SPS after splicing. In other words, instead of signaling whether the output frame rate across the CVS is constant or not, this information may be derived, for example, by decoder 50.

[0223] That is, in another embodiment, the above fourth problem is solved as follows: Instead of informing whether the constant frame rate (constant picture rate) property is preserved after a splicing point, this property is derived as follows.

[0224] For the CVS containing picture n, if Htid is equal to i, fixed_pic_rate_generalwithin_cvs_flag[i] is equal to 1, and picture n is an output picture that is not the last picture (in output order) in the output bitstream, the calculated value for DpbOutputElementalInterval[n] is ClockTick * (elemental_duration_in_tc_minus1[i]+1), where ClockTick is the output order nextPicInOutputOrder specified for use in equation C.16, as specified in equation C.1 (using the value of ClockTick of the CVS containing picture n) when any of the following conditions are true:

[0225] - Picture nextPicInOutputOrder is in the same CVS as picture n.

[0226] - The picture nextPicInOutputOrder is in another CVS, and in the CVS that contains the picture nextPicInOutputOrder, fixed_pic_rate_generalwithin_cvs_flag[i] is equal to 1, the value of ClockTick is the same in both CVSs, and the value of elemental_duration_in_tc_minus1[i] is the same in both CVSs. One or more of the following conditions are true:

[0227] - The GOP size is the same -DPB parameters are the same -The sorting parameters in the DPB parameters are the same -nextPicInOutputOrder is not a noOutput picture -nextPicInOutputOrder has no RASL picture associated with it (it is a CRA) -Additional syntax element indicating the output delay of the first AU of the CVS (explained in the following manner in Fehler! Verweisquelle konnte nicht gefunden werden..) -nextPicInOutputOrder has NoOutputOfPriorPicsFlag set equal to ph_no_output_of_prior_pics_flag in the picture header of nextPicInOutputOrder equal to 0. Note that this parameter indicates that at a CVS boundary, previous pictures that are still in the DPB of the previous CVS will not be output.

[0228] For the CVS containing picture n, if Htid is equal to i, fixed_pic_rate_within_cvs_flag[i] is equal to 1, and picture n is an output picture that is not the last picture (in output order) in the output CVS, the calculated value for DpbOutputElementalInterval[n] is ClockTick * (elemental_duration_in_tc_minus1[i]+1), where ClockTick is as specified in equation C.1 (using the value of ClockTick in the CVS containing picture n) when the next picture in the output order nextPicInOutputOrder specified for use in equation C.16 is in the same CVS as picture n.

[0229] The size of a group of pictures (GOP), the DPB parameters, and aspects related to reordering are shown in the following Figure 12. The first number under the picture 26 corresponds to the decode time, and the second number corresponds to the output time (i.e., the numbers under the picture are given as "decode time - output time"). Thus, the difference between the output time and the decode time varies depending on the GOP size. This may be part of the DPB parameters or reordering information added to the bitstream. For example, for GOP4, the value is 2, and for GOP8, the value is 3.

[0230] Thus, according to one embodiment of this subaspect, video data stream 14 is a concatenation of coded video sequences, and encoder 10 is configured to encode, for each coded video sequence 20 of video data stream 14, a parameter set comprising: a fixed picture rate flag indicating whether a picture output includes a fixed picture rate within said respective coded video sequence, and an elementary output picture duration syntax element (e.g., elemental_duration_in_tc_minus1[i]) if said fixed picture rate flag indicates that said picture output includes a fixed picture rate within said respective coded video sequence. According to this embodiment, encoder 10 is configured to signal in said data stream via one or more continuity detectability syntax elements that picture rate continuity is detectable when transitioning from a first coded video sequence to a second coded video sequence if all of the following apply:

[0231] the fixed picture rate flags of the first coded video sequence and the second coded video sequence indicate picture outputs of the output layer sets set as including a fixed picture rate in the first coded video sequence and the second coded video sequence; the element output picture duration syntax element of the first coded video sequence and the second coded video sequence are the same; - one or more conditions of a set of conditions apply, and the set of conditions is The GOP sizes of the first coded video sequence and the second coded video sequence are the same; the first coded video sequence and the second coded video sequence match in a reordering syntax element (e.g., max_num_reorder_pics) that indicates the maximum allowed number of output pictures that precede another output picture in decoding order and follow another output picture in output order; the DPB parameters (e.g., indicating DPB picture removal times) of the first coded video sequence and the second coded video sequence match; the second coded video sequence does not start with an IRAP associated with a RASL picture (e.g. does not start with a CRA), the first coded video sequence and the second coded video sequence match in the output delay syntax element signaled in the video data stream (14), which indicates the output delay of the first access unit of the first coded video sequence and the second coded video sequence (e.g. the first AU of CVS1 and the first AU of CVS2 have the same output delay relative to the decoding time, e.g. in picture time there is a syntax element =3 indicating a time delay of 3 pictures from decoding to output); The first AU of the second coded video sequence is not a non-output picture; The first AU of the second coded video sequence does not indicate that the previous picture from the first coded video sequence is not to be output.

[0232] Issues regarding no-output pictures are explained in detail in section 2.6 of this document. The aspects related to RASL pictures relate to no-output pictures, since RASL pictures are not output when associated with the first AU of a CVS and can therefore be considered no-output pictures.

[0233] 2.5. Derivation of Output Time, e.g., when PT SEI message is not present or is not used The embodiment according to this sub-aspect can solve the fifth problem above.

[0234] According to a first embodiment of sub-aspect 2.5, the decoder 50 is configured to decode, for a given coded video sequence of a multi-layer video data stream, a parameter set that includes: a fixed picture rate flag indicating whether the picture output includes a fixed picture rate within the given coded video sequence; and, if the fixed picture rate flag indicates that the picture output includes a fixed picture rate within the given coded video sequence, an constituent output picture duration syntax element. According to this embodiment, the decoder 50 is configured to determine the output delay (picture output time of a first picture) of the given coded video sequence based on a product having a first coefficient determined by the constituent output picture duration syntax element and a second coefficient determined using a reordering syntax element of the DPB parameters, the product indicating the maximum allowable number of output picture sets that can precede in decoding order and follow in output order any picture in the OLS. Alternatively, the decoder 50 is configured to determine the output delay (picture output time of the first picture) of a given coded video sequence based on a product of a first coefficient determined by a component output picture duration syntax element in the video data stream and a second coefficient indicated by a delay syntax element.

[0235] According to a first embodiment, an encoder 10 is provided that is configured to encode, for a given coded video sequence of a multi-layer video data stream, a parameter set (e.g., a VPS parameter set or an SPS parameter set including HRD and timing information) into a video data stream (14) that includes a fixed picture rate flag (e.g., fixed_pic_rate_within_cvs_flag) indicating whether the picture output includes a fixed picture rate within the given coded video sequence, and, if the fixed picture rate flag indicates that the picture output includes a fixed picture rate within the given coded video sequence, an elemental output picture duration syntax element (e.g., elemental_duration_in_tc_minus1). According to this embodiment, the encoder 10 is configured to signal in the data stream via one or more output delay computable syntax elements (e.g., indicating that PT / BP-...SEI is not currently required or is not included in the video data stream). The output delay (picture output time of the first picture) of a given coded video sequence is calculable based on a product having a first coefficient determined by the constituent output picture duration syntax element and a second coefficient determined using a reordering syntax element of the DPB parameters, and indicating the maximum allowable number of output picture sets that can precede any picture in the OLS in decoding order and follow it in output order. According to this embodiment, the encoder 10 is configured to signal in the data stream, via one or more output delay calculable syntax elements (e.g., ones indicating that a PT / BP-...SEI is not currently required or is not included in the video data stream), that the output delay (output time of the first picture) for a given coded video sequence is calculable based on a product having a first coefficient determined by the constituent output picture duration syntax element and a second coefficient indicated by a delay syntax element in the video data stream (14).

[0236] In other words, the first embodiment according to sub-aspect 2.5 is configured to indicate in the bitstream that timing information can be derived without the PT SEI and BP SEI, and to derive the output time of the first AU (e.g., the first AU of CVS20) according to DPB parameters or additional parameters. Note that the buffering period (BP) SEI message and the picture timing (PT) SEI message contain timing information such as when to remove an AU from the CPB and when to output an AU from the DPB. There are several values ​​(e.g., the highest various temporal IDs present in the bitstream) that can be used to derive when to decode (remove from) an AU and when to output (from the DPB). The output time can be derived without the help of these SEI messages under several conditions, which will be described below. The output time can be derived as one of the following two options:

[0237] If the -DPB parameter is used, the output time value is the ClockTick * (elemental_duration_in_tc_minus1[i]+1) * Derived as NumPics, where NumPics is the number of reordered pictures signaled in the DPB parameter (max_num_reorder_pics), or the maximum allowed number, the number of pictures in the OLS that can precede any picture in the OLS in decode order and follow that picture in output order, and the maximum number of pictures in the OLS that can precede any picture in the OLS in output order and follow that picture when decoded (max_num_reorder_pics + max_latency_increase_plus1), or - Additional signaling related to a fixed picture rate, indicating a specified number of pictures, NumPics, is added to the VPS or SPS, and this syntax is used to signal the ClockTick * (elemental_duration_in_tc_minus1[i]+1) * The output time value, derived as NumPics, is calculated.

[0238] Figure 13 shows an example of an encoder 10, a video bitstream 14, and a decoder 50 according to a second embodiment of sub-aspect 2.5. The encoder 10, the video bitstream 14, and the decoder 50 may optionally correspond to the encoder 10, the video bitstream 14, and the decoder 50 according to Figure 1. An embodiment according to this sub-aspect may also optionally include features and details described with respect to sub-aspect 2.3, for example with respect to Figure 11.

[0239] According to the embodiment of Figure 13, encoder 10 encodes a parameter set 93 for a given coded video sequence 20 into video bitstream 14. That is, parameter set 93 is associated with one of one or more coded video sequences 20 of video bitstream 14. Parameter set 93 includes a fixed picture rate flag 94, which may correspond, for example, to fixed_pic_rate_within_CVS_flag described herein. Fixed picture rate flag 94 indicates whether the picture output includes a fixed picture rate within the given coded video sequence 20. If fixed picture rate flag 93 indicates that the picture output includes a fixed picture rate within the given coded video sequence 20, parameter set 93 further includes an elemental output picture duration syntax element 96, which may correspond, for example, to elemental_duration_in_tc_minus1 described herein. For example, element output picture duration syntax element 96 may indicate the duration of a picture output interval, eg, the duration of picture output interval 63 for a single picture.

[0240] 13, each of the access units 22 of the video bitstream 14 may be associated with an element picture output time 36. The element output picture time 36 may represent the time instance at which the picture of the respective access unit 22 is output by the decoder 50, e.g., the time instance at which the decoder 50 provides the respective access unit (i.e., its picture) to an output buffer. The access unit output interval 37 may indicate the time interval between the element picture output times 36 of consecutive access units and may correspond, e.g., to the access unit output interval 61 described with respect to FIGS. 8 and 10.

[0241] According to the embodiment of Figure 13, encoder 10 encodes one or more syntax elements 66 into video bitstream 14, and decoder 50 can infer that, if one or more syntax elements 66 have a first state, an access unit of coded video sequence 20 referenced by the one or more syntax elements 66 has the first state and pictures of the access unit are not subject to multiplication output (e.g., inferring that pt_display_elemental_periods_minus1 is 0 or elementalOutputs is 1). As a result, decoder 50 can derive element output picture time 36 using element output syntax element 96, for example, by setting access unit output interval 37, 61 equal to the value of the picture output interval indicated by element output picture duration syntax element 96. As a result, decoder 50 may determine element output picture number 36 in the absence of an indication regarding the repetition number, e.g., in the absence of a PT SEI, which may result in the element output picture number being omitted in video bitstream 14.

[0242] For example, decoder 50 may generally derive element output time 36 based on element output picture duration indicated by element output picture duration syntax element 96 and the number of picture repetitions of the respective access unit, e.g., indicated by a picture output multiplication syntax element or a further picture output multiplication syntax element (see Section 2.3). To this end, decoder 50 sets the picture output interval, e.g., the duration of picture output interval 63, represented by variable DpbOutputElementalInterval, equal to the value indicated by elemental_duration_in_tc_minus1 (the indicated value may correspond to the value actually written to the bitstream plus 1, optionally scaled by a clock tick duration that may be signaled in video bitstream 14, e.g., by syntax element ClockTick, e.g., DpbOutputElementalInterval=ClockTick * (elemental_duration_in_tc_minus1 + 1), DpbOutputElementalInterval (e.g., equation (113) in the definition of elemental_duration_in_tc_minus1 in the introduction to Section 2, where the variable elementalOutputs represents the number of repetitions), and the number of repetitions of each picture, which is inferred to be 1 when syntax element 66 has the first state. As a result, when one or more syntax elements 66 have the first state, decoder 50 may derive the variable DpbOutputInterval (e.g., the duration of access unit output interval 61) using elemental_duration_in_tc_minus1 (e.g., DpbOutputInterval = ClockTick *In other words, in this case, when one or more syntax elements 66 have a first state, the decoder 50 can interpret the elementary picture duration syntax element 96 as referring to the duration of the access unit output interval 37, 61.

[0243] For example, parameter set 93, including fixed picture rate flag 94 and optional component output picture duration syntax element 96, may be a sequence parameter set (SPS) that may be globally associated with coded video sequence 20, where the SPS includes HRD and timing information. Alternatively, parameter set 93 may be a video parameter set (VPS) that may be globally associated with video bitstream 14.

[0244] According to the embodiment of Figure 13, encoder 10 encodes one or more syntax elements 66 into video bitstream 14. When one or more syntax elements 66 have a first state, encoder 10 encodes video bitstream 14 in a manner that, for each access unit 22 of coded video sequence 20 (or video bitstream 14), can be estimated such that the respective access unit 22 is not subject to multiplication output. For example, the variable element output duration described in the introduction of Section 2 and also described with respect to Section 2.3 can be estimated to be 1. Thus, element picture output times 36 of coded video sequence 20 can be determined based on the element output picture duration syntax elements.

[0245] In other words, for example, if one or more syntax elements 66 have a first state, decoder 50 can infer that access unit 22 does not receive a multiplied output, and as a result, can determine the element picture output time of a given access unit by adding the element output picture duration signaled by element output picture duration syntax element 96 to the element picture output time 36 of the preceding access unit of the given access unit.

[0246] For example, when one or more syntax elements 66 have the first state, encoder 10 may provide any picture output multiplication syntax element, such as PT multiplication indicator 74 (or pt_display_elemental_periods_minus1), to signal a single picture output, i.e., no multiplied picture output, if provided. Thus, decoder 50 may infer that when one or more syntax elements 66 have the first state, the picture output multiplication syntax element indicates a non-multiplied output, i.e., a single output.

[0247] According to an embodiment, one or more syntax elements 66 indicate one or more of the bitstream portions of the video bitstream 14, e.g., HRD parameters (or bitstream adaptation parameters) referencing general_nal_hrd_params_present_flag=0, and the coding layer of the video bitstream 14, e.g., video bitstream 14 including (or not including) HRD parameters (or bitstream adaptation parameters) referencing general_vcl_hrd_params_present_flag=0.

[0248] According to an embodiment, the one or more syntax elements 66 may include one or more of a first syntax element and a second syntax element, each of which indicates that the video bitstream 14 does not include (or includes) coded picture buffer (CPB) and bitrate parameters for a respective operation mode of the hypothetical reference decoder, such as NAL operation (e.g., an operation mode that may include SEI NAL units and headers over VCL data) and VCL operation (e.g., an operation mode that may include SEI NAL units and headers over VCL data). For example, the one or more syntax elements 66 may include one or both of a general_nal_hrd_params_present_flag and a general_vcl_hrd_params_present_flag.

[0249] For example, the first state may be a state in which one or more syntax elements indicate that the video bitstream 14 does not include coded picture buffer (CPB) and bitrate parameters for both the NAL and VCL operating modes of the hypothetical reference decoder, e.g., general_nal_hrd_params_present_flag=0 and general_vcl_hrd_params_present_flag=0.

[0250] For example, one or more syntax elements 66 may be encoded into one or more parameter sets in video bitstream 14.

[0251] According to the example embodiment of Figure 13, when one or more syntax elements 66 have a second state, e.g., when one or more syntax elements 66 indicate that video bitstream 14 includes one or more of the parameters described above, encoder 10 may encode, for each access unit 22 of video bitstream 14 or coded video sequence 20, a picture output multiplication syntax element, e.g., picture output multiplication syntax element 74 as described with reference to Figure 11, in a PT SEI message 73 of video bitstream 14. As described with reference to Figure 11, picture output multiplication syntax element 74 may reveal information about whether each access unit 22, i.e., the access unit referenced by PT SEI message 73, undergoes multiplication and, if so, how many consecutive output pictures are generated from each access unit 22. According to this example, element picture output time 36 is determinable based on element output picture duration syntax element 96 and the picture output multiplication syntax element. For example, decoder 50 may decode picture output multiplication syntax elements for each access unit when one or more syntax elements 66 have the second state, and determine elementary picture output times 36 for the access units of coded video sequence 20 based on elementary output picture duration syntax element 96 and the picture output multiplication syntax elements. For example, decoder 50 may multiply, for each access unit 22, the output duration indicated by the elementary output picture duration syntax element by the number of iterations indicated by the picture output multiplication syntax element, e.g., using variable elementalOutputs and equation (113) described above, as described above for the case when the first syntax element has the first case, but the number of iterations is not presumed to be 1.

[0252] In other words, in further embodiments, such as that of FIG. 13, there is an indication in the bitstream, e.g., in the VPS or SPS, that the PT SEI and BP SEI are not present and / or not required, and there is also information that the frame field SEI message is not present or not required, so there is no need to include or consider repetitions, i.e., the decoder can derive elementalOutputPeriods=1. This can be done using the syntax element described above, which indicates that timing can be derived without the PT SEI and BP SEI messages, or alternatively, using bitstream constraints. Note that in this case, if the PT SEI message is not present or if the syntax element is not present, pt_display_elemental_periods_minus1 can be inferred to be 0.

[0253] Bitstream constraints can be added as no_timing_information_sei_message_needed_flag, or the described behavior can be adjusted, for example, when general_nal_hrd_params_present_flag and general_vcl_hrd_params_present_flag, which indicate the presence of CPB and bitrate parameters for NAL or VCL operation, are both 0. In the latter case, no PT SEI or BP SEI messages are needed, and it is possible to simply operate without them by deriving elementalOutputPeriods to 1 and using elemental_duration_in_tc_minus1[I] for the output picture rate to derive the output duration.

[0254] Thus, sub-aspect 2.5 provides a concept for determining element picture output times 36 without the presence of PT SEI messages. As a result, PT SEI messages do not necessarily need to be encoded into video bitstream 14, thus avoiding signaling overhead in video bitstream 14.

[0255] 2.6. Handling of silent pictures and their impact on constant output frame rate As mentioned above, deriving the output time can be complicated if some pictures are not output, are not present in the PT SEI message, or are ignored.

[0256] If a PT SEI message is present, a noOutput picture has an associated output time, but since the picture is not output, such output time is simply ignored. If we count pictures decoded in the bitstream as "occupying" an output time slot, the distance between two output pictures that are actually output will not be equidistant, as shown in Figure 14 below. For example, in the example in Figure 14, picture 26 * is indicated as a noOutput picture, for example, depending on the layer it belongs to. In other words, Figure 8 shows an example of a non-equidistant picture in the case of a no-output picture.

[0257] According to a first embodiment of this sub-aspect, the decoder 50 is configured to decode a parameter set including a fixed picture rate flag, e.g., as in section 2.5, indicating whether the picture output for the video data stream includes a fixed picture rate, and if the fixed picture rate flag indicates that the picture output includes a fixed picture rate, decode a component output picture duration syntax element. According to this embodiment, the decoder 50 decodes, from the video data stream 14, a picture output flag indicating, for each picture, whether the respective picture should be displayed or not. According to this embodiment, the decoder 50 decodes a further picture that has been displayed to not be output, e.g., picture 26 in FIG. 14 .* It is assumed that the picture preceding this picture in output order, for example, picture 26' in FIG. 14, is the target of repeated output.

[0258] In other words, in one embodiment, for example in the example in the previous section, the flag indicating whether a picture is output, i.e. (ph_pic_output_flag), is taken into account for deriving the output time. That is, from the decoder, we are ready to receive pictures that do not need to be displayed as a constant output. In such cases, a constant display rate is achieved, there is a bitstream constraint, and previous pictures in the output order need to be compensated for by repetition.

[0259] According to a second embodiment, the encoder 10 is configured to encode into the video data stream 14 a parameter set including a fixed picture rate flag and, if the fixed picture rate flag indicates that the picture output is with a fixed picture rate, an elementary output picture duration syntax element. According to this embodiment, the encoder 10 is configured to encode into the video data stream 14, for each picture, a picture output flag indicating whether the respective picture should be displayed. According to this embodiment, if the fixed picture rate flag indicates that the picture output of the video data stream 14 includes a fixed picture rate, the encoder 10 sets the picture output flag of each picture to indicate that the respective picture should be displayed. Alternatively or additionally, if the fixed picture rate flag indicates that the picture output for the video data stream 14 is with a fixed picture rate, the encoder 10 sets the picture output flag of each picture that is not the first picture of a coded video sequence of the video data stream 14 to indicate that the respective picture should be displayed. Alternatively or additionally, if the fixed picture rate flag indicates that picture output for video data stream 14 will be with a fixed picture rate, encoder 10 sets a picture output flag as an indication to display the respective picture for each picture that is not the first picture in a coded video sequence of video data stream 14 or that is not exclusively preceded by other no-output pictures in a coded video sequence of video data stream 14. Thus, for example, encoder 10 provides video bitstream 14 such that a decoder can infer that a picture that precedes in output order a further picture indicated not to be output will be subject to repeated output.

[0260] In an example, encoder 10 is configured to encode into video data stream 14 a flag indicating whether the picture output flag of each picture is set to indicate that the respective picture is to be displayed. Alternatively or additionally, the flag indicates whether the picture output flag of each picture that is not the first picture of a coded video sequence of video data stream 14 is set to indicate that the respective picture should be displayed. Alternatively or additionally, the flag indicates whether the picture output flag is set to indicate that the respective picture is to be displayed for each picture that is not the first picture of a coded video sequence of video data stream 14 or a picture in a coded video sequence of video data stream 14 that is not exclusively preceded by another no-output picture.

[0261] In an example, encoder 10 is configured to set a fixed picture rate flag if the flag indicates that the picture output of video data stream 14 includes a fixed picture rate.

[0262] In other words, in another embodiment, for example the second embodiment, if a fixed picture rate is indicated in the bitstream, there is a bitstream constraint that prohibits it.

[0263] Non-output pictures in the bitstream, or At least a non-output picture that is not the first AU of the CVS, or If there is an output picture in CVS, there can be no more pictures without outputs following that output picture in CVS.

[0264] In another embodiment, a bitstream constraint is added that not only applies to a fixed picture rate, but also indicates that the output picture is unconstrained as a broader concept.

[0265] [Table 4] If general_no_no_output_pics_constraint_flag is equal to 1, then ph_pic_output_flag shall be equal to 1. If general_no_no_output_pics_constraint_flag is equal to 0, then no such constraint shall be imposed.

[0266] Furthermore, if a fixed picture rate is used, there is a bitstream constraint that a constraint flag must be set.

[0267] If fixed_pic_rate_general_flag[i] is equal to 1 for any value of i, it is a requirement for bitstream conformance that general_no_no_output_pics_constraint_flag is equal to 1.

[0268] 3. Further Embodiments In the previous section, some aspects have been described as features in the context of an apparatus, but it will be apparent that such description may also be considered as a description of corresponding features of a method. While some aspects have been described as features in the context of a method, it will be apparent that such description may also be considered as a description of corresponding features with respect to the functionality of the apparatus.

[0269] Some or all of the method steps may be performed by (or using) a hardware apparatus, such as, for example, a microprocessor, a programmable computer, or electronic circuitry. In some embodiments, one or more of the most important method steps may be performed by such an apparatus.

[0270] The inventive encoded image signal may be stored on a digital storage medium or transmitted over a transmission medium, such as a wireless or wired transmission medium, such as the Internet.

[0271] Depending on certain implementation requirements, embodiments of the present invention may be implemented in hardware or software, or at least partially in hardware or at least partially in software. This implementation may be performed using a digital storage medium, such as a floppy disk, DVD, Blu-ray, CD, ROM, PROM, EPROM, EEPROM, or FLASH memory, having electronically readable control signals that may be stored thereon and that cooperate (or may be able to cooperate) with a programmable computer system to perform the respective methods. Thus, the digital storage medium may be computer-readable.

[0272] Some embodiments according to the invention include a data carrier having electronically readable control signals that can cooperate with a programmable computer system to perform one of the methods described herein.

[0273] Generally, embodiments of the present invention may be implemented as a computer program product having program code operable to perform one of the methods when the computer program product runs on a computer, which may for example be stored on a machine-readable carrier.

[0274] Other embodiments comprise the computer program for performing one of the methods described herein, stored on a machine readable carrier.

[0275] In other words, an embodiment of the inventive methods is, therefore, a computer program having a program code for performing one of the methods described herein, when the computer program runs on a computer.

[0276] A further embodiment of the inventive method is therefore a data carrier (or digital storage medium, or computer-readable medium) comprising, recorded on it, the computer program for performing one of the methods described herein. The data carrier, digital storage medium, or recorded medium is typically tangible and / or non-transitory.

[0277] A further embodiment of the inventive method is, therefore, a data stream or a sequence of signals representing the computer program for performing one of the methods described herein, The data stream or the sequence of signals can be adapted to be transferred via a data communication connection, such as, for example, via the Internet.

[0278] A further embodiment comprises a processing means, such as for example a computer, or a programmable logic device, configured to or adapted to perform one of the methods described herein.

[0279] A further embodiment comprises a computer having installed thereon the computer program for performing one of the methods described herein.

[0280] Further embodiments according to the invention include an apparatus or system configured to transfer (e.g., electronically or optically) a computer program for performing one of the methods described herein to a receiver. The receiver may be, for example, a computer, a mobile device, a memory device, etc. The apparatus or system may, for example, include a file server for transferring the computer program to the receiver.

[0281] In some embodiments, a programmable logic device (e.g., a field programmable gate array) may be used to perform some or all of the functions of the methods described herein. In some embodiments, a field programmable gate array may cooperate with a microprocessor to perform one of the methods described herein. In general, the methods are preferably performed by any hardware apparatus.

[0282] The devices described herein may be implemented using a hardware device, or using a computer, or using a combination of a hardware device and a computer.

[0283] The methods described herein may be performed using a hardware apparatus, or using a computer, or using a combination of a hardware apparatus and a computer.

[0284] In the foregoing Detailed Description, it will be appreciated that various features are grouped together into examples for the purpose of streamlining the disclosure. This method of disclosure should not be interpreted as reflecting an intention that the claimed examples require more features than are expressly recited in each claim. Rather, as the following claims reflect, subject matter may lie in fewer than all features of a single disclosed example. Accordingly, the following claims are hereby incorporated into the Detailed Description, with each claim standing on its own as a separate example. While each claim may stand on its own as a separate example, and a dependent claim may refer to a specific combination with one or more other claims in the claim, it should be noted that other examples may include combinations of the dependent claim with the subject matter of each other dependent claim, or combinations of each feature with other dependent or independent claims. Such combinations are suggested herein unless it is stated that a particular combination is not intended. Furthermore, it is intended to include features of other independent claims, even if the claims are not directly dependent on those independent claims.

[0285] The above-described embodiments are merely illustrative of the principles of the present disclosure. It is understood that modifications and variations of the arrangements and details described herein will be apparent to those skilled in the art. It is therefore intended to be limited only by the scope of the appended claims and not by the specific details presented by way of description and illustration of the embodiments herein.

Claims

1. 1. A method for decoding a video data stream, comprising: decoding a parameter set including a first syntax element indicating whether a network abstraction layer (NAL) hypothetical reference decoder (HRD) parameter is present and a second syntax element indicating whether a video coding layer (VCL) hypothetical reference decoder (HRD) parameter is present; setting a flag indicating the absence of a picture timing supplemental extension information message applicable to one or more access units based at least in part on determining that the first syntax element indicates that a NALHRD parameter is not present and that the second syntax element indicates that a VCLHRD parameter is not present; and estimating, based at least in part on the first syntax element and the second syntax element, that one or more pictures corresponding to the one or more access units are not subject to multiplication output.

2. The method of claim 1 , further comprising: determining a picture output interval based on estimating that the one or more pictures corresponding to the one or more access units are not subject to multiplication output.

3. decoding a flag indicating whether a picture output in the coded video sequence includes a fixed picture rate; 10. The method of claim 1, further comprising: decoding an element duration syntax element based on the flag indicating that the coded video sequence has a fixed picture rate.

4. The method of claim 1 , further comprising determining a picture output interval by scaling a value of an element duration syntax element by a clock tick duration.

5. 1. A method for encoding a video data stream, comprising: encoding a parameter set including a first syntax element indicating whether a network abstraction layer (NAL) hypothetical reference decoder (HRD) parameter is present and a second syntax element indicating whether a video coding layer (VCL) hypothetical reference decoder (HRD) parameter is present; signaling, via the first syntax element and the second syntax element, that the video data stream does not include a picture timing supplemental extension information message applicable to one or more access units based at least in part on the absence of the NALHRD parameter and the absence of the VCLHRD parameter; and signaling, based at least in part on the first syntax element and the second syntax element, that one or more pictures corresponding to the one or more access units are not subject to multiplication output.

6. 6. The method of claim 5, further comprising: signaling that a picture output interval is derived based on one or more pictures corresponding to the one or more access units that are not subject to the multiplication output.

7. encoding a flag within the coded video sequence indicating whether the picture output includes a fixed picture rate; The method of claim 5 , further comprising: encoding an element duration syntax element based on the coded video sequence having a fixed picture rate.

8. The method of claim 5 , wherein the picture output interval is determined by scaling the value of an element duration syntax element by the clock tick duration.

9. A non-transitory computer readable medium storing a computer program for performing the method of any one of claims 1 to 8 when the program is run on a computer or signal processor.

10. A coding device for decoding or encoding a video data stream, said coding device comprising one or more hardware processors configured to perform the method of any one of claims 1 to 8.

Citation Information

Patent Citations

  • Low-latency video encoding system and its operation method

    JP2015526970A

  • electronic device for signaling subpicture buffer parameters

    JP2016500204A