Various signaling concepts for multi-layer video bitstreams and for output timing derivation

By disabling RASL images as interlayer references in the encoder and handling them with decode refresh and sequence end identifiers, the problem of RASL images being undecoded in multi-layer video bitstreams is solved, ensuring correct image output and decoding order, and improving the reliability and efficiency of the decoder.

CN116171575BActive Publication Date: 2025-11-21FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202180056306.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-06-10
Filing Date
2021-06-09
Publication Date
2025-11-21
Estimated Expiration
2041-06-09

AI Technical Summary

Technical Problem

In multi-layer video bitstreams, independently coded images (RASL images) may lack a reference image, resulting in incorrect decoding. This, in turn, affects other images that depend on that image, making it difficult for existing technologies to effectively handle this situation.

Method used

By disabling RASL images as interlayer references in the encoder and using decoding refresh and sequence end identifiers, we ensure that RASL images are not output or used as interlayer references. Combined with the image output timing signaling mechanism, we ensure the accuracy and reliability of image output timing.

Benefits of technology

It effectively prevents decoding errors caused by the inability to decode RASL images, ensures the correct output and decoding order of images in multi-layer video bitstreams, and improves the reliability and efficiency of the decoder.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116171575B_ABST
    Figure CN116171575B_ABST
Patent Text Reader

Abstract

The first aspect provides concepts for handling coded layer video sequence boundaries in multi-layer video bitstreams with inter-layer references. The second aspect provides concepts for handling, signaling and derivation of picture output timing, e.g. access unit specific signaling and output layer specific signaling related to the number of repetitions of a picture output.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of this disclosure relate to video encoders, video decoders, methods for encoding video sequences into video bitstreams, and methods for decoding video sequences from video bitstreams. Further embodiments relate to video bitstreams. Background Technology

[0002] Video can be encoded into a video bitstream in units of one or more coded video sequences. Each coded video sequence includes a sequence of access units, each access unit comprising one or more pictures of a common time frame of the video. In the case of multi-layer video bitstreams, where video data is encoded into multiple layers of the video bitstream, these layers can include individual coded layer video sequences. Different layers of coded layer video sequences do not necessarily need to start / stop with the same access unit. A coded layer video sequence can begin with an independent coded picture, such as an IRAP picture, which can be decoded independently of access units that are different from those of the access units that depend on the coded picture. There can be a first type of independent coded picture, such as a CRA picture that can be associated with a picture in the same layer, which is after the first type of picture in the encoding order but is presented before the first type of picture. Such a picture can be called a RASL picture. Such a RASL picture can have a reference to a picture that precedes the first type of picture associated with the RASL picture in the decoding order. In other words, a RASL picture can have a reference to pictures in a previous coded layer video sequence. Therefore, there may be situations where the reference image for the RASL image is not present in the video bitstream, making it impossible to decode the RASL image correctly. In such cases, there might be instructions to exclude images preceding the first type of image from the output, based on their presentation order. However, the RASL image can be used as an inter-layer reference image for images at different layers. If the RASL image is undecodeable, images that depend on it may also fail to decode correctly. Summary of the Invention

[0003] A first aspect of this disclosure provides a concept for handling the boundaries of coded layer video sequences in a multi-layer video bitstream. Embodiments of the first aspect prevent images indicating inter-layer references to RASL images that are undecodeable or excluded from the output from being used in the output. Therefore, such images having inter-layer references to RASL images that cannot be correctly decoded are not present in the video bitstream or are excluded from the output.

[0004] According to an embodiment of the first aspect, when encoding the first and second layers into a multi-layer video bitstream such that the first layer depends on the second layer, an encoder encodes within the next access unit in a subsequent access unit. This next access unit follows the access unit in the second layer that includes a sequence end identifier in the encoding order. This next access unit is the access unit closest to the access unit in the second layer that has a sequence end identifier. The image in the first layer is encoded into this next access unit. The image to be encoded into the first layer is decoded and refreshed without any leading image output. Since the first image of the next access unit is encoded using a decoded refresh and without a leading image output, for example, an image with an interlayer reference to a RASL image of the second layer also needs to be a RASL image, and therefore output can be prevented by indicating that no leading image is output. Indicating that no leading image is output can prevent those images, which also indirectly depend on the images that are part of the encoded layer video sequence of the second layer, which is indicated to end at the sequence end identifier of the second layer, from being output.

[0005] The second aspect of this disclosure relates to the output timing of a decoded image, that is, the output time of the decoded image from the decoder's output buffer. The derivation of the image output timing can be, for example, signaled at the access unit level, such as by supplementing the image timing with enhancement information PTSEI. Additionally or alternatively, the image output timing can be signaled at the output layer level, i.e., the individual output layer referencing the video bitstream. For example, information regarding the image output timing for access units and / or for output layers may include information about the number of times the image is to be output (i.e., repeated).

[0006] According to the first sub-aspect of the second aspect, a gating flag is provided in the video bitstream, which signals whether the PT SEI message included in the video bitstream includes a picture output multiplication syntax element. The picture output multiplication syntax element indicates whether the picture of the access unit it references undergoes multiplication picture output, and if so, how many output pictures to generate from the picture of the access unit. The gating flag provides a way to distinguish whether information about multiple picture outputs is retrieved from the PT SEI message or by other means (e.g., by referencing the frame field SEI message of an individual output layer). Therefore, signaling the gating flag allows signaling at different frame rates in different output layers of the output layer set. In other words, the gating flag allows signaling of multiplication pictures output separately for different pictures within an access unit.

[0007] A second sub-aspect of the second aspect provides a concept for combining a frame-field syntax element to use a multiplicative image output of an access unit, the frame-field syntax element indicating where the images of the image sequence represent a field or frame, such as interlaced or progressive scan images. Therefore, embodiments of the second sub-aspect allow for the signaling of image output timing when using frames or fields.

[0008] A third sub-aspect of the second aspect provides a concept for signaling the number of image outputs via a picture output multiplication syntax element of the PT SEI referring to an access unit of the video bitstream and a further picture output multiplication syntax element of the frame field SEI referencing the output layer of the video bitstream. According to one embodiment, the picture output multiplication syntax element is equal to or less than the further picture output multiplication syntax element; for example, the further picture output multiplication syntax element is an integer multiple of the picture output multiplication syntax element. The picture output multiplication syntax element included in the PT SEI message can be access unit specific and thus allows for the determination of the picture refresh interval for the output implemented as a result of repetition or multiplication indicated by the picture output multiplication syntax element. Therefore, with the aid of the picture output multiplication syntax element, timing information of the picture output associated with the access unit can be determined, and in the case of multiplication, the spacing between the pictures presented therein can be determined. The further picture output multiplication syntax element signaled in the frame field SEI message can provide layer-specific information about the frequency at which the pictures need to be repeated in order to present content at the picture refresh interval (i.e., the refresh interval determined from the picture output multiplication syntax element). Requires that the further image output multiplication syntax element be equal to or greater than, for example, an integer multiple of the image output multiplication syntax element. This ensures that the image refresh interval signaled by the image output multiplication syntax element is achieved through the number of image outputs signaled by the further image output multiplication syntax element. For example, in the cases of the first and second scenes, the further image output multiplication syntax element can signal a multiplication value corresponding to twice the multiplication value signaled by the image output multiplication syntax element.

[0009] The fourth sub-aspect of the second aspect provides a concept for deriving from the video bitstream whether the output frame rate is constant outside the boundaries between subsequent coded video sequences, for example, without explicitly signaling this information in the video bitstream. Inferring this information instead of explicit signaling offers the advantage that the corresponding information does not necessarily have to be modified or checked to ensure it is still two when splicing video bitstreams.

[0010] The fifth sub-aspect of the second aspect is for deriving the element image output time (e.g., the output time of an access unit) of an encoded video sequence based on a pixel output image duration syntax element (e.g., elemental_duration_in_tc_minus1), which can be part of a parameter set, such as a video parameter set or sequence parameter set with HRD and timing information. This concept relies on the idea that if one or more syntax elements encoded into the video bitstream have a first state, it is guaranteed that the access unit can be inferred not to undergo multiplication output. Therefore, given information that can be used to infer whether the image of an access unit undergoes multiplication output, this concept allows determining the pixel image output time without needing to signal the PT SEI message containing that information. For example, in this case, the element output image time can be derived based on information about the output time of an individual image, as it can be provided, for example, by the pixel output image duration syntax element. Therefore, this concept allows deriving the pixel image output time in the absence of a PT SEI message and / or allows omitting the signaling of the PT SEI message.

[0011] The sixth sub-aspect of the second aspect provides a concept for handling situations where there are no output pictures in a video bitstream transmitted at a fixed picture rate. According to the sixth sub-aspect, if a fixed picture rate is indicated for the video bitstream, the pictures preceding the point where no output picture is received are repeated; that is, the pictures indicated to be ignored from the output are repeated. Therefore, a fixed picture rate can be maintained even when there are no output pictures. Attached Figure Description

[0012] The embodiments and preferred implementations of this disclosure are described in more detail below with reference to the accompanying drawings, wherein:

[0013] Figure 1 The diagram illustrates the encoder, decoder, and video bitstream according to an embodiment.

[0014] Figure 2 The illustration shows an example of a two-layer bitstream with different IRAP periods.

[0015] Figure 3 The illustration shows an example of randomly accessing two layers of video bitstreams without an end-of-sequence indication.

[0016] Figure 4 An example of an encoded video sequence with an aligned sequence end indication according to an embodiment of the first aspect is illustrated.

[0017] Figure 5 An example of a dependency layer according to an embodiment of the first aspect is illustrated.

[0018] Figure 6An example of a time sublayer is illustrated.

[0019] Figure 7 The illustration shows an example of bitstream splicing.

[0020] Figure 8 The illustration shows an example of frame repetition.

[0021] Figure 9 The illustration shows an example of two-layer bitstreams with different frame rates.

[0022] Figure 10 The illustration shows an example of a two-layer bitstream with repeating output frames in one layer.

[0023] Figure 11 The encoder and video bitstream according to an embodiment of sub-aspect 2.3 are illustrated.

[0024] Figure 12 The illustration shows an example of GOP size, DPB parameters, and reordering.

[0025] Figure 13 Examples of encoders, decoders, and video bitstreams according to an embodiment of sub-aspect 2.5 are illustrated.

[0026] Figure 14 The illustration shows an example of a bitstream that includes images that are not output. Detailed Implementation

[0027] In the following description, embodiments are discussed in detail; however, it should be understood that the embodiments provide many applicable concepts that can be embodied in a wide variety of video coding concepts. The specific embodiments discussed are merely illustrative of specific ways of implementing and using the concepts and do not limit the scope of the embodiments. In the following description, numerous details are set forth to provide a more thorough explanation of embodiments of the present disclosure. However, it will be apparent to those skilled in the art that other embodiments can be practiced without these specific details. In other instances, well-known structures and devices are shown in block diagram form without detail to avoid obscuring the examples described herein. Furthermore, unless otherwise specifically indicated, features of the different embodiments described herein can be combined with each other.

[0028] In the following description of the embodiments, identical or similar elements or elements having the same functionality are provided with the same reference numerals or are identified by the same name, and repeated descriptions of elements provided with the same reference numerals or are identified by the same name are generally omitted. Therefore, the descriptions provided for elements having the same or similar reference numerals or being identified by the same name can be interchanged or applied to each other in different embodiments.

[0029] The detailed description of embodiments of the disclosed concepts begins with a description of examples of encoders, decoders, and video bitstreams, which provide a framework in which embodiments of the invention can be embedded. Hereinafter, the description of embodiments of the concepts of the invention, along with instructions on how to construct such concepts... Figure 1 The descriptions in the encoder and decoder are presented together. Although, information about subsequent... Figure 2 The described embodiments are used to form unbased on the embodiments regarding Figure 1 The described framework is used to operate the encoder and decoder. Further note that the encoder and decoder can be implemented separately from each other, but they are integrated... Figure 1 The terms are described together for illustrative purposes. It should be further noted that the encoder and decoder can be combined within a single device, or one of them can be implemented as part of the other. Additionally, see references... Figure 1 Some embodiments of the present invention are described.

[0030] 0. According to Figure 1 The encoder 10, decoder 50, and video bitstream 14

[0031] Figure 1 An example of encoder 10 and decoder 50 is illustrated. Encoder 10 (which may also be referred to as an encoding device) encodes video sequence 12 into video bitstream 14 (which may also be referred to as a bitstream, data stream, video data stream, or stream). Video sequence 12 includes a sequence of pictures 13, which are arranged in a presentation order or picture order 17. In other words, each of the pictures 13 may represent a frame of video sequence 12 and may be associated with a temporal instant in the presentation order of video sequence 12. Based on video sequence 12, encoder 10 may encode encoded video sequence 20 into video bitstream 14. Encoder 10 may form encoded video sequence 20 in the form of access units 22, each of which has encoded video data belonging to a common temporal instant. In other words, each access unit 22 may have encoded one of the frames of video sequence 12 into it. Encoder 10 encodes encoded video sequence 20 into video bitstream 14 according to encoding order 19, which may be different from the picture order 17 of video sequence 12.

[0032] Encoder 10 can encode the encoded video sequence 20 into one or more layers. That is, the video bitstream 14 can be a single-layer or a multi-layer video bitstream including one or more layers. Each access unit 22 includes one or more encoded pictures 26 (e.g., Figure 1 Images 260 and 261 in the diagram are used to refer to specific images (where apostrophes and asterisks are used to refer to specific images, and subscript indices indicate the layer to which the image belongs). Note that the encoded images will be referred to simply as images below. Each of images 26 belongs to one of the layers 24 of the encoded video sequence, for example... Figure 1Layers 240 and 241. Figure 1 The diagram illustrates an exemplary number of two layers, namely a first layer 241 and a second layer 240. In embodiments according to the disclosed concept, the encoded video sequence 20 and video bitstream 14 do not necessarily include multiple layers, but may include one, two, or more layers. Figure 1 In the example, each access unit 22 includes an encoded image 261 of the first layer 241 and an encoded image 260 of the second layer 240. However, it should be noted that each access unit 22 may (but not necessarily) include encoded images for each layer used to encode the video sequence 20. For example, layers 240, 241 may have different frame rates (or image rates) and / or may include images of complementary subsets of the access units for access unit 22.

[0033] As previously mentioned, images 260 and 261 of one of the access units represent image content at the same instant. For example, images 260 and 261 of the same access unit 22 can represent the same image content at different qualities, such as resolution or fidelity. In other words, layer 240 can represent a first version of the encoded video sequence 20, while layer 241 can represent a second version of the encoded video sequence 20. Therefore, a decoder or extractor, such as decoder 50, can select between different versions of the encoded video sequence 20 to be decoded or extracted from the video bitstream 14. For example, layer 240 can be decoded independently of other layers of the encoded video sequence to provide a first-quality decoded video sequence, while joint decoding of the first layer 241 and the second layer 240 can provide a second-quality decoded video sequence that is higher than the first quality. For example, the first layer 241 can be encoded based on the second layer 240. In other words, the second layer 240 can be a reference layer for the first layer 241. For example, in this scenario, the first layer 241 can be referred to as an enhancement layer, and the second layer 240 can be referred to as a base layer. Image 260 can have a smaller, equal, or larger image size than image 261. For example, image size can refer to the number of samples in a two-dimensional array of images. It is important to note that images 260 and 261 do not necessarily represent the same image content; rather, image 261 can represent an excerpt of the image content of image 260. For example, in some scenarios, different layers of video bitstream 14 can include different sub-images of images encoded into the video bitstream, which can be encoded independently of each other. Therefore, in another example, layers 240 and 241 can be encoded independently of each other into video bitstream 14.

[0034] Encoder 10 encodes access units 22 into bitstream portions 16 of video bitstream 14. For example, each of the access units 22 can be encoded into one or more bitstream portions 16. For example, image 26 can be subdivided into slices of tiles, and each slice can be encoded into a bitstream portion 16. The bitstream portion 16 into which image 26 is encoded can be referred to as a Video Coding Layer (VCL) NAL unit. Video bitstream 14 may also include non-VCL NAL units into which descriptive data is encoded, such as bitstream portions 23, 29. The descriptive data can provide information for decoding or information about the encoded video sequence 20. The bitstream portions into which descriptive data is encoded can be associated with individual bitstream portions, for example, they can refer to individual slices, or they can be associated with one of image 26 or one of access units 22, or they can be associated with a sequence of access units, i.e., with the encoded video sequence 20. It should be noted that video 12 can be encoded into a sequence of encoded video sequence 20.

[0035] Decoder 50 (which may also be referred to as a decoding device) decodes video bitstream 14 to obtain decoded video sequence 51. It should be noted that the video bitstream 14 provided to decoder 50 does not necessarily correspond to the video bitstream 14 provided by encoder, but may have been extracted from the video bitstream provided by encoder, such that the video bitstream decoded by decoder 50 may be a sub-bitstream of the video bitstream encoded by encoder such as encoder 10. As mentioned earlier, decoder 50 may decode the entire encoded video sequence 20 encoded into video data stream 14, or may decode a portion thereof, such as a subset of the layers of encoded video sequence 20 and / or a temporal subset of encoded video sequence 20 (i.e., a video sequence with a frame rate lower than the maximum frame rate provided by encoded video sequence 20). Therefore, decoded video sequence 51 does not necessarily correspond to video sequence 12 encoded by encoder 10. It should also be noted that decoded video sequence 51 may further differ from video sequence 12 due to encoding losses such as quantization loss. For each frame of the decoded video sequence, decoded video sequence 51 comprises one or more decoded pictures 53 decoded from the various layers of encoded video sequence 20. In other words, in this example, the decoded video sequence 51 may include one or more layers, similar to the encoded video sequence 20. The decoded image 53 may be output according to the output order 18, which in this example may correspond to the image order 17. However, the decoded video sequence 51 does not necessarily include all frames of the video sequence 12, and may also include multiple instances of an image, i.e., repetition, which will be explained in detail in Part 2.

[0036] Image 26 can be encoded using a prediction tool that predicts signals or coefficients representing images in video bitstream 14 from previously encoded images. That is, encoder 10 can use the prediction tool to encode a predetermined image 26*—e.g., an image to be encoded using previously encoded images. Correspondingly, decoder 50 can use the prediction tool to predict the image 26* to be decoded from previously decoded images. In the following description, predetermined images or blocks, such as the currently encoded image or block, will be referenced using the reference marker (*). For example, Figure 1 Image 261* in the image is considered to be the current encoded image, where the current encoded image 26* can be equivalently referred to as the current encoded image encoded by encoder 10 and the current decoded image in the decoding process performed by decoder 50.

[0037] Predicting images from other images in the encoded video sequence 20 can also be referred to as inter-frame prediction. For example, image 261* can be encoded using temporal inter-frame prediction from image 261', which belongs to a different access unit in access unit 22 than image 261*. Therefore, image 261* can include an intra-layer reference 32 for image 261', which belongs to the same layer as image 261* but to a different access unit. Additionally or alternatively, image 261* can be predicted using inter-layer (inter-layer) prediction from another layer—for example, a lower layer (which can be downgraded by means of a layer index that can be associated with each of layers 24). For example, image 261* can include an inter-layer reference 34 for image 260', which belongs to the same access unit but to a different layer. In other words, in Figure 1 In this context, images 261′ and 260′ can be examples of possible reference images used for encoding image 261*. Note that prediction can be used, for example, to predict the coefficients of the image itself when determining the transform coefficients transmitted by signals in the video bitstream 14, or it can be used to predict the syntax elements used in the encoding of the image.

[0038] The embodiments described herein can be implemented in the context of Universal Video Coding (VVC) or other video codecs.

[0039] In the following text, reference will be made to Figure 1 And about Figure 1The features described herein are used to describe several concepts and embodiments. It should be noted that the features described regarding encoders, video bitstreams, or decoders should be understood as descriptions of other entities among these entities. For example, a feature described as existing in a video data stream should be understood as a description of an encoder configured to encode that feature into the video bitstream and a decoder or extractor configured to read that feature from the video bitstream. Furthermore, inferences based on information indicating that the feature is encoded into the video bitstream can be performed similarly on the encoder and decoder sides. It should also be noted that the aspects described in the following sections can be combined with each other.

[0040] 1. Meaning of End of Sequence (EOS) in Multi-Level Bitstream

[0041] This section refers to Figure 1 The embodiments according to the first aspect are described. The details described in Part 0 may optionally be applied to the embodiments according to the first aspect.

[0042] According to the embodiment of the first aspect, the video bitstream 40 is a multi-layer video bitstream, for example, such as Figure 1As illustrated in the figure. As described in Part 0, image 26 can be encoded with inter-layer prediction such that information from another image in the same layer is needed for decoding. Alternatively or additionally, images can be encoded using inter-layer prediction such that information from another image in a different layer is needed for decoding. In contrast, independently encoded or randomly accessible images can be images that do not depend on images belonging to access units 22 that are different from their own access units. In other words, independently encoded images are encoded without using temporal inter-frame prediction. For example, intra-frame random access point (IRAP) images can be independently encoded images. Examples of IRAP images are instantaneous decode refresh (IDR) images and clean random access (CRA) images. As mentioned in Part 0, the encoding order 19 does not necessarily correspond to the image order and presentation order (also known as the output order). Dependent encoded images that depend on preceding images but are preceding in the encoding and presentation order can be referred to as trailing images. Another example of dependent encoded images is an image that depends on an image of a previously encoded access unit, but that dependent encoded image precedes the image it depends on in the presentation order 19. An example of such an image could be a Random Access Skip Preamble (RASL) image. A RASL image can be associated with independently encoded images (e.g., CRA images) that it may depend on, which are encoded before the RASL image but rendered after it. Furthermore, a RASL image can depend on one or more other images (i.e., including references to them), which include one or more images encoded before the associated independently encoded (e.g., CRA) images. In cases where an independently encoded image is the start image of a encoded video (layer) sequence (e.g., because it is the first in the bitstream or the first after a sequence end indication), this could mean that images encoded before the independently encoded image are cleared from the buffer, and therefore the RASL image might be excluded from the output because it may be incorrectly decoded due to a lack of reference.

[0043] Encoded video sequence 21 may include one or more coded layer video sequences in each of layers 24. A coded layer video sequence may begin with a starting image—for example, an independently coded image—and may include all images from each layer starting with that starting image, up to the next starting image of the coded layer video sequence in coding order 19 (excluding it), or up to the end of that coded layer video sequence. It should be noted that each of layers 24 may have a different number and / or different arrangement of coded layer video sequences. In other words, the starting images of coded layer video sequences from different layers are not necessarily aligned within the same access unit.

[0044] Figure 2An example of two layers 240 and 241 with different periods of IRAP images is illustrated.

[0045] When the bitstream contains multiple layers, the IRAP images in each layer do not need to be aligned, for example, the lower layer L0 (e.g. Figure 1 Layer 240) may have a higher dependency layer L1 with more frequent IRAP images (e.g. Figure 1 Layer 241). When the lower layer CLVS21′ stops at each such IRAP image, for example by having IDR AU at the lower layer instead of the higher layer, or as Figure 2 As shown in the diagram, the higher-level CLVS can continue through the sequence termination (EOS) NAL unit.

[0046] There exists a situation where the bitstream needs to include the so-called end-of-sequence (EOS) NAL unit 41 (at CLVSS image 260*) before the start of a new CLVS 21″. Figure 2 The CLVSS image (CRA) then has a NoOutputBeforeRecoveryFlag equal to 1, and the RASL image 260′, which cannot be correctly reconstructed due to the lack of references from before the CLVSS image 260*, is omitted from the output. However, when other images in AU 22′ of RASL image 260′ use RASL as a reference for prediction (e.g., samples or grammars such as MV), the other images ( Figure 2 The L1 Trail in the image (e.g., image 261′) will also be incorrectly reconstructed and subsequently output.

[0047] According to a first embodiment of the first aspect, encoder 10 is configured to encode non-RASL images in the images 261 of the first layer 241 in a manner that does not involve prediction from the RASL images 260' of the second layer 240, for example... Figure 2 Image 261' in the middle, these non-RASL images are compared with the RASL images in the second layer 240 (e.g. Figure 2 Image 260' in the image is temporally aligned. Furthermore, encoder 10 uses the RASL image of the second layer 240 as an inter-layer prediction reference for the RASL image of the first layer 241 to encode the RASL image of the first layer 241, which is temporally aligned with the RASL image of the second layer 240. For example, see... Figure 2Assuming image 261' is a RASL image, encoder 10 will use the RASL image 260' of layer 240 as an inter-layer prediction reference to encode image 261', for example, referred to as the first RASL image. The RASL image 260' of layer 240 can be referred to as the second RASL image. Figure 2 In the case shown, where image 261′ is a non-RASL image, the encoder 10 according to the first embodiment encodes the first RASL image 261′ in a manner that does not predict from the third RASL image 260′.

[0048] Using an image as an inter-layer prediction reference can indicate that, in vector-based inter-frame prediction and / or motion vector prediction, the same image is considered when recruiting previously encoded images to form a list of reference images for inter-frame prediction of the current encoded image. Alternatively, using an image as an inter-layer prediction reference can additionally indicate that the same image is considered as being referenced in the list of reference images by the reference index of the inter-frame prediction block of the current encoded image.

[0049] Encoding an image without making predictions from a specific image can foreshadow the avoidance of a specific image in vector-based inter-frame prediction and / or motion vector prediction when recruiting previously encoded images to form a reference image list for inter-frame prediction of the current encoded image, and / or preventing that specific image from being referenced in the reference image list of the current encoded image through the reference index of the inter-frame prediction block of the current encoded image.

[0050] In other words, according to the first embodiment of the first aspect, when the reference image is not a RASL image, the RASL image is prohibited from being used as an interlayer reference image, thereby allowing the L1 tail image in the same access unit as the RASL image to be correctly reconstructed. The constraint discussed herein means that RASL images are not used as references because they either do not exist in the RPL (Reference Image List) or are not selected from the RPL; that is, they may be inactive or unused references in the RPL.

[0051] However, this constraint is unnecessarily strict and may not be a problem in some cases, such as when there is no ESNUT, as... Figure 3 As shown in the diagram.

[0052] Figure 3 The illustration shows an example of a random access image without an ongoing EOS instruction.

[0053] In this case, the reference from L1 Trail to L0 RASL will only become a problem when it is loaded (random access) at CRA location 22*, but the decoder will skip all pictures in enhancement layer L1 anyway until it encounters an IRAP picture in the corresponding enhancement layer.

[0054] exist Figure 3 In the example, image 260* of the second layer 240 is a CRA image; for example, one or more RASL images can depend on independently encoded images. Figure 3 In the example, with Figure 2 Unlike other images, the second layer 240 image (which precedes CRA image 260* in coding order 19) does not have a sequence end indication 41. Nevertheless, CRA image 260* can be the start image of the coding layer video sequence, for example, it can be indicated that, in this case, there is no RASL image dependent on CRA image 260*, such as RASL image 260′ is indicated for output. For example, if the no output before recovery relaxation mentioned below is set to 1 and the HandleCraAsClvsStartFlag mentioned below is set to 1, then CRA image 260′ can be the start image of the coding layer video sequence.

[0055] According to a second embodiment of the first aspect, encoder 10 can use the RASL picture of second layer 240, implying that RASL picture 260' serves as an inter-layer prediction reference picture for a picture of first layer 241 (e.g., a portion of the same access unit) that is temporally aligned with the RASL picture of second layer 240—if the picture of first layer is a RASL picture, and if the CRA picture 260' associated with the RASL picture 260' of second layer 240 is the starting picture of the coded layer video sequence. In other words, in Figure 3In the example, images 261' and 260' are time-aligned, meaning they reside within the same access unit 22'. Image 260' of the second layer 240 is a RASL image associated with CRA image 260*. According to the second embodiment, if image 261' is a RASL image, encoder 10 uses image 260' as an interlayer reference image for image 261'. If image 261' is a non-RASL image, encoder 10 uses image 260' as an interlayer reference image for image 261'—if CRA image 260* does not form the beginning of the coded layer video sequence 21″. Otherwise, that is, if image 261' is a non-RASL image and CRA image 260* forms the beginning of the coded layer video sequence 21″, encoder 10 encodes image 261' without using image 260' as an interlayer reference image. If CRA image 260* forms the beginning of the coding layer video sequence 21″, the previous images in coding sequence 19 may be unavailable, and therefore RASL image 260′ may not be decoded correctly.

[0056] In other words, according to the example of the second embodiment of the first aspect, the first embodiment above (where RASL images are prohibited from being used as inter-layer reference images when the reference image is not a RASL image) is subject to the following conditions:

[0057] • RASL image associated with the CRA following the EOS NAL unit (see Figure 2 )

[0058] • RASL images associated with CRA (see Figure 3 To address this, we externally set HandleCraAsClvsStartFlag to 1 and NoOutputBeforeRecoveryFlag to 1.

[0059] The latter case is where the API notifies the decoder that any CRA is considered the start of a new CLVS, and therefore the processing is the same as when there is an EOS NAL unit but no such NAL unit exists.

[0060] In other words, the constraint (disallowing RASL as an ILRP reference image) can be conditional on the associated CRA where NoOutputBeforeRecoveryFlag is set to 1 (regardless of whether EOS NAL or external is present), as follows:

[0061] - The following constraints apply to the image referenced by each ILRP entry (if present) in RefPicList[0] or RefPicList[1] of the slice of the current image:

[0062] The image should be located in the same AU as the current image.

[0063] The image should exist in the DPB.

[0064] The image should have a nuh_layer_idrefPicLayerId that is smaller than the nuh_layer_id of the current image.

[0065] The image should not be a RASL image if the associated CRA has a NoOutputBeforeRecoveryFlag set to 1 and the current image is not a RASL image.

[0066] Any of the following constraints apply:

[0067] The image should be an IRAP image.

[0068] The TemporalId of the image should be less than or equal to Max(0, vps_max_tid_il_ref_pics_plus1[currLayerIdx][refLayerIdx]-1), where currLayerIdx and refLayerIdx are equal to GeneralLayerIdx[nuh_layer_id] and GeneralLayerIdx[refpicLayerId], respectively.

[0069] Figure 4 The illustration shows an example of encoding video sequence 20 when it can be encoded into video data stream 14. Encoded video sequence 20 includes a first layer 240 and a second layer 241, the second layer being a reference layer for the first layer. In access unit 22', the second layer 240 has a sequence end indication 41, which indicates that access unit 22' is the last access unit in the encoding order 19 of the encoded layer video sequence 210' of the second layer 240, and indicates that a new encoded layer video sequence 210' of the second layer 240 begins in access unit 22'' following access unit 22' in the encoding order 19.

[0070] According to a third embodiment of the first aspect, the encoder 10 is configured to insert each such access unit 22' having a sequence end indication 41 (also referred to as a sequence end identifier) ​​and the sequence end indication 41 into the first layer 241, such as Figure 4 As shown in the diagram.

[0071] Therefore, since the first layer 241 has a sequence end indicator 41 in the access unit 22', the end of the coding layer video sequence 211' of the first layer 24' is with the access unit 22', and thus serves as the coding layer video sequence 210' of the second layer 240 in the same access unit. Accordingly, in layer 241, a new coding layer video sequence 211″ begins with the access unit 22″, synchronized with the start of the next coding layer video sequence 210″ of the second layer 240. Due to the sequence end indicator 41 in the first layer 241, the RASL image in the next coding layer video sequence 211″ will not appear, thus avoiding the aforementioned problem of asynchronous coding layer video sequence boundaries.

[0072] about Figure 5 An alternative to the third embodiment is described. Figure 5 The diagram illustrates the information based on... Figure 4 The video bitstream of the described scene. However, according to this embodiment, encoder 10 does not necessarily insert the sequence end indication 41 (which it may do in the example) at access unit 22' in the first layer 241. According to an alternative third embodiment, encoder 10 is configured to encode within the next access unit 22″, 22″' after access unit 22' with sequence end indication 41 in the second layer 240, wherein image 261″ will be encoded into access units 22″, 22″', image 261″ will be refreshed by decoder and encoded into the first layer 241 without a preceding image output. In other words, the next access units 22″, 22″' are access units that follow access unit 22' in the encoding order 19 and into which images of the first layer are to be encoded, and among these access units, the next access units 22″, 22″' are the access units that follow access unit 22' and are closest to access unit 22'. Figure 5 The diagram illustrates two examples of the first layer, referenced using reference markers 241 and 242. Note that these two examples of the first layer are shown together for illustrative purposes. Figure 5 The following is shown, but can represent a standalone example. Therefore, in the example, one or more such first layers may exist in the encoded video sequence 20. In layer 242, the next access unit is access unit 22″′, into which image 262″′ will be encoded (because access unit 22″ has no image in the first layer), and in layer 241, the next access unit is 22″, into which image 261″ will be encoded.

[0073] By encoding images 261″ and 261″′ of the first layer 241 and 242 through decoding refresh, which is independent of the second layer 240 and has no preceding image output, the result is that there will be no images of the first layer 241 and 242 that depend on the second layer 240 as a RASL image, which in the example could be image 260″′, since they could also be RASL images.

[0074] Encoding images 261″ and 261″′ using a decode refresh can, for example, imply encoding images 261″ and 261″′ without referencing images belonging to access units different from their respective access units (i.e., access units 22″ and 22″′). For example, images 261″ and 261″′ could be IDR or CRA images. The term "preceding image" for images 261″ and 261″′ can refer to an image that follows images 261″ and 261″′ in encoding order 19 but precedes images 261″ and 262″′ in presentation order 18, which depend on images preceding images 261″ and 262″′ in encoding order 19 (i.e., including references to them). An example of a preceding image could be a RASL image. Therefore, encoding images 261″, 261″′ without outputting a preamble image can, for example, indicate that the images in the first layer 241, 242 that are after images 261″, 262″′ in encoding order 19 but before images 261″, 262″′ in presentation order do not exist (e.g., no RASL image), or are indicated as not being output. In other words, according to an example of an alternative third embodiment, encoder 10 is configured to encode the first layer such that image 261″ in the first layer 241 to be encoded into the next access unit 22″, 22″′ is the starting image of the encoded layer video sequence.

[0075] For example, encoding images 261″ and 261″′ as IDRs might mean that there are no preceding images in the first layers 241 and 242, such as RASL images following images 261″ and 261″′ in the encoding order. In other words, image 261″ being an IDR image might prohibit image 261″′ from being a preceding image, and therefore also prohibit image 261″′ from having inter-layer references to RASL images, such as... Figure 3 The reference between images 261′ and 260′. Therefore, an example of encoding images 261″ and 261″′ using a decode refresh and without a leading image output is to encode images 261″ and 261″′ as an IDR.

[0076] Alternatively, images 261″ and 261″′ can be encoded as CRA, and the output without a preceding image can be accomplished, for example, using the sequence end indication 41 in the access unit 22′ of the first layers 241 and 242. In this case, in the example, the preceding image following images 261″ and 261″′ in the encoding order may exist in the first layers 241 and 242, but can be excluded from the output so that incorrect decoding will not take effect. Therefore, even if image 261″′ of layer 241 depends on image 260″′ of layer 240 and image 260″′ is a RASL image, image 260″′ is not output.

[0077] Regarding Figure 4 Compared to the constraints of the first alternative in the explained third embodiment, regarding Figure 5 The second alternative example has the advantage that encoder 10 does not necessarily have to insert the sequence end indication 41 in the first layer 241.

[0078] In other words, according to the example of the third embodiment, the bitstream constraint is: when layer k depends on layer 1 and layer 1 contains EOS NAL units, layer k also contains EOS NAL units at the same location, or the next AU must contain a CLVSS picture (e.g., IDR) for layer k.

[0079] 2. Image output timing

[0080] This section describes an embodiment according to the second aspect, which includes the first to sixth sub-aspects. (See also...) Figure 1 An embodiment according to the second aspect has been described, the details and features of which may optionally be included in the embodiment according to the second aspect. Furthermore, the details described in Part 1 may optionally be applied to the embodiment of the second aspect, such as details regarding the encoding type of images, dependencies between images, encoded video layer sequences, etc.

[0081] Such as about Figure 1 As described, decoder 50 decodes video bitstream 14 or a portion thereof to provide a decoded video sequence 51. Decoder 50 may provide a decoded picture 53 of the decoded video sequence 51 in a decoded picture buffer (DPB). Some embodiments according to the second aspect may involve the output timing of the decoded picture 53 from the decoded picture buffer, which, for example, may be provided from the decoded picture buffer to a display for presentation. Decoder 50 may provide the decoded picture 53 to the decoded picture buffer according to a presentation order 18. (See also: Regarding...) Figure 1As described, decoder 50 does not necessarily have to decode and / or output all the images in video bitstream 14, but it can decode and / or output a subset of images encoded into video bitstream 14. This subset of images can be defined by the layers associated with the images and / or by the definition of temporal subsets of the images. Temporal subsets can be defined by temporal layers, as described regarding... Figure 6 As described.

[0082] Figure 6 The illustration shows an example of the temporal layer of the encoded video sequence 20. Figure 6 The illustration shows layer 240, which includes image 26 from a first time layer 250. Another layer 241 includes images from a second time layer 251. The first time layer 250 and the second time layer 251 can have the same frame rate, such as... Figure 6 As illustrated in the diagram, the image in the first time sublayer 250 may belong to a different access unit 22 than the image in the second time sublayer 251. In other words, the image in the first time sublayer 250 may belong to a different time instant than the image in the second time sublayer 251. Note that in... Figure 6 In the image 26, the arrangement follows the presentation order 18 rather than the encoding order 19. Figure 6 Another example of a layer is illustrated, namely layer 242, which includes images for each of the first time sublayer 250 and the second time sublayer 251. Therefore, layer 242 has a higher frame rate than layers 240 and 241. Note that... Figure 6 The combination of layers illustrated is an illustrative example, and the encoded video sequence can include any combination of layers, each of which can include one or more temporal sublayers. Therefore, each of the access units 22 can include images of one or more layers 24.

[0083] Time sub-layers can have a hierarchical order, which can be defined, for example, by an index associated with the time sub-layer. For example, a second time sub-layer 251 can be higher than a first time sub-layer 250 in the hierarchical order.

[0084] Decoder 50 can select one or more of the time sub-layers included in the video bitstream 14 for decoding, for example, by selecting the maximum time sub-layer for decoding. That is, the decoder can decode all time sub-layers that are equal to or lower than the maximum time sub-layer in hierarchical order.

[0085] For example, decoder 50 can receive an instruction indicating which temporal sub-layer of video bitstream 14 to decode. In other examples, decoder 50 can determine the maximum temporal sub-layer to be decoded itself. In other words, the temporal subset of image 26 to be decoded by decoder 50 can be defined by selecting the maximum temporal sub-layer to be decoded.

[0086] As mentioned above, the subset of pictures 26 to be decoded can be further defined by selecting a subset of layers 24 included in the video bitstream 14 for decoding. The video bitstream 14 can provide several options for a decodeable bitstream. For example, a single layer of the video bitstream 14 can represent a decodeable bitstream that can be decoded by the decoder 50 independently of other layers. Alternatively, a combination or all layers of the video bitstream 14 can represent a decoded bitstream and can be selected for decoding. For example, the video bitstream 14 can include OLS indications, such as descriptive data 23 of the video bitstream 14. An OLS indication can indicate one or more sets of output layers (OLS). Each OLS can indicate that one or more of the layers 24 of the video bitstream 14 belong to the OLS. In other words, an OLS can include one or more or all of the layers 24. In the example, one or more or all of the layers in the OLS can be indicated as output layers of the OLS. The OLS may optionally further include non-output layers. For example, in the case of a quality-scalable bitstream, the reference layer of the output layer of the OLS can be included in the OLS because the output layer that references the reference layer may need the reference layer for decoding, although the reference layer itself does not necessarily have to be the output layer.

[0087] Decoder 50 can select an OLS from those indicated in the OLS specification of the video bitstream 14 for decoding, for example, based on instructions provided externally. In other examples, decoder 50 can choose the OLS to be decoded itself. Therefore, the bitstream to be decoded can be defined by selecting the OLS and the maximum time sublayer used for decoding.

[0088] For example, decoder 50 can provide decoded image 53 to decoded image buffer for each image that is part of the output layer of the OLS to be decoded and included in the temporal subset of the images (e.g., defined by the maximum temporal sublayer).

[0089] For example, the maximum time sublayer to be decoded can be represented by the variable Htid, which can be provided to decoder 50 or derived by decoder 50.

[0090] Decoder 50 can output decoded images 53 from the output buffer (i.e., the decoded image buffer) at the output time. In other words, decoder 50 can determine the output time for each decoded image 53. For example, the output time of the images can be provided to decoder 50 within the descriptive data of video bitstream 14. For example, the output time of the images can be provided by the Picture Timing (PT) Supplemental Enhancement Information (SEI) message in video bitstream 14. For example, a PT SEI message can be provided for each of the access units 22. However, video bitstream 14 does not necessarily have to provide such output timing information. Instead, decoder 50 can determine the output timing for the images. For example, if the bitstream decoded by decoder 50 has a constant output image rate, decoder 50 can deduce the output timing of the images itself.

[0091] The current specification (such as VVC) contains the following text to express that the bitstream has a constant output image rate.

[0092] For a CVS containing image n, when Htid equals i and fixed_pic_rate_general_flag[i] equals 1, and image n is the output image rather than the last image in the output bitstream (in output order), the value calculated for DpbOutputElementalInterval[n] should be equal to ClockTick*(elemental_duration_in_tc_minus1[i]+1), where ClockTick is specified as in Equation C.1 (using the value of ClockTick of the CVS containing image n) if one of the following conditions is true for the image nextPicInOutputOrder specified in Equation C.16:

[0093] The image nextPicInOutputOrder is in the same CVS as image n.

[0094] - The image nextPicInOutputOrder is in a different CVS and fixed_pic_rate_general_flag[i] is equal to 1 in the CVS containing the image nextPicInOutputOrder, the value of ClockTick is the same for both CVSs, and the value of elemental_duration_in_tc_minus1[i] is the same for both CVSs.

[0095] For a CVS containing image n, when Htid equals i and fixed_pic_rate_within_cvs_flag[i] equals 1, and image n is an output image rather than the last image in the output CVS (in output order), the value calculated for DpbOutputElementalInterval[n] should be equal to ClockTick*(elemental_duration_in_tc_minus1[i]+1), where ClockTick is specified in Equation C.1 (using the value of ClockTick of the CVS containing image n) when the following images are in the same CVS as image n according to the output order nextPicInOutputOrder specified in Equation C.16:

[0096] In summary, there are two control flags (sequence parameter sets, such as descriptive data associated with each coded video sequence, e.g., coded video sequence (CVS)20) in the SPS as part of the hypothetical reference decoder (HRD) parameters. One control flag is `fixed_pic_rate_within_cvs_flag`, which indicates that all output pictures within a CLVS (or CVS) have equidistant output times. The other control flag is `fixed_pic_rate_general_flag`, which indicates that, provided the ClockTick values ​​of two CVSs are the same and the `elemental_duration_in_tc_minus1` values ​​are the same, CVS padding starting from the first AU referencing the SPS also has equidistant output times between the output pictures at the boundaries of the CVS and the previous CVS.

[0097] Note that a signaling (fixed_pic_rate_within_cvs_flag) is given for each sub-layer (e.g., temporal sub-layer 25) indicating whether the output rate is constant. This means that if the generated bitstream allows for temporal scalability, this property (constant picture rate) is signaled for each possible frame rate achievable when receiving different numbers of sub-layers. HTid refers to the highest time ID present in the bitstream. For example, if the initial bitstream has four sub-layers with time IDs from 0 to 3, and the highest one is discarded, then HTid becomes 2, and the parameter fixed_pic_rate_within_cvs_flag with HTid = 2 is considered at the decoder to evaluate whether the output rate is constant.

[0098] The problem is, when... Figure 7When performing splicing or editing involving CVS cascading as instructed in the documentation, this solution requires modification of the SPS (e.g., fixed_pic_rate_general_flag).

[0099] Figure 7 The illustration shows an example of SPS modification, which can be an encoded video sequence of a previously encoded video sequence. Figure 7 The upper panel illustrates the first case, in which the splicing of the first video sequence 201 and the second video sequence 202 produces a bitstream, wherein the picture rate, defined by the timing interval 71 between consecutive pictures, is constant at the boundary between the first video sequence 201 and the second video sequence 202. Figure 7 The lower panel illustrates the second case, where the boundary between the timing interval 71 and the first video sequence 201 and the second video sequence 202 is not constant at the boundary.

[0100] It is important to note that, Figure 7 and the following Figures 8-10 and Figure 12 In the image, the images belonging to the common time sublayer 25 are illustrated on a common horizontal plane with respect to their vertical positions.

[0101] One advantage of indicating this information in the bitstream (e.g., fixed_pic_rate_general_flag) is that, in addition to signaling a constant output frame rate in the bitstream, it allows the output time to be derived using the constant output frame rate attribute, rather than using overly complex HRD parameters such as buffer period SEI messages or picture timing SEI messages (or ignoring them). That is, for example, PT SEI messages and / or BP SEI messages can be omitted, i.e., not present in the video bitstream 14, or ignored by the decoder 50 when deriving the output time of the decoded picture.

[0102] One problem with the specification having to derive output times in this way (e.g., without using PT and / or BP SEI messages) is that the output time of the first AU in the CVS cannot be determined. If one knew the output time of the first AU of the CVS, then the output times of other frames could be easily determined when a flag indicating that the output time within the CVS is constant is added by simply incrementing the signal (ClockTick*(elemental_duration_in_tc_minus1[i]+1)) to the previous output frame.

[0103] Please also note that the current specification indicates the following:

[0104] `elemental_duration_in_tc_minus1[i]` incremented by 1 (if present) specifies the time interval in clock ticks between pixels of consecutive images whose HRD output time is specified in the output order below when `Htid` equals `i`. The value of `elemental_duration_in_tc_minus1[i]` should be in the range of 0 to 2047 (inclusive).

[0105] For a CVS containing image n, the value of the variable DpbOutputElementalInterval[n] is specified as follows when Htid equals i and fixed_pic_rate_general_flag[i] equals 1, and image n is the output image rather than the last image in the output bitstream (in output order):

[0106] DpbOutputElementalInterval[n]=DpbOutputInterval[n]÷elementalOutputPeriods (113)

[0107] Where DpbOutputInterval[n] is specified in formula C.16, and elementalOutputPeriods is specified as follows:

[0108] - If a PT SEI message exists for image n, then elementalOutputPeriods equals the value of pt_display_elemental_periods_minus1+1.

[0109] Otherwise, elementalOutputPeriods equals 1.

[0110] This means that a constant output rate does not necessarily apply to decoded images, but it does apply to images that are displayed / output; that is, it does not apply to DpbOutputInterval[n], but it does apply to DpbOutputElementalInterval[n]. In other words, a constant output rate includes the repetition of a frame, meaning that elementalOutputPeriods not equal to 1 means that an image is repeated. Figure 8 An example is given.

[0111] Figure 8 An example of a video bitstream is illustrated, for example... Figure 7An example of the upper panel, where there is no image in access unit 22*. In situations where this may occur, such as when an image is lost during transmission, incorrectly decoded, excluded from output, or is not present, image 26* from a previous access unit can be repeated to achieve a constant frame rate. For example, an indication may be present in the access unit to which image 26* belongs, indicating that the image in the access unit should be repeated. For example, the indication may indicate the number of repetitions.

[0112] For example, DpbOutputInterval[n] can represent the duration of the time interval for the output of access unit 22, that is, the duration of content belonging to a common time frame, such as the repeated output of the image of the access unit, for example. Figure 8 The access unit output interval is 61. In contrast, DpbOutputElementalInterval[n] can represent the time interval for the output of a single element, i.e., for an image or for repetitions of an image, for example... Figure 8 The image output interval is 63.

[0113] The syntax element used for repetition (pt_display_elemental_periods_minus1) is not always a frame repetition. It may also be used for interlaced content when a frame that is encoded and decoded as a frame is displayed as a field in the display step.

[0114] Please refer to the following specification text:

[0115] When sps_field_seq_flag equals 0 and fixed_pic_rate_within_cvs_flag[TemporalId] equals 1, a value greater than 0 for pt_display_elemental_periods_minus1 can be used to indicate the frame repetition period of a display using a fixed frame refresh interval equal to DpbOutputElementalInterval[n], as shown in Equation 113.

[0116] The following grammatical examples and their semantics are illustrative and should be easy to understand:

[0117]

[0118]

[0119] The PT SEI message provides the AU associated with the SEI message with information on CPB removal delay and DPB output delay.

[0120] If the bp_nal_hrd_params_present_flag or bp_vcl_hrd_params_present_flag of the BP SEI message applicable to the current AU is equal to 1, then the variable CpbDpbDelaysPresentFlag is set to 1; otherwise, CpbDpbDelaysPresentFlag is set to 0.

[0121] The presence of the PT SEI message is specified as follows:

[0122] - If CpbDpbDelaysPresentFlag equals 1, then the PT SEI message should be associated with the current AU.

[0123] Otherwise (CpbDpbDelaysPresentFlag equals 0), there should be no PT SEI message associated with the current AU.

[0124] The TemporalId in the PT SEI message syntax is the TemporalId of the SEI NAL unit containing the PT SEI message.

[0125] The increment of 1 in pt_cpb_removal_delay_minus1[i] is used to calculate the number of clock ticks between the nominal CPB removal time of the AU associated with the PT SEI message and the preceding AU containing the BP SEI message in the decoding order when Htid equals i. This value is also used to calculate the earliest possible time when AU data arrives at the CPB of the HSS. The length of pt_cpb_removal_delay_minus1[i] is bp_cpb_removal_delay_length_minus1+1 bits.

[0126] pt_cpb_alt_timing_info_present_flag equal to 1 specifies the syntax element.

[0127] The following syntax elements can exist in the PT SEI message: pt_nal_cpb_alt_initial_removal_delay_delta[i][j], pt_nal_cpb_delay_offset_delta[i][j], pt_nal_cpb_delay_offset[i], pt_nal_dpb_delay_offset[i], pt_vcl_cpb_alt_initial_removal_delay_delta[i][j], pt_vcl_cpb_alt_initial_removal_offset_delta[i][j], pt_vcl_cpb_delay_offset[i], and pt_vcl_dpb_delay_offset[i]. A pt_cpb_alt_timing_info_present_flag value of 0 indicates that these syntax elements do not exist in the PT SEI message. When the associated image is a RASL image, the value of pt_cpb_alt_timing_info_present_flag should be equal to 0.

[0128] Note 1 - For more than one AU following the IRAP image in the decoding order, the value of pt_cpb_alt_timing_info_present_flag may be equal to 1. However, the alternative timing is only applied to the first AU that has pt_cpb_alt_timing_info_present_flag equal to 1 and is following the IRAP image in the decoding order.

[0129] pt_nal_cpb_alt_initial_removal_delay_delta[i][j] specifies the alternative initial CPB removal delay increment for the i-th sublayer of the j-th CPB in the NAL HRD, in units of 90kHz clock. The length of pt_nal_cpb_alt_initial_removal_delay_delta[i][j] is bp_cpb_initial_removal_delay_length_minus1+1 bits.

[0130] When pt_cpb_alt_timing_info_present_flag is equal to 1 and pt_nal_cpb_alt_initial_removal_delay_delta[i][j] does not exist for any i value less than bp_max_sublayers_minus1, its value is inferred to be equal to 0.

[0131] pt_nal_cpb_alt_initial_removal_offset_delta[i][j] specifies the alternative initial CPB removal offset increment for the i-th sublayer of the j-th CPB of the NAL HRD, in units of 90kHz clock. The length of pt_nal_cpb_alt_initial_removal_offset_delta[i][j] is bp_cpb_initial_removal_delay_length_minus1+1 bits.

[0132] When pt_cpb_alt_timing_info_present_flag is equal to 1 and pt_nal_cpb_alt_initial_removal_offset_delta[i][j] does not exist for any i value less than bp_max_sublayers_minus1, its value is inferred to be equal to 0.

[0133] `pt_nal_cpb_delay_offset[i]` specifies the offset used to derive the nominal CPB removal time of the AU associated with the PT SEI message and the AU that follows it in decoding order for the i-th sublayer of the NAL HRD. The length of `pt_nal_cpb_delay_offset[i]` is `bp_cpb_removal_delay_length_minus1+1` bits. When it does not exist, the value of `pt_nal_cpb_delay_offset[i]` is inferred to be 0.

[0134] `pt_nal_dpb_delay_offset[i]` specifies the offset used when deriving the DPB output time of the IRAP AU associated with the BPSEI message for the i-th sublayer of the NAL HRD, when the AU associated with the PT SEI message directly follows the IRAP AU associated with the BP SEI message in decoding order. The length of `pt_nal_dpb_delay_offset[i]` is `bp_dpb_output_delay_length_minus1+1` bits. When it does not exist, the value of `pt_nal_dpb_delay_offset[i]` is inferred to be equal to 0.

[0135] `pt_vcl_cpb_alt_initial_removal_delay_delta[i][j]` specifies the replacement initial CPB removal delay increment for the i-th sublayer of the j-th CPB of the VCL HRD, in units of 90kHz clock. The length of `pt_vcl_cpb_alt_initial_removal_delay_delta[i][j]` is `bp_cpb_initial_removal_delay_length_minus1+1` bits.

[0136] When pt_cpb_alt_timing_info_present_flag is equal to 1 and pt_vcl_cpb_alt_initial_removal_delay_delta[i][j] does not exist for any i value less than bp_max_sublayers_minus1, its value is inferred to be equal to 0.

[0137] `pt_vcl_cpb_alt_initial_removal_offset_delta[i][j]` specifies the alternative initial CPB removal offset increment for the i-th sublayer of the j-th CPB of the VCL HRD, in units of 90kHz clock. The length of `pt_vcl_cpb_alt_initial_removal_offset_delta[i][j]` is bp_cpb_initial_removal_delay_length_minus1+1 bits.

[0138] When pt_cpb_alt_timing_info_present_flag is equal to 1 and pt_vcl_cpb_alt_initial_removal_offset_delta[i][j] does not exist for any i value less than bp_max_sublayers_minus1, its value is inferred to be equal to 0.

[0139] `pt_vcl_cpb_delay_offset[i]` specifies the offset used to derive the nominal CPB removal time of the AU associated with the PT SEI message and the AU that follows it in decoding order for the i-th sublayer of the VCL HRD. The length of `pt_vcl_cpb_delay_offset[i]` is `bp_cpb_removal_delay_length_minus1+1` bits. When it does not exist, the value of `pt_vcl_cpb_delay_offset[i]` is inferred to be equal to 0.

[0140] `pt_vcl_dpb_delay_offset[i]` specifies the offset used when deriving the DPB output time of the IRAP AU associated with the BPSEI message for the i-th sublayer of the VCL HRD, when the AU associated with the PT SEI message directly follows the IRAP AU associated with the BP SEI message in decoding order. The length of `pt_vcl_dpb_delay_offset[i]` is `bp_dpb_output_delay_length_minus1+1` bits. When it does not exist, the value of `pt_vcl_dpb_delay_offset[i]` is inferred to be equal to 0.

[0141] The variable BpResetFlag for the current image is derived as follows:

[0142] - If the current image is associated with a BP SEI message, then BpResetFlag is set to 1.

[0143] Otherwise, BpResetFlag is set to 0.

[0144] The value of pt_sublayer_delays_present_flag[i] equal to 1 indicates that pt_cpb_removal_delay_delta_idx[i] or pt_cpb_removal_delay_minus1[i], and pt_du_common_cpb_removal_delay_increment_minus1[i] or pt_du_cpb_removal_delay_delta_minus1[][] exist in the sublayer where TemporalId equals i. The sublayer_delays_present_flag[i] being equal to 0 indicates that pt_cpb_removal_delay_delta_idx[i] and pt_cpb_removal_delay_minus1[i], as well as pt_du_common_cpb_removal_delay_increment_minus1[i] and pt_du_cpb_removal_delay_increment_minus1[], do not exist for the sublayer with TemporalId equal to i. The value of pt_sublayer_delays_present_flag[bp_max_sublayers_minus1] is inferred to be equal to 1. When it does not exist, the value of pt_sublayer_delays_present_flag[i] is inferred to be equal to 0 for any i in the range from 0 to bp_max_sublayers_minus1-1 (inclusive).

[0145] A value of 1 for pt_cpb_removal_delay_delta_enabled_flag[i] indicates that pt_cpb_removal_delay_delta_idx[i] exists in the PT SEI message. A value of 0 for pt_cpb_removal_delay_delta_enabled_flag[i] indicates that pt_cpb_removal_delay_delta_idx[i] does not exist in the PT SEI message. When it does not exist, the value of pt_cpb_removal_delay_delta_enabled_flag[i] is inferred to be 0.

[0146] `pt_cpb_removal_delay_delta_idx[i]` specifies the index in the list `bp_cpb_removal_delay_delta_val[j]` of the CPB removal increment applied to an Htid equal to `i`, where `j` ranges from 0 to `bp_num_cpb_removal_delay_deltas_minus1` (inclusive). The length of `pt_cpb_removal_delay_delta_idx[i]` is `Ceil(Log2(bp_num_cpb_removal_delay_deltas_minus1+1))` bits. The value of `pt_cpb_removal_delay_delta_idx[i]` is inferred to be 0 when `pt_cpb_removal_delay_delta_idx[i]` does not exist and `pt_cpb_removal_delay_delta_enabled_flag[i]` is equal to 1.

[0147] The variables CpbRemovalDelayMsb[i] and CpbRemovalDelayVal[i] for the current image are derived as follows:

[0148] - If the current AU is the AU that initializes the HRD, then both CpbRemovalDelayMsb[i] and CpbRemovalDelayVal[i] are set to 0, and the value of cpbRemovalDelayValTmp[i] is set to pt_cpb_removal_delay_minus1[i]+1.

[0149] - Otherwise, let the image prevNonDiscardablePic be the preceding image with TemporalId equal to 0 in the decoding order. It is not RASL or RADL. For the image prevNonDiscardablePic, let prevCpbRemovalDelayMinus1[i], prevCpbRemovalDelayMsb[i], and prevBpResetFlag be set to the values ​​of cpbRemovalDelayValTmp[i]-1, CpbRemovalDelayMsb[i], and BpResetFlag, respectively, and the following applies:

[0150] -CpbRemovalDelayMsb[i] is derived as follows:

[0151] cpbRemovalDelayValTmp[i]=pt_cpb_removal_delay_delta_enabled_flag[i]?

[0152] pt_cpb_removal_delay_minus1[bp_max_sublayers_minus1]+1+

[0153] bp_cpb_removal_delay_delta_val[pt_cpb_removal_delay_delta_idx[i]]:

[0154] pt_cpb_removal_delay_minus1[i]+1

[0155] If (prevBpResetFlag)

[0156] CpbRemovalDelayMsb[i]=0

[0157] Otherwise, if (cpbRemovalDelayValTmp[i]) <prevCpbRemovalDelayMinus1[i])

[0158] CpbRemovalDelayMsb[i]=prevCpbRemovalDelayMsb[i]+2 bp _cpb_removal_delay_length_minus1+1 (D.1)

[0159] otherwise

[0160] CpbRemovalDelayMsb[i]=prevCpbRemovalDelayMsb[i]

[0161] -CpbRemovalDelayVal is derived as follows:

[0162] If (pt_sublayer_delays_present_flag[i])

[0163] CpbRemovalDelayVal[i]=CpbRemovalDelayMsb[i]+cpbRemovalDelayValTmp[i](D.2)

[0164] otherwise

[0165] CpbRemovalDelayVal[i]=CpbRemovalDelayVal[i+1]

[0166] The value of CpbRemovalDelayVal[i] should be between 1 and 2. 32 Within the range (including the endpoints).

[0167] The variable AuDpbOutputDelta[i] is derived as follows:

[0168] AuDpbOutputDelta[i]=CpbRemovalDelayVal[i]-

[0169] (pt_cpb_removal_delay_minus1[bp_max_sublayers_minus1]+1)-(D.3)

[0170] (i==bp_max_sublayers_minus1? 0: bp_dpb_output_tid_offset[i])

[0171] The value of bp_dpb_output_tid_offset[i] is found in the associated BP SEI message.

[0172] `pt_dpb_output_delay` is used to calculate the DPB output time of the image. It specifies how many clock ticks to wait after removing the AU from the CPB before decoding the image from the DPB output.

[0173] Note 2 - When the decoded image is still labeled as "used for short-term reference" or "used for long-term reference", the decoded image will not be removed from the DPB at its output time.

[0174] The length of pt_dpb_output_delay is bp_dpb_output_delay_length_minus1 + 1 bits. When max_dec_pic_buffering_minus1[Htid] equals 0, the value of pt_dpb_output_delay should be equal to 0.

[0175] The output time of any image output from the decoder that conforms to the output timing, deduced from pt_dpb_output_delay, should precede the output time of all images in any subsequent CVS, in the order of decoding.

[0176] The order of image output established by the value of this syntax element should be the same as the order established by the value of PicOrderCntVal.

[0177] For images not output during the "collision" process, since they precede CLVSS images whose ph_no_output_of_prior_pics_flag is equal to 1 or inferred to be equal to 1 in the decoding order, the output time derived from pt_dpb_output_delay should increase with the value of PicOrderCntVal relative to all images within the same CVS.

[0178] `pt_dpb_output_du_delay` is used to calculate the DPB output time of the image when `DecodingUnitHrdFlag` equals 1. It specifies how many sub-clock ticks to wait after removing the last DU from an AU in the CPB before decoding the image from the DPB output.

[0179] The length of the syntax element pt_dpb_output_du_delay is given in bits by bp_dpb_output_delay_du_length_minus1+1.

[0180] The output time of any image output from the decoder that conforms to the output timing, deduced from pt_dpb_output_du_delay, should precede the output time of all images in any subsequent CVS, in the order of decoding.

[0181] The order of image output established by the value of this syntax element should be the same as the order established by the value of PicOrderCntVal.

[0182] For images not output during the "collision" process, since they precede CLVSS images whose ph_no_output_of_prior_pics_flag is equal to 1 or inferred to be equal to 1 in the decoding order, the output time derived from pt_dpb_output_du_delay should increase with the value of PicOrderCntVal relative to all images within the same CVS.

[0183] For any two images in CVS, the difference in output time between the two images when DecodingUnitHrdFlag=1 should be the same as the difference when DecodingUnitHrdFlag=0.

[0184] The increment of 1 in pt_num_decoding_units_minus1 specifies the number of DUs in the AU associated with the PT SEI message. The value of pt_num_decoding_units_minus1 should be in the range of 0 to PicSizeInCtbsY-1 (inclusive).

[0185] A value of 1 for `pt_du_common_cpb_removal_delay_flag` indicates that the syntax element `pt_du_common_cpb_removal_delay_increment_minus1[i]` exists. A value of 0 for `pt_du_common_cpb_removal_delay_flag` indicates that the syntax element `pt_du_common_cpb_removal_delay_increment_minus1[i]` does not exist. When it does not exist, `pt_du_common_cpb_removal_delay_flag` is inferred to be equal to 0.

[0186] `pt_du_common_cpb_removel_delay_encrement_minus1[i]` plus 1 specifies the duration, in clock sub-ticks, between the nominal CPB removal times of any two consecutive DUs in the AU associated with the PT SEI message in decoding order when Htid equals i. This value is also used to calculate the earliest possible time in the CPB of the HSS for DU data to arrive, as specified in Annex C. The length of this syntax element is `bp_du_cpb_removal_delay_increment_length_minus1+1` bits.

[0187] When pt_du_common_cpb_removal_delay_increment_minus1[i] does not exist for any i value less than bp_max_sublayers_minus1, its value is inferred to be equal to

[0188] pt_du_common_cpb_removal_delay_increment_minus1[bp_max_sublayers_minus1].

[0189] The increment of 1 in pt_num_nalus_in_du_minus1[i] specifies the number of NAL cells in the i-th DU of the AU associated with the PT SEI message. The value of pt_num_nalus_in_du_minus1[i] should be in the range of 0 to PicSizeInCtbsY-1 (inclusive).

[0190] The first DU of an AU consists of the first pt_num_nalus_in_du_minus1[0]+1 consecutive NAL units in the AU arranged in decoding order. The i-th (i>0) DU of an AU consists of the pt_num_nalus_in_du_minus1[i]+1 consecutive NAL units following the last NAL unit in the DU immediately preceding the AU in decoding order. Each DU should contain at least one VCL NAL unit. All non-VCL NAL units associated with a VCL NAL unit should be included in the same DU as the VCL NAL unit.

[0191] `pt_du_cpb_removal_delay_increment_minus1[i][j]` plus 1 specifies the duration, in clock sub-ticks, between the nominal CPB removal times of the (i+1)th DU and the ith DU in the AU associated with the PT SEI message, in decoding order, when Htid equals j. This value is also used to calculate the earliest possible time for DU data to arrive at the CPB of the HSS, as specified in Annex C. The length of this syntax element is `bp_du_cpb_removal_delay_increment_length_minus1+1` bits.

[0192] When pt_du_cpb_removal_delay_increment_minus1[i][j] does not exist for any j value less than bp_max_sublayers_minus1, its value is inferred to be equal to

[0193] pt_du_cpb_removal_delay_increment_minus1[i][bp_max_sublayers_minus1].

[0194] The value of pt_delay_for_concatenation_ensured_flag equal to 1 specifies the difference between the final arrival time and the CPB removal time of the AU associated with the PTSEI message, when followed by an AU with a BP SEI message, where bp_concatenation_flag equals 1 and InitCpbRemovalDelay[][][Ed.(YK): Check if the use of "InitCpbRemovalDelay[Htid][ScIdx]" here is accurate.] is less than or equal to the value of bp_max_initial_removal_delay_for_concatenation, using the nominal removal time of the subsequent AU from the CPB calculated with bp_cpb_removal_delay_delta_minus1.

[0195] The value of pt_delay_for_concatenation_ensured_flag being 0 indicates that the difference between the final arrival time of the AU associated with the PT SEI message and the CPB removal time may or may not exceed the value of max_val_initial_removal_delay_for_splicing.

[0196] Increment pt_display_elemental_periods_minus1 by 1. When sps_field_seq_flag equals 0 and fixed_pic_rate_within_cvs_flag[TemporalId] equals 1, it indicates the number of pixel image period intervals occupied by the current encoded image for the display model.

[0197] When fixed_pic_rate_within_cvs_flag[TemporalId] equals 0 or sps_field_seq_flag equals 1, the value of pt_display_elemental_periods_minus1 should be equal to 0.

[0198] When sps_field_seq_flag equals 0 and fixed_pic_rate_within_cvs_flag[TemporalId] equals 1, a value greater than 0 for pt_display_elemental_periods_minus1 can be used to indicate the frame repetition period of a display using a fixed frame refresh interval equal to DpbOutputElementalInterval[n], as given in Equation 112.

[0199] Let's continue our discussion of the problem addressed in this article.

[0200] Another issue that needs to be addressed is that sometimes a similar result is required, even if the PT SEI message may not exist, i.e., allowing repetition even when the PT SEI message is not present, since the PT SEI message is optional.

[0201] Also note the interaction with information in the Frame Field Information SEI message (required when `sps_field_seq_flag` equals 1, and optional when it equals 0) and the PT SEI message (`pt_display_elemental_periods_minus1`). This SEI (Frame Field Information SEI) also has syntax elements with the same values ​​as the PT SEI message. That is:

[0202] `display_elemental_periods_minus1` increments by 1. When it exists (it may be encoded only when `field_pic_flag` is off, or it may not be encoded when `field_pic_flag` is on) and `FixedPicRateWithinCvsFlag` is equal to 1, it indicates the number of pixel image period intervals occupied by the currently encoded image for the display model. The value of `display_elemental_periods_minus1` should be equal to `DisplayElementalPeriods - 1` and is constrained as follows:

[0203] - If display_fields_from_frame_flag equals 1, then display_elemental_periods_minus1 should equal 1 or 2.

[0204] Otherwise, when FixedPicRateWithinCvsFlag equals 0, display_elemental_periods_minus1 should equal 0.

[0205] Table 14 specifies the interpretation of combinations of field_pic_flag (in the SEI frame field; should be equal to sps_field_seq_flag), FixedPicRateWithinCvsFlag, bottom_field_flag, display_fields_from_frame_flag, top_field_first_flag, and display_elemental_periods_minus1 (via DisplayElementalPeriods), where non-existent syntax elements are indicated by "-". Combinations of syntax elements not listed in Table 14 are reserved for future use by ITU-T|ISO / IEC and should not exist in bitstreams conforming to this version of the specification.

[0206] Note 1 - When FixedPicRateWithinCvsFlag equals 1, the indicated display time is constrained to interpret the duration for a display model that follows the display pattern indicated by the value of the syntax element of the Frame Fields Information (SEI) message (although the display process is beyond the scope of this specification). Although the video decoder model may be specified to output only the decoded image of the entire crop, the modeled display behavior sometimes includes other steps, such as repeating a frame over multiple time intervals when display_fields_from_frame_flag equals 0, or sequentially displaying the fields of a frame when display_fields_from_frame_flag equals 1.

[0207] Note 2 - Frame doubling can be used to enhance display, for example, displaying 25Hz progressive scan video on a 50Hz progressive scan monitor, or 30Hz progressive scan video on a 60Hz progressive scan monitor. Alternating combinations of frame doubling and frame tripling every other frame can be used to enhance the display of 24Hz progressive scan video on a 60Hz progressive scan monitor.

[0208] Table 14 - Explanation of Frame Field Information Syntax Elements

[0209]

[0210] Please note that multi-layer problems arise when some output layers contain fields (e.g., in interlaced video, images can be partitioned into a first field and a second field, which are output in consecutive time instances). Therefore, images from the first field can be considered to belong to a first time sublayer, e.g., 250, and images from the second field can be considered to belong to a second time sublayer, e.g., 251; video sequences with fields in the layer thus have a higher frame rate, e.g., double frame rate), while some layers do not. Because the two sets together (i.e., a bitstream comprising two layers, one with fields and one without) will have different output frame rates; this means that, more generally, problems will arise for any multi-layer bitstream with output layers having different output frame rates, as follows: Figure 9 As shown in the image.

[0211] Figure 9 The diagram illustrates a video bitstream with a first layer 241 and a second layer 240. The first layer 241 comprises three temporal sublayers: sublayer 250, sublayer 251, and sublayer 252. In contrast, the second layer 240 includes the first temporal sublayer 250 and the image sublayer of the second temporal sublayer 251, but does not include the images of the third temporal sublayer 252. Therefore, the second layer 240 serves as a lower frame rate than the first layer, for example, half the frame rate.

[0212] in other words, Figure 9 The illustration shows an example of two-layer bitstreams with different frame rates. For example, the higher layer 241 may have fields while the lower layer 240 may have progressive frames (e.g., non-interlaced scanning).

[0213] exist Figure 9 In the example, the highest layer 241 has twice the frame rate of the lower layer 240 and there is no repetition. However, if a fixed frame rate is signaled for OLS, repetition in the lowest layer 240 may actually be expected, as follows: Figure 10 As shown in the image.

[0214] Figure 10 The diagram shows Figure 9 An example of a video bitstream where output frames in a lower layer 240, which has a lower encoded frame rate than the higher layer 241, are repeated. Due to this repetition of output frames, the output frame rates of the two layers may be equal. Figure 8 Similarly, the access unit output interval 63 of layer 240 includes multiple image output intervals 63.

[0215] In summary, this aspect of the invention solves the following problem, and embodiments thereof are described in the corresponding sections:

[0216] 1. Output layer set with different frame rates in the output layers (e.g., multiple views with fields in the enhancement layer and frames in the base layer).

[0217] 2. Interaction between PT SEI messages and Frame Field Information SEI messages

[0218] 3. Frame / field repetition without PT SEI message

[0219] 4. The constant output frame rate across CVS is derived instead of being sent by signal, which would otherwise require SPS rewriting after splicing.

[0220] 5. Derive the output time in the absence of or without the use of the PT SEI message.

[0221] 6. Handling of images without output and its impact on constant output frame rate: In fact, although not mentioned above, if some images are not output, deriving the output time can be complicated when the PT SEI message is absent or ignored.

[0222] Before describing the sub-aspects of the second aspect, a brief overview of the aforementioned components used to determine the output timing of the decoded picture is provided. For example, video bitstream 14 may include PT SEI messages that convey information about the timing of the picture output. PT SEI messages may include picture output multiplication syntax elements, such as pt_display_elemental_periods_minus1. For example, PT SEI messages can be signaled with information about the access unit level. In other words, PT SEI messages can be associated with access unit 22. That is, PT SEI messages can be valid for all pictures within an access unit 22. The picture output multiplication syntax elements signaled in the PT SEI messages can reveal information about whether the access unit referenced by each PT SEI message has undergone multiplication of the picture output, for example, for... Figure 10 The lower layer 240 is shown. For example, pt_display_elemental_periods_minus1 being zero can indicate that the image output has not been multiplied, while pt_display_elemental_periods_minus1 > 0 can indicate that the image output should be repeated. If the image output is multiplied, then, for example, by the value of pt_display_elemental_periods_minus1, the image output multiplication syntax element can indicate how many output images to generate from one image of the corresponding access unit.

[0223] It is worth noting that the decoder can deduce a variable named elementalOutputPeriods on pt_display_elemental_periods_minus1. For example, decoder 50 can set the cell output period to be equal to pt_display_elemental_periods_minus1+1.

[0224] In other words, encoder 10 can encode, video bitstream 14 can be included, and decoder 50 can decode PTSEI messages.

[0225] Furthermore, the video bitstream 14 may include Frame Field Supplemental Enhancement Information (Frame Field SEI) messages, also known as FFI SEI messages, encoded by encoder 10 and decoded by decoder 50. These messages convey information about the frame field structure of a predetermined access unit, such as whether the picture of the access unit is encoded as a frame or a field, and if encoded as a frame, whether it is output as a frame or a field, and in what order the bottom and top fields (first and second fields) are used for picture output. For example, the FFI SEI message may include further picture output multiplication syntax elements, such as FFI_display_elemental_periods_minus1. These further picture output multiplication syntax elements may indicate, for example, whether the encoded picture is to be repeated for the use of frame fields, i.e., whether the picture referenced by the FFI SEI message undergoes multi-picture output. It is important to note that, in contrast to PT SEI messages, FFI SEI may refer to a single picture rather than the entire access unit.

[0226] In the following text, the above-mentioned picture output multiplication syntax element of the PT SEI message can be referred to as the PT multiplication indicator, while the further picture output multiplication syntax element of the FFI SEI message can be referred to as the FF multiplication indicator.

[0227] For example, after deriving the information that the image will undergo the multiplication image output, decoder 50 can set the number of image outputs for the corresponding access unit (in the case of PT multiplication indication) or the corresponding image (e.g., in the case of FF multiplication indication) according to the number indicated by the image output multiplication syntax element or further image output multiplication syntax element, for example, the repetition or generation of fields from the decoded frame. In other words, decoder 50 can provide one or more repetitions of the corresponding image to the output buffer. It should be noted that, according to embodiments of this disclosure, this may only be applicable to certain situations.

[0228] 2.1 Different frame rates for different output layers in the output layer set

[0229] As discussed, problems can arise when different frame rates exist in different output layers of the output layer set because all HRD SEI messages (BP, PT, and DUI) are globally applied to each corresponding AU, i.e., to every picture within an AU, without differentiation between layers, and therefore it is impossible to indicate different repetition patterns or different values ​​of elementalOutputPeriods with a single value of pt_display_elemental_periods_minus1. Note that this also applies to embodiments of other sub-aspects, such as sub-aspects 2.3 and 2.5.

[0230] Therefore, in one embodiment, a strobe flag (e.g., pt_display_elemental_periods_present_flag) is present in the PT SEI message to indicate whether elementalOutputPeriods is set within the PT SEI message.

[0231]

[0232] Therefore, according to one embodiment, decoder 50 can decode PT SEI messages for access unit 22. Decoder 50 can decode gating flags (e.g., pt_display_elemental_periods_present_flag) from the picture timing supplement enhancement information message, and if the gating flag is in a first state, decode picture output multiplication syntax elements (e.g., pt_display_elemental_periods_minus1) that reveal information about the multiplied picture outputs experienced by the predetermined access unit (e.g., pt_display_elemental_periods_minus1 is 0 or greater than 0), and if so, how many output pictures will be generated from the predetermined access unit (e.g., pt_display_elemental_periods_minus1 > 0).

[0233] The semantics of pt_display_elemental_periods_minus1 when it is inferred: When it does not exist, the value of pt_display_elemental_periods_minus1 is not inferred because it is not used. Instead, the information is obtained through other means as described in 2.3.

[0234] This applies unless the aspects in 2.5 are taken into account, in which case there is a constraint that the value of elementalOutputPeriods is signaled to be constrained to 1 (see 2.5).

[0235] In another embodiment, there is a bitstream constraint that requires pt_display_elemental_periods_present_flag to be equal to 0 if one of the following applies:

[0236] The frame rate of the output layer of OLS is different.

[0237] • The sps_field_seq_flag values ​​of all SPS referenced by the output layer of the OLS corresponding to the bitstream are different.

[0238] 2.2 Interaction between PT SEI messages and Frame Field Information SEI messages

[0239] As discussed, there may be a Frame Field Information (SEI) message indicating further information about how the frame should be output. For this purpose, in one embodiment, the information in the Picture Timing (PT) SEI message is used in a constrained manner along with the Frame Field Information (SEI) message. (Note that, according to the above embodiment, this only applies when a PT SEI message is present and the syntax element pt_display_elemental_periods_present_flag equals 1.)

[0240] When the `sps_field_seq_flag` in the SPS referenced by the VCL NAL unit of this layer is equal to 1 (the bitstream contains fields), `pt_display_elemental_periods_minus1` should be equal to 0 regardless of the value in the frame field information SEI message. Otherwise, if the `sps_field_seq_flag` in the SPS referenced by the VCL NAL unit of this layer is equal to 0 (the bitstream contains frames), and `display_fields_from_frame_flag` is equal to 0 (frames are not displayed as fields), and `fixed_pic_rate_within_cvs_flag[TemporalId]` is equal to 0, then the value of `pt_display_elemental_periods_minus1` should be equal to 0, i.e., in the case of no constant output frame rate and no field output from frames.

[0241] According to an embodiment of sub-aspect 2.2, encoder 10 is configured to encode a PT SEI message into video data stream 14 for a predetermined access unit 22 of the video data stream. This PT SEI message conveys information about a picture output timing unit for the predetermined access. Furthermore, encoder 10 is configured to encode for a picture sequence including pictures from the predetermined access unit, a sequence parameter set (e.g., SPS) indicating whether the pictures in the picture sequence represent fields or frames (e.g., progressive frames), and a fixed picture rate flag (e.g., in SPS or VPS) indicating whether the picture output of the video data stream involves a fixed picture rate. Encoder 10 is configured to set a picture output multiplication syntax element (e.g., pt_display_elemental_periods_minus1) to indicate that there is no multiplied picture output for the predetermined access unit.

[0242] In cases where the frame field syntax element indicates the image representation field of the image sequence, and / or

[0243] In cases where the frame field syntax element indicates that the picture representation frame of the picture sequence is a picture, the frame to field syntax element in the video data stream indicates that the frame is not displayed as a field (e.g., indicated in the frame field supplemental enhancement information message or inferred from its absence, or deduced in the absence of its absence), and the fixed picture rate flag indicates that the picture output does not involve a fixed picture rate.

[0244] According to one embodiment, video data stream 14 is a multi-layer video data stream that includes an output layer set (OLS) of one or more output layers (e.g., those layers whose pictures are output; there may be one or more reference layers that are not output but are used as reference layers). According to this embodiment, a picture timing supplemental enhancement message communicates information about picture output timing related to all output layers of the multi-layer video data stream having pictures encoded into predetermined access units. According to this embodiment, the picture sequence is of predetermined output layers, including pictures of predetermined output layers in predetermined access units, and a frame field syntax element (e.g., sps_field_seq_flag) indicates whether the pictures of the picture sequence of predetermined output layers represent fields or frames. According to this embodiment, a fixed picture rate flag (e.g., in SPS or VPS) indicates whether picture output for the output layer set involves a fixed picture rate related to the output layer set. According to this embodiment, the encoder is configured to set a picture output multiplication syntax element (e.g., pt_display_elemental_periods_minus1) to indicate that there is no multiplied picture output for the predetermined access unit.

[0245] In the case where the frame field syntax element indicates the image representation field of the image sequence of the predetermined output layer, and / or

[0246] In cases where the frame field syntax element indicates a picture representation frame of a picture sequence for a predetermined output layer, the frame to field syntax element (e.g., display_fields_from_frame_flag) indicates that the frame is not displayed as a field (e.g., indicated in a frame field supplemental enhancement message or inferred from its absence, or deduced in the absence of its absence), and the fixed picture rate flag indicates that the picture output for the output layer set does not involve a fixed picture rate.

[0247] According to one embodiment, encoder 10 is configured to encode an FFI SEI message, which conveys information about the frame field structure for the predetermined access unit and includes frame-to-field syntax elements (e.g., display_fields_from_frame_flag), into video data stream 14 for a predetermined access unit of the video data stream.

[0248] According to one embodiment, video data stream 14 is a multi-layer video data stream that includes an output layer set (OLS) of one or more output layers (e.g., those layers whose pictures are output; there may be one or more reference layers that are not output but are used as reference layers). According to this embodiment, a picture timing supplementation enhancement message relates to all output layers of the multi-layer video data stream having pictures encoded into a predetermined access unit and conveys information about picture output timing. According to this embodiment, a frame field supplementation enhancement message is specific to a predetermined output layer of the multi-layer video data stream and conveys information about the frame field structure associated with the predetermined output layer for the predetermined access unit. According to this embodiment, the picture sequence is of the predetermined output layer, including pictures of the predetermined output layer in the predetermined access unit, and a frame field syntax element (e.g., sps_field_seq_flag) indicates whether the pictures of the picture sequence of the predetermined output layer represent fields or frames. According to this embodiment, a fixed picture rate flag (e.g., in SPS or VPS) indicates whether picture output for the output layer set involves a fixed picture rate associated with the output layer set. According to this embodiment, the encoder is configured to set the image output multiplication syntax element (e.g., pt_display_elemental_periods_minus1) to indicate that there is no multiplied image output for a predetermined access unit when the following condition is met.

[0249] In the case where the frame field syntax element indicates the image representation field of the image sequence of the predetermined output layer, and / or

[0250] In cases where the frame-to-field syntax element indicates a picture representation frame of a picture sequence for a predetermined output layer, the frame-to-field syntax element indicates that the frame is not displayed as a field (e.g., indicating supplementary enhancement information messages in the frame field or inferred from their absence, or deduced in the absence of their absence), and the fixed picture rate flag indicates that the picture output for the set of output layers does not involve a fixed picture rate.

[0251] 2.3 Frame or field repetition without PT SEI message

[0252] In another embodiment, the third problem listed above (frame / field duplication without PT SEI messages) is addressed through an external means (e.g., an API) for elementalOutputPeriods, such that it is not only 1 when PT is absent or when a frame field information SEI message is present. It also addresses the problem indicated in 1), because when the output frame rates of different layers are different, the PT SEI message lacks information related to elementalOutputPeriods, while the frame field information SEI, as part of the SEI message for each layer, provides this information.

[0253] `elemental_duration_in_tc_minus1[i]` incremented by 1 (if present) specifies the time interval in clock ticks between pixels of consecutive images whose HRD output time is specified in the following output order when Htid equals i. The value of `elemental_duration_in_tc_minus1[i]` should be in the range of 0 to 2047 (inclusive).

[0254] For a CVS containing image n, the value of the variable DpbOutputElementalInterval[n] is specified as follows when Htid equals i and fixed_pic_rate_general_flag[i] equals 1, and image n is the output image rather than the last image in the output bitstream (in output order):

[0255] DpbOutputElementalInterval[n]=DpbOutputInterval[n]÷elementalOutputPeriods (113)

[0256] Where DpbOutputInterval[n] is specified in formula C.16, and elementalOutputPeriods is specified as follows:

[0257] - If a PT SEI message exists for image n, then pt_display_elemental_periods_present_flag equals 1, and elementalOutputPeriods equals the value of pt_display_elemental_periods_minus1+1.

[0258] - If an external method is provided, then elementalOutputPeriods is set to the value provided via the external method.

[0259] Otherwise (no external method is provided to set the value of elementalOutputPeriods), if a Frame Field Information (SEI) message is provided for a layer with a predefined index, the value of elementalOutputPeriods is set to display_elemental_periods_minus1+1. (Note that throughout the description, display_elemental_periods_minus1 may be another name for the aforementioned ffi_display_elemental_periods_minus1 syntax element.)

[0260] Otherwise, elementalOutputPeriods equals 1.

[0261] Layers with predefined indices used, such as those with the highest frame rate, are identified using one of the following methods:

[0262] • Indicated via additional signaling in HRD SEI, VPS / SPS, or other means (e.g., fixed_pic_rate_layer_index), or

[0263] • The layer containing the highest sublayer identifier value (temporal_id), or

[0264] The display_elemental_periods_minus1 value in the frame field SEI message differs from the pt_display_elemental_periods_minus1 value in the applicable PT SEI message.

[0265] As an alternative to using the layer index to identify which frame field information SEI message to use to determine elementalOuputPeriods, consider one of the following:

[0266] 1) Otherwise, if a Frame Field Information (SEI) message is provided for a layer that has output pictures present in AU n and the next AU (i.e., the AU containing nextPicInOutputOrder) in the order of output, the value of elementalOutputPeriods is set to display_elemental_periods_minus1+1.

[0267] 2) Otherwise, if a Frame Field Information (SEI) message is provided for the output image existing in the AU, the value of elementalOutputPeriods is set to the lowest value of display_elemental_periods_minus1+1 among all output layers.

[0268] Therefore, according to one embodiment of sub-aspect 2.3, decoder 50 is configured to derive the number of picture outputs for a predetermined access unit (e.g., the currently decoded access unit) of the video data stream based on one or more of the following criteria: if the number of picture outputs for the predetermined access unit is provided via the decoder's API, then the number of picture outputs provided via the API is adopted; and / or if a frame field supplemental enhancement information message is present in the video data stream, the frame field supplemental enhancement information message conveying information about the frame field structure for the predetermined access unit and including a further picture output multiplication syntax element (e.g., display_elemental_periods_minus1), then the further picture output multiplication syntax element is decoded from the frame field supplemental enhancement information message, and the number of picture outputs for the predetermined access unit is set according to the further picture output multiplication syntax element.

[0269] It's important to note that, generally, image output multiplication syntax elements can be represented by `pt_display_elemental_periods_minus1+1` or simply as `pt_display_elemental_periods_minus1`, because subtracting one before encoding is merely a choice of symbolization scheme. In other words, the value of an image output multiplication syntax element can correspond to the number of image outputs, or it can correspond to the number of image outputs minus 1. This also applies to other image output multiplication syntax elements.

[0270] According to one embodiment, the video data stream 14 is a multi-layer video data stream and includes PTSEI messages, as previously described, which convey information about the timing of image output in relation to all output layers of the multi-layer video data stream having images encoded into predetermined access units.

[0271] According to one embodiment, the above standard set further includes the following: if a PT SEI message containing a picture output multiplication syntax element (e.g., pt_display_elemental_periods_minus1) exists in the video data stream, the picture output multiplication syntax element is decoded from the PT SEI message, and the number of picture outputs for a predetermined access unit is set according to the picture output multiplication syntax element.

[0272] As mentioned earlier, PT SEI messages can involve all pictures in a predetermined access unit, and FFI SEI messages can be messages encoded into a predetermined access unit that convey information about the frame field structure of the picture in a predetermined output layer.

[0273] For example, when setting the number of image outputs for a predetermined access unit based on the further image output multiplication syntax element, decoder 50 may determine the number of image outputs for a predetermined output layer based on the further image output multiplication syntax element and use the further image output multiplication syntax element to determine the inter-frame output image interval for the predetermined access unit (e.g., the variable DpbOutputElementalInterval introduced above). According to one embodiment, if the predetermined output layer has images encoded into the predetermined access unit and access units that follow immediately after it in output order, decoder 50 performs the selection of the number of outputs.

[0274] According to one embodiment, decoder 50 can set the number of picture outputs for a predetermined access unit according to one or more of the following criteria based on the further picture output multiplication syntax element: According to the first criterion, if the predetermined output layer has pictures encoded into the predetermined access unit and the access units that follow in the output order, the number of picture outputs for the predetermined output layer is set according to the further picture output multiplication syntax element, and the further picture output multiplication syntax element is used to determine the inter-frame output picture interval (DpbOutputElementalInterval) for the predetermined access unit (for example, Alternative 1 in the above alternatives uses the layer index to identify which frame field information SEI message to use to determine elementalOutputPeriods). According to the second standard, if any other output layer has a picture encoded to the predetermined output layer and that other output layer has a frame field supplementation enhancement information message with a further picture output multiplication syntax element, the number of picture outputs for the predetermined output layer is set according to the further picture output multiplication syntax element, and the inter-frame output picture interval for the predetermined access unit is determined using the smaller of the further picture output multiplication syntax element and the further picture output multiplication syntax element (for example, Alternative 2 in the above alternatives uses the layer index to identify which frame field information SEI message to use to determine elementalOuputPeriods).

[0275] According to one embodiment, more than one output layer has images encoded into a predetermined output layer and includes frame field supplementation enhancement information messages, and the further image output multiplication syntax element of the predetermined output layer is minimized. According to this embodiment, when setting the number of image outputs for a predetermined access unit based on the further image output multiplication syntax element, the decoder 50 determines the number of image outputs for the predetermined output layer based on the further image output multiplication syntax element and uses the further image output multiplication syntax element to determine the inter-frame output image interval (DpbOutputElementalInterval) for the predetermined access unit.

[0276] According to one embodiment, if the number of image outputs for a predetermined access unit is set based on the image output multiplication syntax element, the decoder 50 sets the number of image outputs identically for all output layers and uses the image output multiplication syntax element to determine the inter-frame output image interval (DpbOutputElementalInterval) for the predetermined access unit. If the number of image outputs for a predetermined access unit is set based on a further image output multiplication syntax element, the decoder 50 sets the number of image outputs for the predetermined output layer based on the further image output multiplication syntax element and uses the further image output multiplication syntax element to determine the inter-frame output image interval (DpbOutputElementalInterval) for the predetermined access unit.

[0277] According to one embodiment, decoder 50 determines the predetermined output layer based on one of the following:

[0278] Based on signaling in multi-layer video data streams (e.g., specifically instructing a predetermined output layer),

[0279] As the output layer with the highest temporal sublayer (e.g., the decoder determines the sublayers belonging to each output layer and indicates the output layer with the highest temporal layer (in terms of hierarchy, i.e., the output layer that other temporal layers of the output layer do not depend on, as the highest temporal layer), or

[0280] As an output layer of the output layer set, for this output layer, the image output multiplication syntax elements are different from the image output multiplication syntax elements.

[0281] According to one embodiment, encoder 10 can provide signaling (e.g., specifically indicating a predetermined output layer) for a multi-layer video data stream. Alternatively, encoder 10 can select the output layer with the highest temporal sublayer (e.g., the decoder determines the sublayers belonging to each output layer and indicates the output layer with the highest temporal layer (in terms of hierarchy, i.e., the output layer on which other temporal layers of the output layer do not depend, as the highest temporal layer) as the predetermined output layer. Alternatively, encoder 10 can select the predetermined output layer as the output layer of a set of output layers, for which the further picture output multiplication syntax elements differ from the picture output multiplication syntax elements.

[0282] In an alternative to the above embodiments of sub-aspect 2.3, other embodiments address the problem through bitstream constraints, for example, regarding... Figure 11 As described.

[0283] Figure 11 The encoder 10 is illustrated according to an embodiment of sub-aspect 2.3. Figure 11 The encoder 10 may optionally correspond to Figure 1The encoder 10. The video bitstream 14 according to this embodiment can be a single-layer video bitstream or a multi-layer video bitstream, for example, as per [reference to...]. Figure 1 As described in Part 0, the video bitstream 14 has encoded a series of access units 22 therein. Figure 11 Reference marker 22* is used for reference. For example, the predetermined access unit is the currently encoded access unit. The encoder 10 according to this embodiment encodes PT SEI messages into the video bitstream 14 for the predetermined access unit 22*, such as the PT SEI messages previously described in this section. The PT SEI message conveys information about the timing of the picture output for the predetermined access unit 22*. The PT SEI message 73 includes a picture output multiplication syntax element 74, hereinafter also referred to as the PT multiplication indicator 73. For example, the PT multiplication indicator 73 can indicate the number of picture outputs for the picture of the predetermined access unit 22*, as previously described.

[0284] According to this embodiment, encoder 10 also encodes an FFI SEI message 83 into video data stream 14, which conveys information about the frame field structure for predetermined access unit 22*, such as the FFI SEI message described in detail previously in this section. FFI SEI message 83 includes a further picture output multiplication syntax element 84, which may also be referred to hereinafter as an FF multiplication indicator 84. As previously described, FF multiplication indicator 83 can indicate the number of picture outputs for one of the pictures for predetermined access unit 22*. For example, FFI SEI message 83 can reference one of the layers, such as one of the output layers of video bitstream 14.

[0285] The PT multiplication indicator 74 and FF multiplication indicator 84 can be encoded into the video bitstream using a symbolic scheme. For the purpose of encoding perspective syntax elements, the actual value of the corresponding syntax element can be derived by subtracting one from the number of picture outputs represented by the corresponding syntax element. In other words, in the example, the actual values ​​of the PT multiplication indicator 74 and FF multiplication indicator 84 written to the video bitstream 14 may differ from the values ​​represented by the PT multiplication indicator 74 and FF multiplication indicator 84, for example, by a difference of 1. However, the values ​​of the PT multiplication indicator 74 and FF multiplication indicator 84 should be understood as representing the actual number of picture outputs.

[0286] according to Figure 11 In one embodiment, the PT multiplication indicator is equal to or less than the FF multiplication indicator.

[0287] According to an embodiment, the information in the FFI SEI 83 is specific to a layer of the video bitstream 14. For example, the video bitstream 14 may include an FFI SEI message 84 for each output layer of the video bitstream 14. Alternatively, the FFI SEI may be provided at the access unit level G, an FFI SEI message 83 may be provided for a predetermined access unit 22*, and the FFI SEI message 83 includes a corresponding multiplication indicator 84 for each of one or more output layers having a picture encoded into the predetermined access unit 22*. According to this embodiment, the PT multiplication indicator 84 is equal to or less than the FF multiplication indicator 84 of all one or more output layers. For example, an output layer may represent a layer whose picture is considered by the decoder 50 for output, for example, as described in the introductory section of Part 2.

[0288] For example, video bitstream 14 is a multi-layer video bitstream, and encoder 10 provides each FFI SEI message 83 for one or more output layers of video bitstream 14, thereby providing one or more FFI SEI messages 83, each of which includes a corresponding FF multiplication indicator 84. Each of an FFI SEI message 83 may reference one of the layers, where one or more layers may be indicated as the output layer of the OLS indicated in video bitstream 14. According to this embodiment, all FF multiplication indicators 84 signaled in the corresponding FFI SEI messages 83 signaled for the output layer are greater than or equal to PT multiplication indicators 74. Therefore, the minimum value exceeding FF multiplication indicators 84 is greater than or equal to PT multiplication indicators 74.

[0289] According to one embodiment, the PT multiplication indicator 74 for the predetermined access unit 22* is equal to the minimum value of the FF multiplication indicator 84 of all FFI SEI messages in the output layer of the predetermined access unit 22*.

[0290] In other words, alternatively, pt_display_elemental_periods_minus1 in the PT SEI message applied to the AU is equal to the minimum value of display_elemental_periods_minus1 in the frame field information SEI messages of all output layers in the AU.

[0291] According to one embodiment, the FF multiplication indicator 84 is an integer multiple of the PT multiplication indicator 74. Note that this constraint is particularly effective for the actual value of the number of image outputs represented by the corresponding syntax elements.

[0292] For example, by setting the cell output period to PT_display_elemental_periods_minus1+1 and the display cell period to display_elemental_periods_minus1+1, decoder 50 can derive the aforementioned ElementalOutputPeriods and DisplayElementalPeriods. In this case, the above constraints can be applied to the variable cell output period and the display cell period; that is, the display cell period can be an integer multiple of the cell output period.

[0293] According to another embodiment, the video bitstream 14 is a multi-layer video data stream, and according to it, the PT SEI message 83 refers to all output layers of the multi-layer video bitstream 14, and according to it, the FFI SEI message is associated with a predetermined output layer, and therefore with the picture of the predetermined output layer encoded into the predetermined access unit 22*. The encoder 10 is configured to encode the PT multiplication syntax element 74 and the FF multiplication indicator 84 such that the FF multiplication indicator 84 is x times the PT multiplication indicator 74, where x is the distance between the predetermined access unit 22* and the access unit 22 before or after the access unit 22, which has the picture of the predetermined output layer encoded therein.

[0294] In other words, alternatively, the pt_display_elemental_periods_minus1 in the PT SEI message applied to the AU and the display_elemental_periods_minus1 in the SEI message applied to the frame field information of each picture in each layer of the AU do not need to be the same, but there are bitstream constraints as follows:

[0295] For each layer, let picA and picB be two consecutive output images, and let AuA and AuB be the nth and mth output AUs in the output order, respectively. display_elemental_periods_minus1 = ((mn)*(pt_display_elemental_periods_minus1+1))-1.

[0296] 2.4 Derivation of Frame Rate Periodicity

[0297] According to an embodiment of this sub-aspect, the constant output frame rate across CVS is derived rather than signaled, because otherwise it would require SPS rewriting after splicing. In other words, instead of signaling whether the output frame rate across CVS is constant, this information can be derived, for example, through decoder 50.

[0298] In other words, in another embodiment, the fourth problem listed above is solved as follows: Instead of using the signal transmission to determine whether a constant frame rate (fixed image rate) is maintained after the stitching point, the property is derived as follows.

[0299] For a CVS containing image n, when Htid equals i and fixed_pic_rate_within_cvs_flag[i] equals 1, and image n is the output image rather than the last image in the output bitstream (in output order), the value calculated for DpbOutputElementalInterval[n] should be equal to ClockTick*(elemental_duration_in_tc_minus1[i]+1), where ClockTick is specified as in Equation C.1 (using the value of ClockTick of the CVS containing image n) if one of the following conditions is true for the image nextPicInOutputOrder specified in Equation C.16:

[0300] The image nextPicInOutputOrder is in the same CVS as image n.

[0301] - The image nextPicInOutputOrder is in a different CVS and fixed_pic_rate_within_cvs_flag[i] is equal to 1 in the CVS containing the image nextPicInOutputOrder, the value of ClockTick is the same for both CVSs, and the value of elemental_duration_in_tc_minus1[i] is the same for both CVSs, and one or more of the following conditions are true:

[0302] - Same GOP size

[0303] - Same DPB parameters

[0304] The reordering parameters within the -DPB parameters are the same.

[0305] -nextPicInOutputOrder does not mean no output image.

[0306] - No RASL image is associated with nextPicInOutputOrder (it's CRA).

[0307] - The appended syntax element indicates the output delay of the first AU in CVS (described in the following aspects of {REF_Ref42174639\r\h\*MERGEFORMAT}).

[0308] -nextPicInOutputOrder has a value of 0 for NoOutputOfPriorPicsFlag, which is set to the ph_no_output_of_prior_pics_flag in the image header of nextPicInOutputOrder. Note that this parameter indicates at the CVS boundary that the previous image in the DPB of the previous CVS should not be output.

[0309] Information regarding Group of Pictures (GOP) size, DPB parameters, and reordering is provided below. Figure 12 The diagram below image 26 shows the first number corresponding to the decoding time and the second number corresponding to the output time (i.e., the numbers below the image are given in the form of "decoding time - output time"). As can be seen, the difference between the output time and the decoding time varies depending on the GOP size. This can be part of the DPB parameters or some reordering information added to the bitstream. For example, for GOP4, the value is 2, while for GOP8, the value is 3.

[0310] Therefore, according to an embodiment of this sub-aspect, the video data stream 14 is a concatenation of encoded video sequences, and the encoder 10 is configured to encode a set of parameters for each encoded video sequence 20 of the video data stream 14, the set of parameters including a fixed picture rate flag (e.g., fixed_pic_rate_within_cvs_flag) indicating whether the picture output involves a fixed picture rate within the corresponding encoded video sequence 20, and if the fixed picture rate flag indicates that the picture output involves a fixed picture rate within the corresponding encoded video sequence 20, also including a pixel output picture duration syntax element (e.g., elemental_duration_in_tc_minus1[i]). According to this embodiment, the encoder 10 is configured to signal in the data stream 14 via one or more continuity detectability syntax elements if all of the following apply, picture rate continuity is detectable, to be applied to the transition from the first encoded video sequence to the second encoded video sequence:

[0311] - The fixed picture rate flag for the first and second coded video sequences indicates that the picture output for the output layer set involves a fixed picture rate in the first and second coded video sequences.

[0312] - The pixel output image duration syntax elements are the same for the first and second coded video sequences.

[0313] - One or more of the following conditions apply, where the set of conditions includes one or more of the following:

[0314] The first and second coded video sequences are identical in GOP size.

[0315] The first and second encoded video sequences are consistent in the reordering syntax element (e.g., max_num_reorder_pics), which indicates the maximum allowed number of output pictures that can be decoded before and output in the order of the other output picture.

[0316] The first and second coded video sequences are consistent with each other on the DPB parameters (e.g., indicating the DPB image removal time).

[0317] • The second coded video sequence does not begin with an IRAP associated with the RASL image (e.g., does not begin with a CRA).

[0318] The first and second coded video sequences are consistent on the output delay syntax element transmitted in the video data stream. This output delay syntax element indicates the output delay of the first access unit of the first and second coded video sequences (e.g., the first AU in CVS1 and the first AU in CVS2 have the same output delay related to their decoding time. For example, in picture time, both have syntax element = 3, which indicates a 3-picture time delay from decoding to output).

[0319] • The first AU in the second coded video sequence is not a non-output image, and

[0320] • The first AU in the second coded video sequence does not indicate that previous images from the first coded video sequence were not output.

[0321] The issue of no-output images is explained in more detail in Section 2.6 of this document. The aspect related to RASL images is related to no-output images because such RASL images are not output when associated with the first AU of the CVS, and therefore can be considered as no-output images.

[0322] 2.5 Derivation of output time, for example, PT SEI messages not existing or not being used.

[0323] The fifth problem listed above can be solved according to an embodiment of this sub-aspect.

[0324] According to a first embodiment of sub-aspect 2.5, decoder 50 is configured to decode a parameter set for a predetermined encoded video sequence of a multi-layer video data stream. The parameter set includes a fixed picture rate flag indicating whether picture output involves a fixed picture rate within the predetermined encoded video sequence, and if the fixed picture rate flag indicates that picture output involves a fixed picture rate within the predetermined encoded video sequence, it also includes a pixel output picture duration syntax element. According to this embodiment, decoder 50 is configured to determine an output delay (picture output time of the first picture) for the predetermined encoded video sequence based on a product of a first factor determined by the pixel output picture duration syntax element and a second factor determined using a reordering syntax element in the DPB parameters. The reordering syntax element indicates the maximum allowed number of pictures in the output picture set that can precede any picture in the OLS in decoding order and follow that picture in output order. Alternatively, decoder 50 is configured to determine an output delay (picture output time of the first picture) for the predetermined encoded video sequence based on a product of a first factor determined by the pixel output picture duration syntax element and a second factor indicated by a delay syntax element in the video data stream.

[0325] According to the first embodiment, encoder 10 is configured to encode a parameter set (e.g., a VPS parameter set or an SPS parameter set with HRD and timing information) into the video data stream (14) for a predetermined encoded video sequence of a multi-layer video data stream. The parameter set includes a fixed picture rate flag, such as fixed_pic_rate_within_cvs_flag, which indicates whether the picture output involves a fixed picture rate within the predetermined encoded video sequence. If the fixed picture rate flag indicates that the picture output involves a fixed picture rate within the predetermined encoded video sequence, it also includes a pixel output picture duration syntax element, such as elemental_duration_in_tc_minus1. According to this embodiment, encoder 10 is configured to signal in the data stream, based on a product of a first factor determined by a pixel output image duration syntax element and a second factor determined using a reordering syntax element in the DPB parameters, one or more output delay computability syntax elements (e.g., elements currently indicating that the video data stream does not need or does not contain PT / BP-...SEI): the output delay (image output time of the first image) for a predetermined encoded video sequence is computable, and the reordering syntax element indicates the maximum allowed number of images in the output image set that can precede any image in the OLS in decoding order and follow that image in output order. According to this embodiment, encoder 10 is configured to signal in the data stream, based on a product of a first factor determined by a pixel output image duration syntax element and a second factor indicated by a delay syntax element in the video data stream (14), one or more output delay computability syntax elements (e.g., elements currently indicating that the video data stream does not need or does not contain PT / BP-...SEI): the output delay (image output time of the first image) for a predetermined encoded video sequence is computable.

[0326] In other words, the first embodiment according to sub-aspect 2.5 includes indicating in the bitstream that timing information can be derived without PT SEI and BP SEI, and deriving the output time for the first AU (e.g., the first AU of CVS 20), depending on the DPB parameters or additional parameters. Note that the Buffer Period (BP) SEI message and the Picture Timing (PT) SEI message contain timing information on when to remove the AU from the CPB and when to output the AU from the DPB. Several values ​​(e.g., for different highest time IDs present in the bitstream) can be used to derive when to decode the AU (removed from the CPB) and when to output it (from the DPB). Under some conditions explained below, the output time can be derived without the help of these SEI messages. The output time can be derived as one of two options:

[0327] - If the DPB parameter is used, the output time value is derived as ClockTick*(elemental_duration_in_tc_minus1[i]+1)*NumPics, where NumPics is the number of reordered pictures sent by signaling in the DPB parameter (max_num_reorder_pics), or the maximum allowed number of pictures in the OLS that can precede any picture in the OLS in decoding order and follow that picture in output order, plus the maximum number of pictures in the OLS that can precede any picture in the OLS in output order and follow that picture in decoding order (max_num_reorder_pics+max_latency_increase_plus1), or

[0328] - Add additional signaling to the VPS or SPS associated with a fixed image rate, which indicates a given number of images NumPics and this syntax is used to calculate the value of the output time, which is derived as ClockTick*(elemental_duration_in_tc_minus1[i]+1)*NumPics.

[0329] Figure 13 An example of an encoder 10, video bitstream 14, and decoder 50 according to a second embodiment of sub-aspect 2.5 is illustrated. The encoder 10, video bitstream 14, and decoder 50 may optionally correspond to [the embodiment described in the second embodiment]. Figure 1 The encoder 10, video bitstream 14, and decoder 50. Furthermore, embodiments according to this sub-aspect may optionally include elements related to sub-aspect 2.3 (e.g., regarding...). Figure 11 The features and details described.

[0330] according to Figure 13In this embodiment, encoder 10 encodes a parameter set 93 of a predetermined encoded video sequence 20 into a video bitstream 14. That is, parameter set 93 is associated with one or more encoded video sequences 20 of video bitstream 14. Parameter set 93 includes a fixed picture rate flag 94, which may correspond, for example, to the fixed_pic_rate_within_CVS_flag described herein. Fixed picture rate flag 94 indicates whether picture output involves a fixed picture rate within the predetermined encoded video sequence 20. If fixed picture rate flag 93 indicates that picture output involves a fixed picture rate within the predetermined encoded video sequence 20, parameter set 93 also includes a pixel output picture duration syntax element 96, which may correspond, for example, to the elemental_duration_in_tc_minus1 described herein. For example, pixel output picture duration syntax element 96 may indicate the duration of a picture output interval, such as the duration of a picture output interval 63 for a single picture.

[0331] For example, such as Figure 13 As illustrated in the diagram, each access unit 22 of the video bitstream 14 can be associated with a pixel image output time 36. The pixel image output time 36 can represent a time instance in which the image of the corresponding access unit 22 will be output by the decoder 50, for example, when the decoder 50 provides the corresponding access unit (i.e., its image) to the output buffer. The access unit output interval 37 can indicate the time interval between the pixel image output times 36 of consecutive access units, and can, for example, correspond to the time interval related to… Figure 8 and Figure 10 The described access unit output interval is 61.

[0332] according to Figure 13 In one embodiment, encoder 10 encodes one or more syntax elements 66 into video bitstream 14, and if one or more syntax elements 66 have a first state, i.e., the access unit of the encoded video sequence 20 referenced by one or more syntax elements 66 has a first state, decoder 50 can infer that the picture of the access unit does not undergo multiplication output (e.g., infer pt_display_elemental_periods_minus1 is 0 or elementalOutputs is 1). Therefore, decoder 50 can derive pixel output picture time 36 using pixel output picture duration syntax element 96, for example, by setting the access unit output intervals 37, 61 to a value equal to the picture output interval indicated by the pixel output picture duration syntax element 96. Thus, decoder 50 can determine pixel output picture time 36 without indicating the number of repetitions (e.g., without PT SEI), which may therefore be omitted in video bitstream 14.

[0333] For example, generally, decoder 50 can derive pixel image output time 36 based on the pixel output image duration indicated by pixel output image duration syntax element 96 and the number of repetitions of the image of the corresponding access unit indicated, for example, by image output multiplication syntax element or further image output multiplication syntax element (see Part 2.3). To this end, decoder 50 can set the duration of the image output interval (e.g., image output interval 63), for example, represented by the variable DpbOutputElementalInterval, to be equal to the value indicated by elemental_duration_in_tc_minus1 (where the indicated value can correspond to the value actually written to the bitstream plus 1, and can optionally be scaled by a clock tick duration, which can optionally be signaled in the video bitstream 14, for example, by the syntax element ClockTick, e.g., DpbOutputElementalInterval = ClockTick). kTick*(elemental_duration_in_tc_minus1+1).) and by using DpbOutputElementalInterval and the number of repetitions of the corresponding image (e.g., equation (113) in the definition of elemental_duration_in_tc_minus1 given in the introduction of Part 2, where the variable elementalOutputs represents the number of repetitions) to deduce the variable DpbOutputInterval (e.g., the duration of access unit output interval 61), which is deduced to be 1 in the case that syntax element 66 has a first state. Therefore, when one or more syntax elements 66 are in a first state, decoder 50 can set the duration of the access unit output interval 37 (or 61) (e.g., DpbOutputInterval) to be equal to the value indicated by elemental_duration_in_tc_minus1 (e.g., a value derived by adding 1 to the actual value and / or by multiplying by the clock tick duration, e.g., DpbOutputInterval = ClockTick * (elemental_duration_in_tc_minus1 + 1)). In other words, in this case, where one or more syntax elements 66 are in a first state, decoder 50 can interpret the pixel image duration syntax element 96 as a reference to the duration of the access unit output intervals 37, 61.

[0334] For example, the parameter set 93, which includes a fixed picture rate flag 94 and an optional pixel output picture duration syntax element 96, can be a sequence parameter set (SPS) that can be globally associated with the encoded video sequence 20. The SPS includes HRD and timing information. Alternatively, the parameter set 93 can be a video parameter set (VPS) that can be globally associated with the video bitstream 14.

[0335] according to Figure 13 In one embodiment, encoder 10 encodes one or more syntax elements 66 into video bitstream 14. Encoder 10 encodes video bitstream 14 such that if one or more syntax elements 66 have a first state, then for each access unit 22 of the encoded video sequence 20 (or video bitstream 14), the corresponding access unit 22 can be inferred to be not subject to multiplication output. For example, the introduction in Part 2 and the variable pixel output period described in Part 2.3 can be inferred to be one. Therefore, the pixel image output time 36 of the encoded video sequence 20 can be determined based on the pixel output image duration syntax element.

[0336] In other words, for example, if one or more syntax elements 66 have a first state, the decoder 50 can infer that the access unit 22 does not undergo multiplication output and thus the pixel image output time of the predetermined access unit can be determined by adding the pixel image output time 36 of the preceding access unit to the pixel image output time of the predetermined access unit, which is signaled by the pixel output image duration syntax element 96.

[0337] For example, if one or more syntax elements 66 have a first state, then if any picture output multiplication syntax element is provided—such as the PT multiplication indicator 74 (or pt_display_elemental_periods_minus1)—the encoder 10 can provide a picture output multiplication syntax element such that it signals a single picture output, i.e., no multiplied picture output. Therefore, the decoder 50 can infer that, with one or more syntax elements 66 in the first state, the picture output multiplication syntax element indicates no multiplied output, i.e., a single output.

[0338] According to an embodiment, one or more syntax elements 66 indicate one or more HRD parameters (or bitstream consistency parameters) in the video bitstream 14 that include (or do not include) bitstream portions referencing the video bitstream 14, for example, general_nal_hrd_params_present_flag = 0, and the HRD parameters (or bitstream consistency parameters) referencing the coding layer of the video bitstream 14, for example, general_vcl_hrd_params_present_flag = 0.

[0339] According to an embodiment, one or more syntax elements 66 may include one or more of a first syntax element and a second syntax element, each syntax element indicating that the video bitstream 14 does not contain (or does contain) an encoded picture buffer (CPB) and bitrate parameters for a corresponding operating mode for a hypothetical reference decoder, such as NAL operations (e.g., operating modes that may include SEI NAL units and headers on top of VCL data) and VCL operations (e.g., operating modes specifically considered for encoded video data, such as VCL NAL units). For example, one or more syntax elements 66 may include one or both of general_nal_hrd_params_present_flag and general_vcl_hrd_params_present_flag.

[0340] For example, the first state could be a state in which one or more syntax elements indicate that the video bitstream 14 does not contain the encoded picture buffer (CPB) and the bitrate parameters for the NAL and VCL operation modes for the hypothetical reference decoder, such as general_nal_hrd_params_present_flag = 0 and general_vcl_hrd_params_present_flag = 0.

[0341] For example, one or more syntax elements 66 can be encoded into one or more parameter sets of the video bitstream 14.

[0342] according to Figure 13 In an example of an embodiment, if one or more syntax elements 66 have a second state, such as if one or more syntax elements 66 indicate that the video bitstream 14 indicates that it contains one or more of the parameters mentioned above, then the encoder 10 may, for each access unit 22 of the video bitstream 14 or the encoded video sequence 20, output a picture multiplication syntax element (e.g., as per the context of the multiplication syntax element). Figure 11 The described image output multiplication syntax element 74) is encoded into the PT SEI message 73 of the video bitstream 14, for example, as described in the following... Figure 11 As described. (Regarding...) Figure 11As described, the image output multiplication syntax element 74 can reveal information about whether the corresponding access unit 22 (i.e., the access unit referenced by the PT SEI message 73) undergoes multiplication output, and if so, how many sequential output images to be generated from the corresponding access unit 22. According to this example, the pixel image output time 36 can be determined based on the pixel output image duration syntax element 96 and the image output multiplication syntax element. For example, if one or more syntax elements 66 have a second state, the decoder 50 can decode the image output multiplication syntax element for each access unit and determine the pixel image output time 36 for the access unit of the encoded video sequence 20 based on the pixel output image duration syntax element 96 and the image output multiplication syntax element. For example, using the variables elementalOutputs and equation (113) mentioned above, the decoder 50 can, for each access unit 22, multiply the output duration indicated by the pixel output image duration syntax element by the number of repetitions indicated by the image output multiplication syntax element, as described above regarding the case where the first syntax element has a first state, but without inferring a repetition of 1.

[0343] In other words, in another embodiment, for example Figure 13 In the embodiments described, where there are indications in the bitstream, for example in the VPS or SPS, that there is no PT SEI and BP SEI in the bitstream and / or they are not required, and there is information that the frame field SEI message is absent or unnecessary, then repetition does not need to be included or taken into account; that is, the decoder can deduce elementalOutputPeriods as 1. This can be accomplished by the mentioned syntax element, which indicates that timing can be deduced in the absence of PT SEI and BP SEI messages or alternatively using bitstream constraints. Note that in this case, pt_display_elemental_periods_minus1 can be deduced as 0 when the PT SEI message is absent or when the syntax element is absent.

[0344] Bitstream constraints can be added as `no_timing_infomation_sei_message_needed_flag` or, for example, the described operation can be adjusted to the case where `general_nal_hrd_params_present_flag` and `general_vcl_hrd_params_present_flag`, which indicate the presence of CPB and the bitrate parameters for NAL or VCL operations, are both equal to 0. In the latter case, the presence of PT SEI or BP SEI messages is not required, and the operation without them can be simply performed by deriving `elementalOutputPeriods` to 1 and using `elemental_duration_in_tc_minus1[I]` as the output picture rate used to derive the output time.

[0345] Therefore, sub-aspect 2.5 provides a concept for determining the pixel image output time 36 in the absence of a PT SEI message. Thus, the PT SEI message does not necessarily have to be encoded into the video bitstream 14, thereby avoiding signaling overhead in the video bitstream 14.

[0346] 2.6 Processing of images without output and its impact on constant output frame rate

[0347] As discussed above, if some images are not output, deriving the output time can be complex when they are not present in the PT SEI message or when the PT SEI message is ignored.

[0348] When a PT SEI message is present, an image without input has an associated output time, but since the image is not output, such output time is simply ignored. Counting such an image decoded in the bitstream as "occupying" an output time slot will cause the distance between the two actual output images to no longer be equidistant, as shown below. Figure 14 As illustrated in the diagram. For example, in Figure 14 In the example, for instance, through layer dependency relationships, image 26* is indicated as an image with no output. In other words, Figure 8 The illustration shows an example of a non-isolated image without an output image.

[0349] According to a first embodiment of this sub-aspect, decoder 50 is configured to decode a set of parameters including a fixed picture rate flag, such as as described in Section 2.5, indicating whether picture output for the video data stream involves a fixed picture rate, and if the fixed picture rate flag indicates that picture output involves a fixed picture rate, also including a pixel output picture duration syntax element. According to this embodiment, decoder 50 decodes a picture output flag from the video data stream 14 for each picture, indicating whether the corresponding picture should be displayed. According to this embodiment, decoder 50 infers, in order of output, another picture indicated not to be output (e.g., in another picture) Figure 14 Images 26* and earlier (e.g., images from previous images) Figure 14 Image 26' will be repeatedly output.

[0350] In other words, in one embodiment, such as the example in the embodiment in the previous paragraph, a flag indicating whether an image is output, i.e., (ph_pic_output_flag), is considered to derive the output timing, i.e., the decoder is prepared to receive images that do not need to be displayed, i.e., there is no constant output from the decoder. In this case, a constant display rate is achieved, and there is a bitstream constraint, i.e., the preceding images in the output order need to be compensated for by repetition.

[0351] According to a second embodiment, encoder 10 is configured to encode a set of parameters into video data stream 14, the set of parameters including a fixed picture rate flag, and, if the fixed picture rate flag indicates that picture output involves a fixed picture rate, also including a pixel output picture duration syntax element. According to this embodiment, encoder 10 is configured to encode a picture output flag indicating whether the corresponding picture should be displayed into video data stream 14 for each picture. According to this embodiment, if the fixed picture rate flag indicates that picture output for video data stream 14 involves a fixed picture rate, encoder 10 sets the picture output flag for each picture to indicate that the corresponding picture will be displayed. Alternatively or additionally, if the fixed picture rate flag indicates that picture output for video data stream 14 involves a fixed picture rate, encoder 10 sets the picture output flag for each picture that is not the first picture of the encoded video sequence of video data stream 14 to indicate that the corresponding picture should be displayed. Alternatively or additionally, if the fixed picture rate flag indicates that the picture output for video data stream 14 involves a fixed picture rate, then encoder 10 sets the picture output flag for each picture that is not the first picture in the encoded video sequence of video data stream 14, or for each picture within the encoded video sequence of video data stream 14 and not exclusively preceding other pictures without output, to indicate that the corresponding picture should be displayed. Thus, for example, encoder 10 provides video bitstream 14 such that decoder can infer, in the output order, that pictures preceding another picture indicated as not to be output will undergo repeated output.

[0352] In the example, encoder 10 is configured to encode a flag into video data stream 14 that indicates whether the image output flag for each picture is set to indicate that the corresponding picture will be displayed. Alternatively or additionally, the flag indicates whether the image output flag for each picture that is not the first picture in the encoded video sequence of video data stream 14 is set to indicate that the corresponding picture will be displayed. Alternatively or additionally, the flag indicates that the image output flag for each picture that is not the first picture in the encoded video sequence of video data stream 14, or that is within the encoded video sequence of video data stream 14 and not exclusively preceding other pictures without output, is set to indicate that the corresponding picture should be displayed.

[0353] In the example, encoder 10 is configured to set a flag when a fixed picture rate flag indicates that the picture output for video data stream 14 involves a fixed picture rate.

[0354] In other words, in another embodiment, such as the example of the second embodiment, when the bitstream indicates the presence of a fixed image rate, there are bitstream constraints that prohibit the following:

[0355] • Non-output images within the bitstream, or

[0356] • At least one non-output image that is not the first AU in CVS, or

[0357] • Once an output image exists in CVS, no further output images can be created after that output image in CVS.

[0358] In another embodiment, a bitstream constraint is added, which indicates that no output image is a constraint, and as a broader concept, it applies not only to a fixed image rate.

[0359]

[0360] A value of 1 for `general_no_no_output_pics_constraint_flag` specifies that `ph_pic_output_flag` should be equal to 1. A value of 0 for `general_no_no_output_pics_constraint_flag` does not impose this constraint.

[0361] In addition, when using a fixed image rate, there are bitstream constraints that require setting constraint flags.

[0362] The requirement for bitstream consistency is that when fixed_pic_rate_general_flag[i] is equal to 1 for any value of i, general_no_no_output_pics_constraint_flag should be equal to 1.

[0363] 3. Other embodiments

[0364] In the preceding sections, although some aspects have been described as features within the context of the device, it is clear that such descriptions can also be considered descriptions of the corresponding features of the method. Similarly, although some aspects have been described as features within the context of the method, it is clear that such descriptions can also be considered descriptions of the corresponding features concerning the functionality of the device.

[0365] Some or all of the method steps may be performed by (or using) hardware devices, such as, for example, microprocessors, programmable computers, or electronic circuits. In some embodiments, one or more of the most important method steps may be performed by such devices.

[0366] The encoded image signal of the present invention can be stored on a digital storage medium or transmitted on a transmission medium such as a wireless transmission medium or a wired transmission medium such as the Internet.

[0367] Depending on certain implementation requirements, embodiments of the present invention may be implemented in hardware or software, or at least partially in hardware or at least partially in software. Implementation may be performed using a digital storage medium having electronically readable control signals stored thereon, such as a floppy disk, DVD, Blu-ray disc, CD, ROM, PROM, EPROM, EEPROM, or FLASH memory, which cooperates (or is capable of cooperating with) a programmable computer system to perform the corresponding methods. Therefore, the digital storage medium may be computer-readable.

[0368] Some embodiments of the invention include a data carrier having electronically readable control signals, which is capable of cooperating with a programmable computer system to perform one of the methods described herein.

[0369] Typically, embodiments of the present invention can be implemented as a computer program product having program code that, when run on a computer, is operable to perform one of the methods. The program code may, for example, be stored on a machine-readable medium.

[0370] Other embodiments include a computer program stored on a machine-readable medium for performing one of the methods described herein.

[0371] In other words, one embodiment of the method of the present invention is therefore a computer program having program code that, when run on a computer, performs one of the methods described herein.

[0372] Therefore, another embodiment of the method of the present invention is a data carrier (or digital storage medium or computer-readable medium) comprising a computer program recorded thereon for performing one of the methods described herein. Data carriers, digital storage media, or recording media are generally tangible and / or non-transitory.

[0373] Therefore, another embodiment of the method of the present invention represents a data stream or signal sequence for performing one of the methods described herein. The data stream or signal sequence may, for example, be configured to be transmitted via a data communication connection (e.g., via the Internet).

[0374] Another embodiment includes a processing component, such as a computer or programmable logic device, configured or adapted to perform one of the methods described herein.

[0375] Another embodiment includes a computer on which a computer program is installed for performing one of the methods described herein.

[0376] Another embodiment of the invention includes an apparatus or system configured to transmit (e.g., electronically or optically) a computer program for performing one of the methods described herein to a receiver. The receiver may be, for example, a computer, a mobile device, a memory device, etc. The apparatus or system may, for example, include a file server for transmitting the computer program to the receiver.

[0377] In some embodiments, a programmable logic device (e.g., a field-programmable gate array) may be used to perform some or all of the functionality of the methods described herein. In some embodiments, the field-programmable gate array may cooperate with a microprocessor to perform one of the methods described herein. Generally, the methods are preferably performed by any hardware device.

[0378] The apparatus described herein can be implemented using hardware devices, a computer, or a combination of hardware devices and a computer.

[0379] The methods described herein can be performed using hardware devices, computers, or a combination of hardware devices and computers.

[0380] As can be seen in the foregoing detailed description, various features have been grouped together in the examples for the purpose of simplification of this disclosure. This method of disclosure should not be construed as reflecting an intention to require more features than expressly recited in each claim. Rather, as reflected in the appended claims, the subject matter may comprise fewer features than in any single disclosed example. Therefore, the following claims are thus incorporated into the detailed description, each of which can stand alone as a separate example. While each claim can stand alone as a separate example, it should be noted that although dependent claims may refer to a specific combination with one or more other claims in the claim, other examples may also include combinations of dependent claims with the subject matter of each other dependent claim or combinations of each feature with other dependent or independent claims. Such combinations are presented herein unless it is expressly stated that a particular combination is not contemplated. Furthermore, even if a claim is not directly dependent on an independent claim, it is intended to include the features of that claim in any other independent claim.

[0381] The above embodiments are merely illustrative of the principles of this disclosure. It should be understood that modifications and variations of the arrangements and details described herein will be readily apparent to those skilled in the art. Therefore, the intent is limited only by the scope of the pending patent claims and not by the specific details presented in the description and explanation of the embodiments herein.

Claims

1. A decoder (50) for decoding a video data stream (14), the decoder (50) being configured to A predetermined encoded video sequence (20) for the video data stream (14). Decode the parameter set (93), which includes A fixed picture rate flag (94) indicates whether the picture output involves a fixed picture rate within the predetermined coded video sequence, and if the fixed picture rate flag indicates that the picture output involves a fixed picture rate within the predetermined coded video sequence, it also includes a pixel output picture duration syntax element (96), and Decode one or more syntax elements (66) from the video data stream (14), and If the one or more syntax elements (66) have a second state, For each access unit (22) of the video data stream (14), a picture output multiplication syntax element is decoded from the picture timing supplementation enhancement information message of the video data stream (14). This picture output multiplication syntax element reveals information about whether the corresponding access unit (22) has undergone multiplication output, and if so, how many sequential output pictures are generated from the corresponding access unit (22). The output delay for the predetermined encoded video sequence is determined based on the pixel output image duration syntax element and the image output multiplication syntax element; If the one or more syntax elements (66) have a first state, The pixel output image duration syntax element (96) is used to determine the pixel image output time (36) for the predetermined encoded video sequence. as well as For each access unit (22) of the encoded video sequence (20), the picture output multiplication syntax element (74) in the picture timing supplemental enhancement information message (73) of the video data stream (14) is inferred to indicate non-multiplicative output, the picture output multiplication syntax element (74) revealing information about whether the corresponding access unit (22) has undergone multiplicative output, and if so, how many sequential output pictures to generate from the corresponding access unit (22).

2. The decoder (50) according to claim 1, wherein The one or more syntax elements (66) indicate one or more of the following The video data stream (14) may or may not contain image periodic supplementation and enhancement information messages. The video data stream (14) may or may not require periodic supplementation of image enhancement information messages. The video data stream (14) may or may not contain frame field supplementation enhancement information messages. The video data stream (14) may or may not include frame field supplementation enhancement information messages. The video data stream (14) may or may not contain buffer period supplementation enhancement information messages. The video data stream (14) may or may not include supplemental enhancement information messages.

3. The decoder (50) according to claim 1, wherein The one or more syntax elements (66) indicate one or more of the following The video data stream (14) may or may not contain CPB and bitrate parameters for NAL operation. The video data stream (14) may or may not contain CPB and bit rate parameters for VCL operation.

4. The decoder (50) according to claim 1 is configured to decode the one or more syntax elements (66) from one or more parameter sets of the video data stream (14).

5. An encoder (10) for encoding a video data stream (14), the encoder (10) being configured to For the predetermined encoded video sequence (20) of the video data stream (14). The parameter set (93) is encoded into the video data stream (14), the parameter set (93) including A fixed picture rate flag (94) indicates whether the picture output involves a fixed picture rate within the predetermined coded video sequence, and if the fixed picture rate flag indicates that the picture output involves a fixed picture rate within the predetermined coded video sequence, it also includes a pixel output picture duration syntax element (96), and Encode one or more syntax elements (66) into the video data stream (14) so ​​that If the one or more syntax elements (66) have a second state, For each access unit (22) of the video data stream (14), the image output multiplication syntax element (74) is encoded into the image timing supplementation enhancement information message (73) of the video data stream (14), the image timing supplementation enhancement information message (73) revealing whether the corresponding access unit (22) undergoes multiplication output, and if so, how many sequential output images are generated from the corresponding access unit (22) such that The pixel image output time for the predetermined encoded video sequence is determined based on the pixel output image duration syntax element and the image output multiplication syntax element (36); The encoder (10) is configured as follows: If the one or more syntax elements (66) have a first state, The pixel output image duration syntax element (96) is used to determine the pixel image output time (36) for the predetermined encoded video sequence, and for each access unit (22) of the encoded video sequence (20), the corresponding access unit (22) can be inferred to be not subject to multiplication output.

6. The encoder (10) according to claim 5, wherein The one or more syntax elements (66) indicate one or more of the following The video data stream (14) may or may not contain image periodic supplementation and enhancement information messages. The video data stream (14) may or may not require periodic supplementation of image enhancement information messages. The video data stream (14) may or may not contain frame field supplementation enhancement information messages. The video data stream (14) may or may not include frame field supplementation enhancement information messages. The video data stream (14) may or may not contain buffer period supplementation enhancement information messages. The video data stream (14) may or may not contain additional enhancement information messages.

7. The encoder (10) according to claim 5, wherein The one or more syntax elements (66) indicate one or more of the following The video data stream (14) may or may not contain CPB and bitrate parameters for NAL operation. The video data stream (14) may or may not contain CPB and bit rate parameters for VCL operation.

8. The encoder (10) according to claim 5 is configured to encode the one or more syntax elements into one or more parameter sets of the video data stream (14).

9. A method for decoding (50) a video data stream (14), wherein the method comprises: For the predetermined encoded video sequence of the video data stream (14), Decode the parameter set, which includes A fixed picture rate flag, indicating whether the picture output involves a fixed picture rate within the predetermined coded video sequence, and if the fixed picture rate flag indicates that the picture output involves a fixed picture rate within the predetermined coded video sequence, then a pixel output picture duration syntax element is also included. Decode one or more syntax elements from the video data stream (14). If the one or more syntax elements (66) have a second state, For each access unit (22) of the video data stream (14), a picture output multiplication syntax element is decoded from the picture timing supplementation enhancement information message of the video data stream (14). This picture output multiplication syntax element reveals information about whether the corresponding access unit (22) has undergone multiplication output, and if so, how many sequential output pictures are generated from the corresponding access unit (22). The output delay for the predetermined encoded video sequence is determined based on the pixel output image duration syntax element and the image output multiplication syntax element; If the one or more syntax elements (66) have a first state, The pixel image output time (36) for the predetermined encoded video sequence is determined based on the pixel output image duration syntax element (96); and For each access unit (22) of the encoded video sequence (20), it is inferred that the corresponding access unit (22) is an output that does not undergo multiplication.

10. A method for encoding a video data stream (14), wherein the method includes: For the predetermined encoded video sequence (20) of the video data stream (14). The parameter set (93) is encoded into the video data stream (14), the parameter set (93) comprising: A fixed picture rate flag (94) indicates whether the picture output involves a fixed picture rate within the predetermined coded video sequence, and if the fixed picture rate flag indicates that the picture output involves a fixed picture rate within the predetermined coded video sequence, it also includes a pixel output picture duration syntax element (96), and Encode one or more syntax elements (66) into the video data stream (14) so ​​that If the one or more syntax elements (66) have a second state, For each access unit (22) of the video data stream (14), the image output multiplication syntax element (74) is encoded into the image timing supplementation enhancement information message (73) of the video data stream (14), the image timing supplementation enhancement information message (73) revealing whether the corresponding access unit (22) undergoes multiplication output, and if so, how many sequential output images are generated from the corresponding access unit (22) such that The pixel image output time for the predetermined encoded video sequence is determined based on the pixel output image duration syntax element and the image output multiplication syntax element (36); If the one or more syntax elements (66) have a first state, The pixel image output time (36) for the predetermined encoded video sequence is determined based on the pixel output image duration syntax element (96); and For each access unit (22) of the encoded video sequence (20), it is inferred that the corresponding access unit (22) is an output that does not undergo multiplication.