Codec method and video codec for decoding bit stream from picture
By introducing extensible nestable supplementary enhanced information messages into the video data stream, the shortcomings of video encoding and decoding in the prior art in parallel processing and HRD consistency management are solved, efficient encoding and decoding of multi-layer videos are realized, and sub-bit stream processing and buffer management are optimized.
Patent Information
- Application Number
- CN202080088792.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-12-20
- Filing Date
- 2020-12-18
- Publication Date
- 2025-08-29
- Estimated Expiration
- 2040-12-18
AI Technical Summary
The existing video encoding and decoding technologies have shortcomings in parallel processing capabilities and HRD consistency management, especially when scalable bitstream pruning and subbit stream processing, it is difficult to effectively support the timing and buffer management of multi-layer video encoding and decoding.
By introducing extensible nestable supplementary enhancement information messages into the video data stream, indicating the timing information and buffer period information of the output layer set, efficient management of the subbit stream, including replacing non-scalable nestable timing information to ensure HRD consistency, and managing multi-layer video encoding and decoding through access units and diffusion factors.
It improves the parallel processing capabilities of video encoding and decoding, ensures HRD consistency, supports efficient pruning and sub-bit stream processing of multi-layer video, and optimizes the timing and buffer management of video data streams.
Smart Images

Figure CN114788281B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to video encoding and video decoding, and in particular to a video encoder, to a video decoder, to methods for encoding and decoding, and to a video data stream for implementing advanced video coding concepts. Background Art
[0002] H.265 / HEVC (HEVC = High Efficiency Video Coding) is a video codec that already provides tools for improving or even enabling parallel processing at the encoder and / or decoder. For example, HEVC supports the subdivision of a picture into an array of tiles that are coded independently of each other. Another concept supported by HEVC relates to WPP, according to which CTU rows or CTU lines of a picture can be processed in parallel (e.g. in the form of slices) from left to right, as long as a certain minimum CTU offset (CTU = Coding Tree Unit) is respected in the processing of consecutive CTU rows. However, it would be advantageous to have a video codec with capabilities to support parallel processing of a video encoder and / or video decoder even more efficiently.
[0003] Typically, in video coding, the decoding process of picture samples requires smaller partitions, where samples are divided into rectangular areas for joint processing, such as prediction or transform coding. Therefore, pictures are divided into blocks of a specific size that remains constant during the encoding of a video sequence. In the H.264 / AVC standard, fixed-size blocks of 16x16 samples, so-called macroblocks, are used (AVC = Advanced Video Coding).
[0004] In the state-of-the-art HEVC standard (see [1]), there is a maximum size of Coding Tree Block (CTB) or Coding Tree Unit (CTU) of 64×64 samples. In the further description of HEVC, the more general term CTU is used for this type of block.
[0005] CTUs are processed in raster scan order, starting with the top-left CTU and processing the CTUs in the image row by row, working down to the bottom-right CTU.
[0006] The decoded CTU data is organized into containers called slices. Originally, in previous video coding standards, a slice referred to a segment consisting of one or more consecutive CTUs of a picture. Slices were adopted for segmenting decoded data. From another perspective, a complete picture can also be defined as a large segment, and therefore, historically, the term slice still applies. In addition to the decoded picture samples, a slice also includes additional information related to the decoding process of the slice itself, which is placed in the so-called slice header.
[0007] According to the current state of the art, the VCL (Video Coding Layer) also includes techniques for fragmentation and spatial segmentation. For example, such segmentation can be applied in video coding for various reasons, such as processing load balancing in parallelization, CTU size matching in network transmission, and error mitigation.
[0008] As specified in the video coding standard, the bitstream has information associated with HRD conformance. This conformance consists of a hypothetical reference decoder (HRD) that includes a buffer model that assumes that NAL units enter the coded picture buffer (CPB) before the decoder and are removed from it at specific times to ensure that the CPB size is not exceeded (buffer overrun) or that NAL units arrive no later than they need to be removed (buffer underrun). In addition, the model consists of a decoded picture buffer (DPB) from which decoded pictures are output when they are no longer needed for prediction, and in many implementations, the size of the decoded pictures is also limited. The timing information of the HRD is conveyed in the bitstream via so-called SEI messages, specifically the buffering period (BP) SEI message that defines specific timing information for the buffering period (a number of access units, or AUs), the picture timing (PT) SEI message that conveys timing information for a single associated AU, and the decoding unit information (DUI) SEI message that conveys timing information for an associated subset of AUs (i.e., decoding units, or DUs).
[0009] The bitstream as specified in the video coding standard has information associated with HRD (Hypothetical Reference Decoder) conformance. This conformance consists of an assumed buffer model that assumes that NAL units enter the coded picture buffer (CPB) and are removed from it at specific times to ensure that the CPB size is not exceeded (buffer overrun) or that NAL units arrive no later than they need to be removed (buffer underrun).
[0010] When the bitstream is scalable, pruning can be performed to obtain sub-bitstreams that are also conformant. For example, when there is an output layer set (OLS) (hereinafter referred to as B3) containing three layers with resolution scalability (e.g., a 480p base layer, a 720p first enhancement layer, and a 1080p second enhancement layer), two sub-bitstreams can be obtained: one with two layers (480p and 720p) B2, and the other with one layer (480p) B1. Similarly, the OLS can be used for temporal scalability, where B3, B2, and B1 have the same resolution but different frame rates.
[0011] Obviously, such bitstreams B3, B2 and B1 have different HRD consistency because their required CPB size, code rate and timing information may be different.
[0012] Different CPB sizes and bitrates are indicated in the VPS as features of the defined set of output layers (3 in the described example). The different timing information is provided by so-called nesting SEI messages. The nesting SEI message may contain a nesting buffering period SEI and a picture timing SEI, which are applied to a sub-bitstream obtainable by bitstream pruning (bitstream extraction). When this operation (extraction or pruning) is performed, the buffering period SEI message and the picture timing SEI message of, for example, the input bitstream B3 are then removed from the bitstream and from the NAL units also belonging to the 2nd enhancement layer. Furthermore, the buffering period SEI message and the picture timing SEI message corresponding to the bitstream B2 carried in the nesting SEI message are placed in the bitstream outside the nesting SEI message, thereby replacing the removed bitstream. Summary of the Invention
[0013] It is an object of the present invention to provide improved concepts for video encoding and video decoding.
[0014] The objects of the invention are solved by the subjects of the independent claims.
[0015] Preferred embodiments are provided in the dependent claims.
[0016] According to an embodiment, a video data stream having video encoded therein is provided. The video data stream includes an indication of whether one or more scalable nested supplemental enhancement information messages are present within the video data stream, the one or more scalable nested supplemental enhancement information messages including timing information for each of one or more output layer sets.
[0017] Furthermore, according to an embodiment, a video encoder for encoding a video into a video data stream is provided, such that the video data stream has the video encoded therein. The video encoder is configured to generate the video data stream such that the video data stream includes an indication of whether one or more scalable nested supplemental enhancement information messages are present within the video data stream, the one or more scalable nested supplemental enhancement information messages including timing information for each of one or more output layer sets.
[0018] Furthermore, according to an embodiment, an apparatus is provided for receiving a video data stream as an input bitstream, wherein the video data stream has video encoded therein, processing the input bitstream to obtain sub-bitstreams, and providing an indication of whether one or more scalable nested supplemental enhancement information messages are present within the video data stream, the one or more scalable nested supplemental enhancement information messages including timing information for each of one or more output layer sets.
[0019] Furthermore, according to an embodiment, a method for encoding video into a video data stream is provided, such that the video data stream has the video encoded therein. The method includes generating the video data stream such that the video data stream includes an indication of whether one or more scalable nested supplemental enhancement information messages are present within the video data stream, the one or more scalable nested supplemental enhancement information messages including timing information for each of one or more output layer sets.
[0020] Furthermore, according to an embodiment, a method for receiving a video data stream as an input bitstream is provided, wherein the video data stream has video encoded therein. The method includes processing the input bitstream to obtain a sub-bitstream. Indicating whether one or more scalable nested supplemental enhancement information messages are present within the video data stream, the one or more scalable nested supplemental enhancement information messages including timing information for each of one or more output layer sets.
[0021] Furthermore, a computer program is provided for implementing the method described above when executed on a computer or a signal processor.
[0022] Furthermore, according to an embodiment, a video data stream having video encoded therein is provided. An indication within the video data stream indicates whether timing information for a sub-bitstream is to be obtained from one or more non-scalable nested picture timing supplemental enhancement information messages of the video data stream.
[0023] Furthermore, according to an embodiment, a video data stream having video encoded therein is provided. A first indication within the video data stream indicates whether timing information for a sub-bitstream is to be obtained from one or more non-scalable nested buffering period supplemental enhancement information messages of the video data stream. And / or a second indication within the video data stream indicates whether timing information for the sub-bitstream is to be obtained from one or more non-scalable nested decoding unit supplemental enhancement information messages of the video data stream.
[0024] In addition, according to an embodiment, a video data stream includes one or more non-scalable nested supplementary enhancement information messages including timing information. If the video data stream includes a scalable nested supplementary enhancement information message including timing information, this indicates that, depending on the scalable nested supplementary enhancement information message, all of the one or more non-scalable nested timing information supplementary enhancement information messages are to be replaced with scalable nested supplementary enhancement information messages including timing information (e.g., the timing information is picture timing information, buffering period information, or decoding unit information). Or, a subset of at least one of the one or more non-scalable nested timing information enhancement information messages is to be replaced with a scalable nested supplementary enhancement information message including timing information (e.g., the timing information is picture timing information, buffering period information, or decoding unit information).
[0025] Furthermore, according to an embodiment, a video encoder for encoding video into a video data stream is provided, such that the video data stream has the video encoded therein. The video encoder is configured to generate the video data stream such that an indication within the video data stream indicates whether timing information for a sub-bitstream is to be obtained from one or more non-scalable nested picture timing supplemental enhancement information messages of the video data stream.
[0026] Furthermore, according to an embodiment, a video encoder for encoding video into a video data stream is provided, such that the video data stream has the video encoded therein. A first indication within the video data stream indicates whether timing information for a sub-bitstream is to be obtained from one or more non-scalable nested buffering period supplemental enhancement information messages of the video data stream. And / or a second indication within the video data stream indicates whether timing information for the sub-bitstream is to be obtained from one or more non-scalable nested decoding unit supplemental enhancement information messages of the video data stream.
[0027] In addition, according to an embodiment, a video encoder for encoding a video into a video data stream is provided, so that the video data stream has the video encoded therein. The video encoder is used to generate a video data stream so that the video data stream includes one or more non-scalable nested supplementary enhancement information messages including timing information. If the video data stream includes a scalable nested supplementary enhancement information message including timing information, this indicates that, depending on the scalable nested supplementary enhancement information message, all of the one or more non-scalable nested timing information supplementary enhancement information messages are to be replaced with scalable nested supplementary enhancement information messages including timing information (for example, the timing information is picture timing information or buffering period information or decoding unit information). Or: a subset of at least one of the one or more non-scalable nested timing information supplementary enhancement information messages is to be replaced with a scalable nested supplementary enhancement information message including timing information (for example, the timing information is picture timing information or buffering period information or decoding unit information).
[0028] Furthermore, according to an embodiment, an apparatus is provided for receiving a video data stream as an input bitstream, wherein the video data stream has video encoded therein. The apparatus is configured to process the input bitstream to obtain a sub-bitstream. An indication within the video data stream indicates whether timing information for the sub-bitstream is to be obtained from one or more non-scalable nested picture timing supplemental enhancement information messages of the video data stream.
[0029] Furthermore, according to an embodiment, an apparatus is provided for receiving a video data stream as an input bitstream, wherein the video data stream has video encoded therein. The apparatus is configured to process the input bitstream to obtain a sub-bitstream. A first indication within the video data stream indicates whether timing information for the sub-bitstream is to be obtained from one or more non-scalable nested buffering period supplemental enhancement information messages of the video data stream. And / or a second indication within the video data stream indicates whether timing information for the sub-bitstream is to be obtained from one or more non-scalable nested decoding unit supplemental enhancement information messages of the video data stream.
[0030] In addition, according to an embodiment, an apparatus for receiving a video data stream as an input bitstream is provided, wherein the video data stream has video encoded therein. The apparatus is configured to process the input bitstream to obtain a sub-bitstream. The video data stream includes one or more non-scalable nested supplementary enhancement information messages including timing information. If the video data stream includes a scalable nested supplementary enhancement information message including timing information, the apparatus relies on the scalable nested supplementary enhancement information message to: replace all one or more non-scalable nested timing information supplementary enhancement information messages with scalable nested supplementary enhancement information messages including timing information (e.g., the timing information is picture timing information or buffering period information or decoding unit information). Or, replace a subset of at least one of the one or more non-scalable nested timing information supplementary enhancement information messages with a scalable nested supplementary enhancement information message including timing information (e.g., the timing information is picture timing information or buffering period information or decoding unit information).
[0031] Furthermore, according to an embodiment, a method for encoding a video into a video data stream is provided, such that the video data stream has the video encoded therein. The method includes generating the video data stream such that an indication within the video data stream indicates whether timing information for a sub-bitstream is to be obtained from one or more non-scalable nested picture timing supplemental enhancement information messages of the video data stream.
[0032] Furthermore, according to an embodiment, a method is provided for receiving a video data stream as an input bitstream, wherein the video data stream has video encoded therein. The method includes processing the input bitstream to obtain a sub-bitstream. An indication within the video data stream indicates whether timing information for the sub-bitstream is to be obtained from one or more non-scalable nested picture timing supplemental enhancement information messages of the video data stream.
[0033] Furthermore, according to an embodiment, a method for encoding a video into a video data stream is provided, such that the video data stream has the video encoded therein. A first indication within the video data stream indicates whether timing information for a sub-bitstream is to be obtained from one or more non-scalable nested buffering period supplemental enhancement information messages of the video data stream. And / or a second indication within the video data stream indicates whether the timing information for the sub-bitstream is to be obtained from one or more non-scalable nested decoding unit supplemental enhancement information messages of the video data stream.
[0034] Furthermore, according to an embodiment, a method for receiving a video data stream as an input bitstream is provided, wherein the video data stream has video encoded therein. The method includes processing the input bitstream to obtain a sub-bitstream. A first indication within the video data stream indicates whether timing information for the sub-bitstream is to be obtained from one or more non-scalable nested buffering period supplemental enhancement information messages of the video data stream. And / or a second indication within the video data stream indicates whether the timing information for the sub-bitstream is to be obtained from one or more non-scalable nested decoding unit supplemental enhancement information messages of the video data stream.
[0035] Furthermore, according to an embodiment, a method for encoding a video into a video data stream is provided, such that the video data stream has the video encoded therein. The method includes generating the video data stream such that the video data stream includes one or more non-scalable nested supplemental enhancement information messages including timing information. If the video data stream includes a scalable nested supplemental enhancement information message including the timing information, this indicates that at least one of the one or more non-scalable nested picture timing supplemental enhancement information messages is to be replaced by a scalable nested supplemental enhancement information message including the timing information.
[0036] Furthermore, according to an embodiment, a method for receiving a video data stream as an input bitstream is provided, wherein the video data stream has video encoded therein. The method includes processing the input bitstream to obtain a sub-bitstream. The video data stream includes one or more non-scalable nested supplemental enhancement information messages including timing information. If the video data stream includes a scalable nested supplemental enhancement information message including the timing information, the method includes replacing at least one of the one or more non-scalable nested picture timing supplemental enhancement information messages with a scalable nested supplemental enhancement information message including the timing information.
[0037] Furthermore, a computer program is provided for implementing as described above when executed on a computer or a signal processor.
[0038] Furthermore, according to an embodiment, a video data stream having video encoded therein is provided. The video data stream includes a plurality of access units. For each access unit in the plurality of access units, if the access unit includes two or more scalable nested supplementary enhancement information messages, and the two or more scalable nested supplementary enhancement information messages include buffering period information and / or picture timing information for an output layer set, then the buffering period information and / or picture timing information for the output layer set in all two or more scalable nested supplementary enhancement information messages in the access unit are equal.
[0039] Furthermore, according to an embodiment, a video data stream having video encoded therein is provided. The video data stream includes a plurality of access units. For each access unit in the plurality of access units, if the access unit includes two or more scalable nested supplementary enhancement information messages, and the two or more scalable nested supplementary enhancement information messages include buffering period information and / or picture timing information for an output layer set, the video data stream includes an indication indicating whether the buffering period information and / or picture timing information for the output layer set in all two or more scalable nested supplementary enhancement information messages for the access unit are equal.
[0040] In addition, according to an embodiment, a video data stream having a video encoded therein is provided. The video data stream includes a plurality of access units. For each access unit of the plurality of access units, if the access unit includes three or more scalable nested supplementary enhancement information messages, two or more scalable nested supplementary enhancement information messages include buffering period information and / or picture timing information for an output layer set, then the buffering period information and / or picture timing information for the output layer set appears only in one scalable nested supplementary enhancement information message among the three or more scalable nested supplementary enhancement information messages, and in another scalable nested supplementary enhancement information message among the three or more scalable nested supplementary enhancement information messages that immediately follows the one scalable nested supplementary enhancement information message among the three or more scalable nested supplementary enhancement information messages.
[0041] In addition, according to an embodiment, a video data stream having video encoded therein is provided. The video data stream includes multiple access units. For each access unit in the multiple access units, if the access unit includes three or more scalable nested supplementary enhancement information messages, two or more scalable nested supplementary enhancement information messages include buffering period information and / or picture timing information for an output layer set, then the video data stream includes an indication indicating whether the buffering period information and / or picture timing information for the output layer set appears only in one of the three or more scalable nested supplementary enhancement information messages, and in another scalable nested supplementary enhancement information message of the three or more scalable nested supplementary enhancement information messages that immediately follows the one of the three or more scalable nested supplementary enhancement information messages.
[0042] Furthermore, according to an embodiment, a video encoder for encoding a video into a video data stream is provided, such that the video data stream has the video encoded therein. The video encoder is configured to generate the video data stream such that the video data stream includes a plurality of access units. For each access unit of the plurality of access units, if the access unit includes two or more scalable nested supplementary enhancement information messages, and the two or more scalable nested supplementary enhancement information messages include buffering period information and / or picture timing information for an output layer set, the video encoder is configured to generate the video data stream such that the video data stream includes an indication of whether the buffering period information and / or picture timing information for the output layer set in all two or more scalable nested supplementary enhancement information messages of the access unit are equal.
[0043] In addition, according to an embodiment, a video encoder for encoding a video into a video data stream is provided, so that the video data stream has the video encoded therein. The video encoder is used to generate the video data stream so that the video data stream includes multiple access units. For each access unit of the multiple access units, if the access unit includes three or more scalable nested supplementary enhancement information messages, two or more scalable nested supplementary enhancement information messages include buffering period information and / or picture timing information for an output layer set, then the video encoder is used to generate the video data stream so that the buffering period information and / or picture timing information for the output layer set appears only in one scalable nested supplementary enhancement information message among the three or more scalable nested supplementary enhancement information messages, and in another scalable nested supplementary enhancement information message among the three or more scalable nested supplementary enhancement information messages that immediately follows the one scalable nested supplementary enhancement information message among the three or more scalable nested supplementary enhancement information messages.
[0044] In addition, according to an embodiment, a video encoder for encoding a video into a video data stream is provided, so that the video data stream has the video encoded therein. The video encoder is used to generate the video data stream so that the video data stream includes a plurality of access units. For each access unit of the plurality of access units, if the access unit includes three or more scalable nested supplementary enhancement information messages, two or more scalable nested supplementary enhancement information messages include buffering period information and / or picture timing information for an output layer set, then the video encoder is used to generate the video data stream so that the video data stream includes an indication indicating whether the buffering period information and / or picture timing information for the output layer set appears only in one scalable nested supplementary enhancement information message among the three or more scalable nested supplementary enhancement information messages and in another scalable nested supplementary enhancement information message among the three or more scalable nested supplementary enhancement information messages that immediately follows the one scalable nested supplementary enhancement information message among the three or more scalable nested supplementary enhancement information messages.
[0045] Furthermore, according to an embodiment, an apparatus is provided for receiving a video data stream as an input bitstream, wherein the video data stream has video encoded therein. The input bitstream includes a plurality of access units. The apparatus is configured to process the access units of the input bitstream to obtain sub-bitstreams. For each access unit of the plurality of access units, if the access unit includes two or more scalable nested supplementary enhancement information messages, the two or more scalable nested supplementary enhancement information messages including buffering period information and / or picture timing information for an output layer set, then the buffering period information and / or picture timing information for the output layer set in all two or more scalable nested supplementary enhancement information messages of the access unit are equal.
[0046] Furthermore, according to an embodiment, an apparatus for receiving a video data stream as an input bitstream is provided, wherein the video data stream has video encoded therein. The input bitstream includes a plurality of access units. The apparatus is configured to process an access unit of the input bitstream to obtain a sub-bitstream. For each access unit of the plurality of access units, if the access unit includes two or more scalable nested supplementary enhancement information messages, the two or more scalable nested supplementary enhancement information messages including buffering period information and / or picture timing information for an output layer set, the video data stream includes an indication of whether the buffering period information and / or picture timing information for the output layer set in all two or more scalable nested supplementary enhancement information messages of the access unit are equal.
[0047] In addition, according to an embodiment, an apparatus for receiving a video data stream as an input bitstream is provided, wherein the video data stream has video encoded therein. The input bitstream includes a plurality of access units, and the apparatus is configured to process the access units of the input bitstream to obtain sub-bitstreams. For each access unit of the plurality of access units, if the access unit includes three or more scalable nested supplementary enhancement information messages, two or more scalable nested supplementary enhancement information messages include buffering period information and / or picture timing information for an output layer set, then the buffering period information and / or picture timing information for the output layer set appears only in one scalable nested supplementary enhancement information message among the three or more scalable nested supplementary enhancement information messages, and in another scalable nested supplementary enhancement information message among the three or more scalable nested supplementary enhancement information messages that immediately follows the one scalable nested supplementary enhancement information message among the three or more scalable nested supplementary enhancement information messages.
[0048] In addition, according to an embodiment, an apparatus for receiving a video data stream as an input bitstream is provided, wherein the video data stream has video encoded therein. The input bitstream includes a plurality of access units. The apparatus is configured to process the access units of the input bitstream to obtain a sub-bitstream. For each access unit of the plurality of access units, if the access unit includes three or more scalable nested supplementary enhancement information messages, two or more scalable nested supplementary enhancement information messages include buffering period information and / or picture timing information for an output layer set, then the video data stream includes an indication of whether the buffering period information and / or picture timing information for the output layer set appears only in one of the three or more scalable nested supplementary enhancement information messages, and in another scalable nested supplementary enhancement information message of the three or more scalable nested supplementary enhancement information messages that immediately follows the one of the three or more scalable nested supplementary enhancement information messages.
[0049] Furthermore, according to an embodiment, a method for encoding a video into a video data stream is provided, such that the video data stream has the video encoded therein. The method includes generating the video data stream such that the video data stream includes a plurality of access units. For each access unit in the plurality of access units,
[0050] If an access unit includes two or more scalable nested supplementary enhancement information messages, and the two or more scalable nested supplementary enhancement information messages include buffering period information and / or picture timing information for an output layer set, the method includes generating a video data stream so that the buffering period information and / or picture timing information for the output layer set are equal in all two or more scalable nested supplementary enhancement information messages of the access unit.
[0051] Furthermore, according to an embodiment, a method for encoding a video into a video data stream is provided, such that the video data stream has the video encoded therein. The method includes generating the video data stream such that the video data stream includes a plurality of access units. For each access unit in the plurality of access units,
[0052] If an access unit includes two or more scalable nested supplementary enhancement information messages, and the two or more scalable nested supplementary enhancement information messages include buffering period information and / or picture timing information for an output layer set, the method includes generating a video data stream so that the video data stream includes an indication of whether the buffering period information and / or picture timing information for the output layer set in all two or more scalable nested supplementary enhancement information messages of the access unit are equal.
[0053] Furthermore, according to an embodiment, a method for encoding a video into a video data stream is provided, such that the video data stream has the video encoded therein. The method includes generating the video data stream such that the video data stream includes a plurality of access units. For each access unit in the plurality of access units,
[0054] If the access unit includes three or more scalable nested supplementary enhancement information messages, and two or more scalable nested supplementary enhancement information messages include buffering period information and / or picture timing information for the output layer set, the method includes generating a video data stream so that the buffering period information and / or picture timing information for the output layer set appears only in one scalable nested supplementary enhancement information message among the three or more scalable nested supplementary enhancement information messages, and in another scalable nested supplementary enhancement information message among the three or more scalable nested supplementary enhancement information messages that immediately follows the one scalable nested supplementary enhancement information message among the three or more scalable nested supplementary enhancement information messages.
[0055] Furthermore, according to an embodiment, a method for encoding a video into a video data stream is provided, such that the video data stream has the video encoded therein. The method includes generating the video data stream such that the video data stream includes a plurality of access units. For each access unit in the plurality of access units,
[0056] If an access unit includes three or more scalable nested supplementary enhancement information messages, and two or more scalable nested supplementary enhancement information messages include buffering period information and / or picture timing information for an output layer set, the method includes generating a video data stream so that the video data stream includes an indication of whether the buffering period information and / or picture timing information for the output layer set appears only in one scalable nested supplementary enhancement information message among the three or more scalable nested supplementary enhancement information messages and in another scalable nested supplementary enhancement information message among the three or more scalable nested supplementary enhancement information messages that immediately follows the one scalable nested supplementary enhancement information message among the three or more scalable nested supplementary enhancement information messages.
[0057] Furthermore, according to an embodiment, a method for receiving a video data stream as an input bitstream is provided, wherein the video data stream has video encoded therein. The input bitstream includes a plurality of access units. The method includes processing the access units of the input bitstream to obtain sub-bitstreams. For each access unit of the plurality of access units, if the access unit includes two or more scalable nested supplementary enhancement information messages, the two or more scalable nested supplementary enhancement information messages including buffering period information and / or picture timing information for an output layer set, then the buffering period information and / or picture timing information for the output layer set in all two or more scalable nested supplementary enhancement information messages of the access unit are equal.
[0058] Furthermore, according to an embodiment, a method for receiving a video data stream as an input bitstream is provided, wherein the video data stream has video encoded therein. The input bitstream includes a plurality of access units. The method includes processing an access unit of the input bitstream to obtain a sub-bitstream. For each access unit of the plurality of access units, if the access unit includes two or more scalable nested supplementary enhancement information messages, the two or more scalable nested supplementary enhancement information messages including buffering period information and / or picture timing information for an output layer set, the video data stream includes an indication indicating whether the buffering period information and / or picture timing information for the output layer set in all two or more scalable nested supplementary enhancement information messages of the access unit are equal.
[0059] In addition, according to an embodiment, a method for receiving a video data stream as an input bitstream is provided, wherein the video data stream has video encoded therein. The input bitstream includes a plurality of access units. The method includes processing the access units of the input bitstream to obtain a sub-bitstream. For each access unit of the plurality of access units, if the access unit includes three or more scalable nested supplementary enhancement information messages, two or more scalable nested supplementary enhancement information messages include buffering period information and / or picture timing information for an output layer set, then the buffering period information and / or picture timing information for the output layer set appears only in one scalable nested supplementary enhancement information message of the three or more scalable nested supplementary enhancement information messages, and in another scalable nested supplementary enhancement information message of the three or more scalable nested supplementary enhancement information messages that immediately follows the one scalable nested supplementary enhancement information message of the three or more scalable nested supplementary enhancement information messages.
[0060] In addition, according to an embodiment, a method for receiving a video data stream as an input bitstream is provided, wherein the video data stream has video encoded therein. The input bitstream includes a plurality of access units. The method includes processing the access units of the input bitstream to obtain a sub-bitstream. For each access unit of the plurality of access units, if the access unit includes three or more scalable nested supplementary enhancement information messages, two or more scalable nested supplementary enhancement information messages include buffering period information and / or picture timing information for an output layer set, then the video data stream includes an indication indicating whether the buffering period information and / or picture timing information for the output layer set appears only in one of the three or more scalable nested supplementary enhancement information messages, and in another scalable nested supplementary enhancement information message of the three or more scalable nested supplementary enhancement information messages that immediately follows the one of the three or more scalable nested supplementary enhancement information messages.
[0061] Furthermore, a computer program is provided for implementing the method as described above when executed on a computer or a signal processor.
[0062] Furthermore, according to an embodiment, a video data stream having video encoded therein is provided. The video data stream includes a plurality of access units. The video data stream includes a spreading factor that depends on the number of sub-bitstreams in the video data stream, or the video data stream includes a clock sub-tick value that depends on the highest sub-bitstream in the sub-bitstreams of the video data stream.
[0063] Furthermore, according to an embodiment, a video encoder is provided for encoding video into a video data stream, such that the video data stream has the video encoded therein. The video encoder is configured to generate the video data stream such that the video data stream includes a plurality of access units. Furthermore, the video encoder is configured to generate the video data stream such that the video data stream includes a spreading factor that depends on the number of sub-bitstreams in the video data stream, or such that the video data stream includes a clock sub-tick value that depends on the highest sub-bitstream in the sub-bitstreams of the video data stream.
[0064] Furthermore, according to an embodiment, a video encoder for encoding a video into a video data stream is provided, such that the video data stream has the video encoded therein. The video encoder is configured to generate the video data stream such that the video data stream includes sub-layer-specific frame rate information for the sub-layer; and / or the video encoder is configured to generate the video data stream such that the video data stream includes sub-layer-specific frame display duration information for the sub-layer.
[0065] Furthermore, according to an embodiment, a video decoder is provided for receiving a video data stream as an input bitstream, wherein the video data stream has video encoded therein. The video data stream includes a plurality of access units. The video decoder is configured to decode the video data stream to decode the video. The video data stream includes a spreading factor, the spreading factor being dependent on the number of sub-bitstreams in the video data stream, wherein the video decoder is configured to use the spreading factor to decode the video; or the video data stream includes a clock sub-tick value, the clock sub-tick value being dependent on the highest sub-bitstream in the sub-bitstreams of the video data stream, wherein the video decoder is configured to use the clock sub-tick value to decode the video.
[0066] Furthermore, according to an embodiment, a video decoder is provided for receiving a video data stream as an input bitstream, wherein the video data stream has video encoded therein. The video decoder is configured to decode the video data stream to decode the video. The video data stream includes sub-layer-specific frame rate information for a sub-layer, and / or wherein the video data stream includes sub-layer-specific frame display duration information for the sub-layer. The decoder is configured to determine a spreading factor using the sub-layer-specific frame rate information for the sub-layer and / or using the sub-layer-specific frame display duration information.
[0067] Furthermore, according to an embodiment, a method for encoding video into a video data stream is provided, wherein the video data stream has the video encoded therein. The method includes generating the video data stream such that the video data stream includes a plurality of access units. The method includes generating the video data stream such that the video data stream includes a spreading factor that depends on the number of sub-bitstreams in the video data stream; or the method includes generating the video data stream such that the video data stream includes a clock sub-tick value that depends on the highest sub-bitstream in the sub-bitstreams of the video data stream.
[0068] Furthermore, according to an embodiment, a method is provided for receiving a video data stream as an input bitstream, wherein the video data stream has video encoded therein. The video data stream includes a plurality of access units. The method includes decoding the video data stream to decode the video. The video data stream includes a spreading factor, the spreading factor being dependent on the number of sub-bitstreams in the video data stream, wherein the method includes using the spreading factor to decode the video; or the video data stream includes a clock sub-tick value, the clock sub-tick value being dependent on the highest sub-bitstream in the sub-bitstreams of the video data stream, wherein the method includes using the clock sub-tick value to decode the video.
[0069] Furthermore, according to an embodiment, a method for receiving a video data stream as an input bitstream is provided, wherein the video data stream has video encoded therein. The method includes decoding the video data stream to decode the video. The video data stream includes sub-layer-specific frame rate information for a sub-layer, and / or wherein the video data stream includes sub-layer-specific frame display duration information for the sub-layer. The method includes determining a spreading factor using the sub-layer-specific frame rate information for the sub-layer and / or using the sub-layer-specific frame display duration information.
[0070] Furthermore, a computer program is provided for implementing the method as described above when executed on a computer or a signal processor.
[0071] According to an embodiment, a video decoder is provided for receiving a video data stream as an input bitstream, wherein the video data stream has video encoded therein. The video decoder is configured to decode the video data stream to decode the video. To decode the video, the video decoder estimates a decoded picture buffer size for a sub-picture based on information within the video data stream indicating a picture buffer size for a current decoding region.
[0072] Furthermore, according to an embodiment, a video decoder is provided for receiving a video data stream as an input bitstream, wherein the video data stream has video encoded therein. The video decoder is configured to decode the video data stream to decode the video. To decode the video, the video decoder is configured to estimate a bitrate for a sub-picture based on information within the video data stream indicating bitrate information of a currently decoded video sequence.
[0073] Furthermore, according to an embodiment, a video decoder is provided for receiving a video data stream as an input bitstream, wherein the video data stream has video encoded therein. The video decoder is configured to decode the video data stream to decode the video. To decode the video, the video decoder is configured to receive a decoded picture buffer size for a sub-picture encoded within the video data stream and to decode the video using the decoded picture buffer size for the sub-picture; and / or to decode the video, the video decoder is configured to receive a bitrate for a sub-picture encoded within the video data stream and to decode the video using the bitrate for the sub-picture.
[0074] Furthermore, according to an embodiment, a video data stream having video encoded therein is provided, wherein the video data stream includes the syntax element cpb_size_value_minus1[i][j] and the syntax element cpb_size_scale, or the video data stream includes the syntax element bit_rate_value_minus1[i][j] and the syntax element bit_rate_scale.
[0075] Furthermore, according to an embodiment, a video data stream having video encoded therein is provided, wherein the video data stream includes an indication indicating whether currently decoded picture buffer size information should be used to estimate a decoded picture buffer size of a sub-picture, and / or the video data stream includes an indication indicating whether currently decoded video sequence bit rate information should be used to estimate a bit rate for the sub-picture.
[0076] Furthermore, according to an embodiment, a video data stream having a video encoded therein is provided, wherein the video data stream includes an indication indicating whether a decoded picture buffer size of a sub-picture is encoded in the video data stream or whether the decoded picture buffer size of the sub-picture should be estimated, and / or the video data stream includes an indication indicating whether a bit rate of the sub-picture is encoded within the video data stream or whether the bit rate of the sub-picture should be estimated.
[0077] Furthermore, according to an embodiment, a video encoder for encoding a video into a video data stream is provided, such that the video data stream has the video encoded therein. The video encoder is configured to generate the video data stream such that the video data stream includes the syntax element cpb_size_value_minus1[i][j] and the syntax element cpb_size_scale. Alternatively, the video encoder generates the video data stream such that the video data stream includes the syntax element bit_rate_value_minus1[i][j] and the syntax element bit_rate_scale.
[0078] Furthermore, according to an embodiment, a video encoder for encoding a video into a video data stream is provided, such that the video data stream has the video encoded therein. The video encoder generates the video data stream such that the video data stream includes an indication indicating whether the decoded picture buffer size information of the sub-picture should be used to estimate the decoded picture buffer size of the sub-picture, and / or the video encoder generates the video data stream such that the video data stream includes an indication indicating whether the bit rate information of the currently decoded video sequence should be used to estimate the bit rate of the sub-picture.
[0079] Furthermore, according to an embodiment, a video encoder for encoding a video into a video data stream is provided, such that the video data stream has the video encoded therein. The video encoder is configured to generate the video data stream such that the video data stream includes an indication indicating whether a decoded picture buffer size for a sub-picture is encoded within the video data stream or whether the decoded picture buffer size for the sub-picture should be estimated, and / or the video encoder is configured to generate the video data stream such that the video data stream includes an indication indicating whether a bit rate for the sub-picture is encoded within the video data stream or whether the bit rate for the sub-picture should be estimated.
[0080] Furthermore, according to an embodiment, a video encoder for encoding video into a video data stream is provided, such that the video data stream has the video encoded therein. The video encoder is configured to generate the video data stream, wherein, if the video data stream includes common decoding unit removal timing information and a plurality of extractable sub-bitstreams, each of the plurality of sub-bitstreams being specific to an output layer set, a hypothetical reference decoder parameter syntax structure specific to each output layer set in a video parameter set or a sequence parameter set or a supplemental enhancement information message of the video data stream includes an absolute value of a beat divisor or a spreading factor for extending the common decoding unit removal timing.
[0081] Furthermore, according to an embodiment, a method is provided for receiving a video data stream as an input bitstream, wherein the video data stream has video encoded therein. The method includes decoding the video data stream to decode the video. To decode the video, the method includes estimating a decoded picture buffer size for a sub-picture based on information within the video data stream indicating picture buffer size information for a current decoding region.
[0082] Furthermore, according to an embodiment, a method is provided for receiving a video data stream as an input bitstream, wherein the video data stream has video encoded therein. The method includes decoding the video data stream to decode the video. To decode the video, the method includes estimating a bitrate for a sub-picture based on information within the video data stream indicating bitrate information for a currently decoded video sequence.
[0083] Furthermore, according to an embodiment, a method is provided for receiving a video data stream as an input bitstream, wherein the video data stream has video encoded therein. The method includes decoding the video data stream to decode the video. To decode the video, the method includes receiving a decoded picture buffer size for a sub-picture encoded within the video data stream, and decoding the video using the decoded picture buffer size for the sub-picture; and / or to decode the video, the method includes receiving a bitrate for a sub-picture encoded within the video data stream, and decoding the video using the bitrate for the sub-picture.
[0084] Furthermore, according to an embodiment, a method for encoding a video into a video data stream is provided, wherein the video data stream has the video encoded therein. The method includes generating the video data stream such that the video data stream includes a syntax element cpb_size_value_minus1[i][j] and a syntax element cpb_size_scale; or the method includes generating the video data stream such that the video data stream includes a syntax element bit_rate_value_minus1[i][j] and a syntax element bit_rate_scale.
[0085] Furthermore, according to an embodiment, a method for encoding a video into a video data stream is provided, such that the video data stream has the video encoded therein. The method includes generating the video data stream such that the video data stream includes an indication indicating whether current decoded picture buffer size information should be used to estimate a decoded picture buffer size for a sub-picture, and / or the method includes generating the video data stream such that the video data stream includes an indication indicating whether current decoded video sequence bit rate information should be used to estimate a bit rate for the sub-picture.
[0086] Furthermore, according to an embodiment, a method for encoding a video into a video data stream is provided, such that the video data stream has the video encoded therein. The method includes generating the video data stream such that the video data stream includes an indication indicating whether a decoded picture buffer size for a sub-picture is encoded within the video data stream or whether the decoded picture buffer size for the sub-picture should be estimated, and / or the method includes generating the video data stream such that the video data stream includes an indication indicating whether a bit rate for the sub-picture is encoded within the video data stream or whether the bit rate for the sub-picture should be estimated.
[0087] Furthermore, a computer program is provided for implementing the method as described above when executed on a computer or a signal processor.
[0088] According to an embodiment, a video data stream having video encoded therein is provided. The video data stream includes a plurality of access units. The video data stream further includes delta time information for each of two or more decoding units of an access unit in the plurality of access units, wherein a decoding unit removal time for each of the two or more decoding units of the access unit is dependent on the access unit removal time for the access unit and on the delta time information for the decoding unit.
[0089] Furthermore, according to an embodiment, a video encoder for encoding a video into a video data stream is provided, such that the video data stream has the video encoded therein. The video encoder is configured to generate the video data stream such that the video data stream includes a plurality of access units. Furthermore, the video encoder is configured to generate the video data stream such that the video data stream includes delta time information for each of two or more decoding units of an access unit in the plurality of access units, wherein a decoding unit removal time for each of the two or more decoding units of the access unit is dependent on the access unit removal time for the access unit and on the delta time information for the decoding unit.
[0090] Furthermore, according to an embodiment, a video decoder for receiving a video data stream as an input bitstream is provided, wherein the video data stream has video encoded therein. The video data stream includes a plurality of access units. The video decoder is configured to decode the video data stream to decode the video. Furthermore, the video data stream includes delta time information for each of two or more decoding units of an access unit in the plurality of access units, wherein a decoding unit removal time for each of the two or more decoding units of the access unit depends on the access unit removal time for the access unit and on the delta time information for the decoding unit, wherein the video decoder is configured to decode the video data stream using the delta time information for each of the two or more decoding units of the access unit.
[0091] Furthermore, according to an embodiment, a method for encoding video into a video data stream is provided, such that the video data stream has the video encoded therein. The method includes generating the video data stream such that the video data stream includes a plurality of access units. The method includes generating the video data stream such that the video data stream includes delta time information for each of two or more decoding units of an access unit in the plurality of access units, wherein a decoding unit removal time for each of the two or more decoding units of the access unit is dependent on the access unit removal time for the access unit and on the delta time information for the decoding unit.
[0092] Furthermore, according to an embodiment, a method for receiving a video data stream as an input bitstream is provided, wherein the video data stream has video encoded therein. The video data stream comprises a plurality of access units. The method comprises decoding the video data stream to decode the video. The video data stream comprises delta time information for each of two or more decoding units of an access unit of the plurality of access units, wherein a decoding unit removal time for each of the two or more decoding units of the access unit depends on the access unit removal time for the access unit and on the delta time information for the decoding unit, wherein the method comprises decoding the video data stream using the delta time information for each of the two or more decoding units of the access unit.
[0093] Furthermore, a computer program is provided for implementing the method as described above when executed on a computer or a signal processor. BRIEF DESCRIPTION OF THE DRAWINGS
[0094] Figure 1 A video encoder for encoding video into a video data stream according to an embodiment is described.
[0095] Figure 2 An apparatus for receiving an input video data stream according to an embodiment is described.
[0096] Figure 3 A video decoder for receiving a video data stream having video stored therein according to an embodiment is described.
[0097] Figure 4 The variation of the removal time is illustrated for the case where there are three decoding units per access unit.
[0098] Figure 5 Two layers and the removal times of access units and decoding units of the two layers are described.
[0099] Figure 6 A video encoder is described.
[0100] Figure 7 A video decoder is described.
[0101] Figure 8 The relationship between the reconstruction signal (eg the reconstructed picture) on the one hand and the combination of the prediction residual signal and the prediction signal as signaled in the data stream on the other hand is illustrated. DETAILED DESCRIPTION
[0102] The following description of the drawings begins with the presentation of a description of an encoder and decoder of a block-based prediction codec for decoding pictures of a video, in order to form an example of a decoding framework into which embodiments of the present invention may be built. Figures 6 to 8 The corresponding encoders and decoders are described. Thereafter, a description of embodiments of the concepts of the invention is given together with information on how these concepts can be implemented into Figure 6 and Figure 7 The descriptions in the encoder and decoder are presented together, although with Figures 1 to 3 The embodiments described below can also be used to form Figure 6 and Figure 7 The encoder and decoder are based on the decoding framework that operates the encoder and decoder.
[0103] Figure 6 A video encoder is shown, an apparatus for predictively decoding a picture 12 into a data stream 14, exemplarily using transform-based residual coding. The apparatus or encoder is denoted by reference numeral 10. Figure 7 A corresponding video decoder 20 is shown, for example, an apparatus 20 configured to predictively decode a picture 12′ from a data stream 14 also using transform-based residual decoding, wherein a prime has been used to indicate that the picture 12′ reconstructed by the decoder 20 deviates from the picture 12 originally encoded by the apparatus 10 with respect to decoding losses introduced by quantization of the prediction residual signal. Figure 6 and Figure 7 Transform-based prediction residual coding is used illustratively, although embodiments of the present application are not limited to this type of prediction residual coding. Figure 6 and Figure 7 The same is true for other details of the description, as will be outlined below.
[0104] The encoder 10 is configured to perform a spatial-to-spectral transform on the prediction residual signal and encode the prediction residual signal obtained thereby into the data stream 14. Likewise, the decoder 20 is configured to decode the prediction residual signal from the data stream 14 and perform a spectral-to-spatial transform on the prediction residual signal obtained thereby.
[0105] Internally, the encoder 10 may include a prediction residual signal former 22 that generates a prediction residual 24 in order to measure the deviation of the prediction signal 26 from the original signal, for example, the deviation from the picture 12. The prediction residual signal former 22 may be, for example, a subtractor that subtracts the prediction signal from the original signal, for example, from the picture 12. The encoder 10 then further includes a transformer 28 that performs a spatial-to-spectral transformation on the prediction residual signal 24 to obtain a spectral-domain prediction residual signal 24′, which is then quantized by a quantizer 32 also included by the encoder 10. The quantized prediction residual signal 24" is thus decoded into the bit stream 14. To this end, the encoder 10 may optionally include an entropy decoder 34 which entropy decodes the prediction residual signal transformed and quantized into the data stream 14. The prediction signal 26 is generated by a prediction stage 36 of the encoder 10 based on the prediction residual signal 24 encoded into the data stream 14 and decodable from the data stream 14. To this end, as Figure 6 As shown in , the prediction stage 36 may internally include a dequantizer 38 that dequantizes the prediction residual signal 24" to obtain a spectral domain prediction residual signal 24'" corresponding to the signal 24' excluding quantization losses, followed by an inverse transformer 40 that inversely transforms (e.g., spectral to spatial transform) the latter prediction residual signal 24'" to obtain a prediction residual signal 24"" corresponding to the original prediction residual signal 24 excluding quantization losses. Then, a combiner 42 of the prediction stage 36 recombines the prediction signal 26 and the prediction residual signal 24"", such as by addition, to obtain a reconstructed signal 46, e.g., a reconstruction of the original signal 12. The reconstructed signal 46 may correspond to the signal 12'. The prediction module 44 of the prediction stage 36 then generates the prediction signal 26 based on the signal 46 by using, for example, spatial prediction (e.g., intra-picture prediction) and / or temporal prediction (e.g., inter-picture prediction).
[0106] Likewise, Figure 7 As shown in FIG, the decoder 20 may be internally composed of components corresponding to the prediction stage 36 and interconnected in a manner corresponding to the prediction stage 36. In particular, the entropy decoder 50 of the decoder 20 may entropy decode the quantized spectral domain prediction residual signal 24" from the data stream, whereupon the dequantizer 52, inverse transformer 54, combiner 56, and prediction module 58, which are interconnected and cooperate in the manner described above with respect to the modules of the prediction stage 36, recover the reconstructed signal based on the prediction residual signal 24", such that Figure 7 As shown in , the output of the combiner 56 results in a reconstructed signal, picture 12 ′.
[0107] Although not specifically described above, it is readily apparent that the encoder 10 can set some decoding parameters, including, for example, prediction modes, motion parameters, etc., according to some optimization scheme, such as, for example, in a manner that optimizes certain rate- and distortion-related criteria (e.g., decoding cost). For example, the encoder 10 and decoder 20, and corresponding modules 44 and 58, can each support different prediction modes, such as intra-frame decoding mode and inter-frame decoding mode. The granularity at which the encoder and decoder switch between these prediction mode types can correspond to subdividing the pictures 12 and 12', respectively, into decoding segments or decoding blocks. For example, in units of these decoding segments, the pictures can be subdivided into blocks that are decoded intra-frame and blocks that are decoded inter-frame. As outlined in more detail below, intra-frame decoded blocks are predicted based on the spatial, already decoded / decoded neighborhood of the corresponding block. There may be several intra-coding modes, and these may be selected for respective intra-coding segments including directional or angular intra-coding modes, according to which the respective segments are filled by extrapolating sample values of a neighborhood into the respective intra-coding segment along a specific direction specific to the respective directional intra-coding mode. The intra-coding modes may, for example, also include one or more other modes, such as a DC coding mode (according to which the prediction of the respective intra-coding block assigns a DC value to all samples within the respective intra-coding segment), and / or a planar intra-coding mode (according to which the prediction of the respective block is approximated or determined as a spatial distribution of sample values described by a two-dimensional linear function at the sample positions of the respective intra-coding block, with a driving tilt and offset of the plane defined by the two-dimensional linear function based on neighboring samples). In contrast, inter-coding blocks may, for example, be predicted temporally. For inter-coded blocks, a motion vector may be signaled within the data stream, the motion vector indicating the spatial displacement of a portion of a previously coded picture of the video to which the picture 12 belongs, at which portion the previously coded / decoded picture was sampled in order to obtain a prediction signal for the corresponding inter-coded block. This means that, in addition to the residual signal coding comprised by the data stream 14 (such as the entropy coded transform coefficient levels representing the quantized spectral domain prediction residual signal 24″), the data stream 14 may also have encoded therein: coding mode parameters for assigning a coding mode to the individual blocks; prediction parameters for some of the blocks (such as motion parameters for inter-coded segments) and optionally other parameters (such as parameters for controlling and signaling the subdivision of the pictures 12 and 12 ′, respectively, into segments). The decoder 20 uses these parameters to subdivide the pictures in the same way as the encoder did, to assign the same prediction modes to the segments and to perform the same prediction resulting in the same prediction signal.
[0108] Figure 8The relationship between the reconstruction signal (e.g. the reconstructed picture 12') on the one hand and the combination of the prediction residual signal 24"" and the prediction signal 26 as signaled in the data stream 14 on the other hand is illustrated. As already indicated above, the combination may be an addition. The prediction signal 26 is Figure 8 1 is illustrated as subdividing the picture region into intra-coded blocks illustratively indicated using shading and inter-coded blocks illustratively indicated using non-shading. The subdivision may be any subdivision, such as a regular subdivision of the picture region into rows and columns of square or non-square blocks, or a subdivision of the picture 12 from a multi-tree root block into a plurality of leaf blocks of different sizes, such as a quadtree subdivision, wherein in Figure 8 Their hybrid is illustrated in
[15] , where the picture region is first subdivided into rows and columns of root blocks and then further subdivided into one or more leaf blocks according to a recursive multitree subdivision.
[0109] Similarly, for intra-coded blocks 80, the data stream 14 may have an intra-coding mode encoded therein that assigns one of several supported intra-coding modes to the corresponding intra-coded block 80. For inter-coded blocks 82, the data stream 14 may have one or more motion parameters encoded therein. In general, inter-coded blocks 82 are not limited to being temporally coded. Alternatively, inter-coded blocks 82 may be any block predicted from a previously coded portion other than the current picture 12 itself, such as a previously coded picture of the video to which picture 12 belongs, or another view or hierarchically lower-level picture if the encoder and decoder are scalable encoders and decoders, respectively.
[0110] Figure 8 The prediction residual signal 24″″ in is also illustrated as subdividing the picture area into blocks 84. These blocks may be referred to as transform blocks to distinguish them from the decoded blocks 80 and 82. In practice, Figure 8 It is illustrated that the encoder 10 and the decoder 20 may use two different subdivisions of the pictures 12 and 12', respectively, into blocks, namely one into coded blocks 80 and 82, respectively, and another into transform blocks 84, respectively. The two subdivisions may be identical, for example, each coded block 80 and 82 may simultaneously form a transform block 84, but Figure 8A case is described in which, for example, the subdivision into transform blocks 84 forms an extension of the subdivision into coded blocks 80 and 82, such that any block boundary between blocks 80 and 82 overlaps a boundary between two blocks 84, or alternatively, each block 80 and 82 coincides with either one of transform blocks 84 or a cluster of transform blocks 84. However, the subdivision can also be determined or selected independently of one another, such that transform block 84 can alternatively straddle a block boundary between blocks 80 and 82. With respect to the subdivision into transform blocks 84, similar statements hold true as those made regarding the subdivision into blocks 80 and 82. For example, block 84 can be the result of a regular subdivision of the picture region into blocks (with or without arrangement into rows and columns), the result of a recursive multitree subdivision of the picture region, or a combination thereof or any other type of blockification. Note, in passing, that blocks 80, 82, and 84 are not limited to quadratic, rectangular, or any other shape.
[0111] Figure 8 It is further illustrated that the combination of the prediction signal 26 and the prediction residual signal 24"" directly results in the reconstructed signal 12'. However, it should be noted that according to alternative embodiments more than one prediction signal 26 may be combined with the prediction residual signal 24"" to result in the picture 12'.
[0112] exist Figure 8 In the embodiment described below, the transform blocks 84 shall have the following meanings. The transformer 28 and the inverse transformer 54 perform their transforms in units of these transform blocks 84. For example, many codecs use some kind of DST or DCT for all transform blocks 84. Some codecs allow skipping of transforms so that the prediction residual signal is decoded directly in the spatial domain for some of the transform blocks 84. However, according to the embodiments described below, the encoder 10 and the decoder 20 are configured in such a way that they support several transforms. For example, the transforms supported by the encoder 10 and the decoder 20 may include:
[0113] o DCT-II (or DCT-III), where DCT stands for Discrete Cosine Transform
[0114] o DST-IV, where DST stands for discrete sine transform
[0115] o DCT-IV
[0116] o DST-VII
[0117] oIdentity Transformation (IT)
[0118] Naturally, while the transformer 28 will support all forward transformed versions of these transforms, the decoder 20 or inverse transformer 54 will support their corresponding backward or inverse versions:
[0119] oInverse DCT-II (or Inverse DCT-III)
[0120] oInverse DST-IV
[0121] oInverse DCT-IV
[0122] o Inverse DST-VII o Identity Transformation (IT)
[0123] The subsequent description provides more details on which transforms can be supported by the encoder 10 and decoder 20. In any case, it should be noted that the set of supported transforms may include only one transform, such as a spectral to spatial or spatial to spectral transform.
[0124] As already outlined above, Figures 6 to 8
[0045] The examples have been presented as examples in which the inventive concepts described further below may be implemented in order to form specific examples for encoders and decoders according to the present application. Figure 6 and Figure 7 The encoder and decoder of may represent possible implementations of the encoder and decoder described below, respectively. However, Figure 6 and Figure 7 However, an encoder according to an embodiment of the present application may use the concepts outlined in more detail below to perform block-based encoding of picture 12, and with Figure 6 The encoder differs, for example, in that it is not a video encoder but a still picture encoder, in that it does not support inter-frame prediction, or in that it uses a different Figure 8 Likewise, a decoder according to an embodiment of the present application may perform block-based decoding of a picture 12' from a data stream 14 using the coding concepts outlined further below, but may be similar to Figure 7 The decoder 20 of the embodiment of the present invention differs in that it is not a video decoder but a still picture decoder, that it does not support intra-frame prediction, or that it decodes in a different manner than that described with respect to FIG. Figure 8 The picture 12 ′ is divided into blocks in the manner described and / or its prediction residual is derived from the data stream 14 , for example not in the transform domain but in the spatial domain.
[0125] Figure 1 A video encoder 100 for encoding a video into a video data stream according to an embodiment is illustrated. The video encoder 100 is configured to generate a video data stream.
[0126] Figure 2An apparatus 200 for receiving an input video data stream according to an embodiment is illustrated. The input video data stream has video encoded therein. The apparatus 200 is configured to generate an output video data stream from the input video data stream.
[0127] Figure 3 A video decoder 300 for receiving a video data stream having video stored therein according to an embodiment is illustrated. The video decoder 300 is configured to decode video from the video data stream.
[0128] Furthermore, a system according to an embodiment is provided. The system comprises Figure 2 The device and Figure 3 Video decoder. Figure 3 The video decoder (300) is configured to receive Figure 2 The output video data stream of the device (200) is provided. Figure 3 The video decoder 300 is configured to Figure 2 The output video data stream of the apparatus 200 is decoded video.
[0129] In an embodiment, the system may further include, for example Figure 1 The video encoder 100. For example, Figure 2 The apparatus 200 may be configured to Figure 1 The video encoder 100 receives a video data stream as an input video data stream.
[0130] The (optional) intermediate device 210 of the apparatus 200 may, for example, be configured to receive a video data stream as an input video data stream from the video encoder 100 and to generate an output video data stream from the input video data stream. For example, the intermediate device may, for example, be configured to modify (header / metadata) information of the input video data stream and / or may, for example, be configured to delete pictures from the input video data stream and / or may be configured to mix / splice the input video data stream with an additional second bitstream having a second video encoded therein.
[0131] The (optional) video decoder 221 may, for example, be configured to decode video from the output video data stream.
[0132] The (optional) hypothetical reference decoder 222 may, for example, be configured to determine timing information of the video from the output video data stream, or may, for example, be configured to determine buffer information of a buffer in which the video or a portion of the video is to be stored.
[0133] The system includes Figure 1 The video encoder 101 and Figure 2 Video decoder 151.
[0134] The video encoder 101 is configured to generate an encoded video signal. The video decoder 151 is configured to decode the encoded video signal to reconstruct pictures of the video.
[0135] Hereinafter, specific embodiments are described.
[0136] In HEVC, the comments in the extraction process specification describe the following processing of nested SEI messages:
[0137] A "smart" bitstream extractor may include appropriate non-scalable nested buffered pictures SEI messages, non-scalable nested picture timing SEI messages, and non-scalable nested decoding unit information SEI messages in the extracted sub-bitstream, provided that the SEI messages applicable to the sub-bitstream appeared as scalable nesting SEI messages in the original bitstream.
[0138] In VVC, the envisaged design properly has normatively specified behaviour, for example as follows with respect to the extraction process as defined in JVET-P2001-vC, to which embodiments of the present invention have been added.
[0139] Sub-bitstream extraction process
[0140] The input to this process is the bitstream inBitstream, the target OLS index targetOlsIdx and the target highest TemporalId value tIdTarget.
[0141] The output of this process is the sub-bitstream outBitstream.
[0142] Bitstream conformance requirements for input bitstreams. Any output sub-bitstream (i.e., the output of the process and bitstream specified in this clause) with targetOlsIdx equal to the index into the list of OLSs specified by the VPS and tIdTarget equal to any value in the range 0 to 6 (inclusive) as input, and meeting the following conditions, shall be a conforming bitstream:
[0143] - The output sub-bitstream contains at least one VCL NAL unit whose nuh_layer_id is equal to each of the nuh_layer_id values in LayerIdInOls[targetOlsIdx].
[0144] - The output sub-bitstream contains at least one VCL NAL unit whose TemporalId is equal to tIdTarget.
[0145] NOTE - A conforming bitstream contains one or more coded slice NAL units with TemporalId equal to 0, but does not necessarily contain a coded slice NAL unit with nuh_layer_id equal to 0.
[0146] The output sub-bitstream OutBitstream is exported as follows:
[0147] - The bitstream outBitstream is set to be the same as the bitstream INBITSTREAM.
[0148] - Remove from outBitstream all NAL units with a TemporalId greater than tIdTarget.
[0149] - Remove from outBitstream all NAL units whose nal_unit_type is not equal to any of VPS_NUT, DPS_NUT and EOB_NUT and whose nuh_layer_id is not included in the LayerIdInOls[targetOlsIdx] list.
[0150] - Remove from outBitstream all SEINAL units containing scalable nesting SEI messages with nesting_ols_flag equal to 1 and no value of i in the range of 0 to nesting_num_olss_minus1 (inclusive) such that NestingOlsIdx[i] is equal to targetOlsIdx.
[0151] - When targetOlsIdx is greater than 0, remove from outBitstream all SEI NAL units containing non-scalable nested SEI messages with payloadType equal to 0 (buffering period), 1 (picture timing), or 130 (decoding unit information).
[0152] According to a specific embodiment:
[0153] - When outBitstream contains a SEINAL unit (containing a scalable nesting SEI message with nesting_ols_flag equal to 1 and applicable to outBitstream (NestingOlsIdx[i] equal to targetOlsIdx), do the following:
[0154] - Extract the appropriate non-scalable nested SEI messages with payloadType equal to 0 (buffering period), 1 (picture timing), or 130 (decoding unit information) from the scalable nested SEI messages and put these messages into outBitstream.
[0155] - Remove all SEI NAL units containing scalable nested SEI messages from outBitstream
[0156] In the following, the presence of scalable nested SEI messages for OLS in the bitstream is described.
[0157] According to an embodiment, a video data stream having video encoded therein is provided. The video data stream includes an indication of whether one or more scalable nested supplemental enhancement information messages are present within the video data stream, the one or more scalable nested supplemental enhancement information messages including timing information for each of one or more output layer sets.
[0158] In an embodiment, the indication is a parameter set flag. The video data stream may, for example, include a parameter set flag that indicates whether one or more scalable nested supplemental enhancement information messages are present in the video data stream, the one or more scalable nested supplemental enhancement information messages including timing information for each of the one or more output layer sets.
[0159] According to an embodiment, a video data stream according to claim 2. The sequence parameter set of the video data stream may for example comprise a parameter set flag.
[0160] In an embodiment, the parameter set flag is sps_ols_nest_timing_present_flag.
[0161] According to an embodiment, the video data stream may, for example, include a further supplemental enhancement information message, which may, for example, include a parameter set flag. The parameter set flag indicates whether one or more scalable nested supplemental enhancement information messages for each of the one or more output layer sets are present in the video data stream.
[0162] In an embodiment, the timing information may include, for example, at least one of picture timing information, buffering period information, and decoding unit information.
[0163] According to an embodiment, the timing information is timing information for a hypothetical reference decoder.
[0164] Furthermore, according to an embodiment, a video encoder for encoding a video into a video data stream is provided, such that the video data stream has the video encoded therein. The video encoder is configured to generate the video data stream such that the video data stream includes an indication of whether one or more scalable nested supplemental enhancement information messages are present within the video data stream, the one or more scalable nested supplemental enhancement information messages including timing information for each of one or more output layer sets.
[0165] According to an embodiment, the indication is a parameter set flag. The video encoder may, for example, be configured to generate a video data stream such that the video data stream may, for example, include a parameter set flag indicating whether one or more scalable nested supplemental enhancement information messages are present within the video data stream, the one or more scalable nested supplemental enhancement information messages including timing information for each of the one or more output layer sets.
[0166] In an embodiment, the video encoder may, for example, be configured to generate a video data stream such that a sequence parameter set of the video data stream may, for example, comprise a parameter set flag.
[0167] According to an embodiment, the video encoder may, for example, be configured to generate the video data stream such that the parameter set flag is sps_ols_nest_timing_present_flag.
[0168] In an embodiment, the video encoder may be configured to generate a video data stream, for example, such that the video data stream may include a further supplemental enhancement information message, which may include a parameter set flag. The video encoder may be configured to generate the video data stream, for example, such that the parameter set flag indicates whether one or more scalable nested supplemental enhancement information messages for each of the one or more output layer sets are present within the video data stream.
[0169] According to an embodiment, the timing information may include, for example, at least one of picture timing information, buffering period information, and decoding unit information.
[0170] In an embodiment, the timing information is timing information for a hypothetical reference decoder.
[0171] Furthermore, according to an embodiment, an apparatus is provided for receiving a video data stream as an input bitstream, wherein the video data stream has video encoded therein, processing the input bitstream to obtain sub-bitstreams, and providing an indication of whether one or more scalable nested supplemental enhancement information messages are present within the video data stream, the one or more scalable nested supplemental enhancement information messages including timing information for each of one or more output layer sets.
[0172] According to an embodiment, the indication is a parameter set flag. Apparatus for processing a video data stream may, for example, include a parameter set flag indicating whether one or more scalable nested supplemental enhancement information messages are present within the video data stream, the one or more scalable nested supplemental enhancement information messages including timing information for each of one or more output layer sets.
[0173] In an embodiment, a sequence parameter set of a video data stream may, for example, comprise a parameter set flag.
[0174] According to an embodiment, the parameter set flag is sps_ols_nest_timing_present_flag.
[0175] In an embodiment, the video data stream may, for example, include a further Supplemental Enhancement Information message, which may, for example, include a parameter set flag. The parameter set flag indicates whether one or more scalable nested Supplemental Enhancement Information messages for each of one or more output layer sets are present in the video data stream. The apparatus may, for example, be configured to process the further Supplemental Enhancement Information message.
[0176] According to an embodiment, the timing information may include, for example, at least one of picture timing information, buffering period information, and decoding unit information.
[0177] According to an embodiment, if one or more scalable nested supplemental enhancement information messages including timing information are present in a video data stream, the timing information may, for example, include picture timing information for each output layer set in one or more output layer sets, then the device may, for example, be configured to replace the picture timing information of a non-scalable nested picture timing supplemental enhancement information message. If one or more scalable nested supplemental enhancement information messages including timing information are present in a video data stream, the timing information includes buffering period information for each output layer set in one or more output layer sets, then the device may, for example, be configured to replace the buffering period information of a non-scalable nested picture timing supplemental enhancement information message. If one or more scalable nested supplemental enhancement information messages including timing information are present in a video data stream, the timing information includes decoding unit information for each output layer set in one or more output layer sets, then the device may, for example, be configured to replace the decoding unit information of the non-scalable nested picture timing supplemental enhancement information message.
[0178] In an embodiment, the timing information is timing information for a hypothetical reference decoder.
[0179] In an embodiment, the apparatus may, for example, be configured to decode a sub-bitstream to decode the video.
[0180] Furthermore, according to an embodiment, a system for encoding video into a video data stream and decoding the video is provided. The system includes: a video encoder as described above and an apparatus as described above. The video encoder can, for example, be configured to encode the video into a video data stream, such that the video data stream has the video encoded therein. The apparatus can, for example, be configured to receive the video data stream as an input bitstream. Furthermore, the apparatus can, for example, be configured to process the input bitstream to obtain a sub-bitstream. Furthermore, the apparatus can, for example, be configured to decode the sub-bitstream to decode the video.
[0181] The VVC draft specification contains the definition of OLS in VPS, which can also be used for HRD-based conformance testing of OLS sub-bitstreams based on the corresponding HRD SEI messages (BP, PT, DUI) in the bitstream in a nested form (extensible nested SEI messages). When defining OLS, it is crucial to ensure that the corresponding HRD SEI messages for these OLS are in the bitstream to enable conformance testing.
[0182] Therefore, part of the invention is to indicate in the bitstream (either a parameter set flag such as eg sps_ols_scal_nest_timing_present_flag in SPS, or a new SEI message containing such a flag) that a scalable nesting SEI message indicating all OLSs should be present within the bitstream.
[0183] Hereinafter, the concept of how HRD SEI is applied to a sub-bitstream is described.
[0184] According to an embodiment, a video data stream having video encoded therein is provided. An indication within the video data stream indicates whether timing information for a sub-bitstream is to be obtained from one or more non-scalable nested picture timing supplemental enhancement information messages of the video data stream.
[0185] In an embodiment, if the indication indicates that timing information for the sub-bitstream is not to be obtained from one or more non-scalable nested picture timing supplemental enhancement information messages, this indicates that the one or more non-scalable nested picture timing supplemental enhancement information messages are to be replaced by one or more scalable nested picture timing supplemental enhancement information messages.
[0186] In an embodiment, the indication is a flag.One of the one or more scalable nested supplemental enhancement information messages may, for example, include a flag indicating whether timing information for a sub-bitstream may be configured to be obtained from one or more non-scalable nested picture timing supplemental enhancement information messages.
[0187] According to an embodiment, the flag is use_orig_pic_timing_flag.
[0188] In an embodiment, the video data stream may include, for example, a video parameter set. The indication is a flag. The video parameter set includes a flag indicating whether timing information for a sub-bitstream can be obtained, for example, from one or more non-scalable nested picture timing supplemental enhancement information messages.
[0189] According to an embodiment, the flag is same_pic_timing_within_ols_flag.
[0190] In an embodiment, the flag is general_same_pic_timing_in_all_ols_flag.
[0191] According to an embodiment, an indication is given that each of one or more non-scalable nested picture timing supplemental enhancement information messages in each access unit of one or more access units applies to an access unit for any output layer set in a video data stream and that no scalable nested picture timing supplemental enhancement information message exists, or an indication is given that the non-scalable nested picture timing supplemental enhancement information message in each access unit of one or more access units may or may not apply to an access unit for any output layer set in a video data stream and that a scalable nested picture timing supplemental enhancement information message may exist.
[0192] In an embodiment, if the indication indication can be configured, for example, to obtain timing information for a sub-bitstream from one or more non-scalable nested picture timing supplemental enhancement information messages, then within an access unit of the bitstream, at least one of the one or more scalable nested picture timing supplemental enhancement information messages appears before the one or more non-scalable nested picture timing supplemental enhancement information messages.
[0193] According to an embodiment, the indication is a constraint flag. The one or more non-scalable nested picture timing supplemental enhancement information messages include a constraint flag indicating whether the one or more non-scalable nested picture timing supplemental enhancement information messages apply to at least one sub-bitstream of the one or more sub-bitstreams.
[0194] In an embodiment, the indication is a first indication. A second indication within the video data stream indicates whether to further obtain timing information for the sub-bitstream from one or more non-scalable nested buffering period supplemental enhancement information messages of the video data stream; and / or a third indication within the video data stream indicates whether to further obtain timing information for the sub-bitstream from one or more non-scalable nested decoding unit supplemental enhancement information messages of the video data stream.
[0195] According to an embodiment, the indication within the video data stream further indicates whether to further obtain timing information for the sub-bitstream from one or more non-scalable nested buffering period supplemental enhancement information messages of the video data stream; and / or the indication within the video data stream further indicates whether to further obtain timing information for the sub-bitstream from one or more non-scalable nested decoding unit supplemental enhancement information messages of the video data stream.
[0196] Furthermore, according to an embodiment, a video data stream having video encoded therein is provided. A first indication within the video data stream indicates whether timing information for a sub-bitstream is to be obtained from one or more non-scalable nested buffering period supplemental enhancement information messages of the video data stream. And / or a second indication within the video data stream indicates whether timing information for the sub-bitstream is to be obtained from one or more non-scalable nested decoding unit supplemental enhancement information messages of the video data stream.
[0197] Furthermore, according to an embodiment, a video data stream includes one or more non-scalable nested supplementary enhancement information messages including timing information. If the video data stream includes a scalable nested supplementary enhancement information message including timing information, this indicates that, depending on the scalable nested supplementary enhancement information message, all of the one or more non-scalable nested timing information supplementary enhancement information messages are to be replaced with scalable nested supplementary enhancement information messages including timing information (e.g., the timing information is picture timing information, buffering period information, or decoding unit information). Or, a subset of at least one of the one or more non-scalable nested timing information enhancement information messages is to be replaced with a scalable nested supplementary enhancement information message including timing information (e.g., the timing information is picture timing information, buffering period information, or decoding unit information).
[0198] According to an embodiment, the one or more non-scalable nested supplementary enhancement information messages including timing information are one or more non-scalable nested picture timing supplementary enhancement information messages, while the scalable nested supplementary enhancement information message including timing information is a scalable nested picture timing supplementary enhancement information message. Alternatively, the one or more non-scalable nested supplementary enhancement information messages including timing information are one or more non-scalable nested buffering period supplementary enhancement information messages, while the scalable nested supplementary enhancement information message including timing information is a scalable nested buffering period supplementary enhancement information message. Alternatively, the one or more non-scalable nested supplementary enhancement information messages including timing information are one or more non-scalable nested decoding unit supplementary enhancement information messages, while the scalable nested supplementary enhancement information message including timing information is a scalable nested decoding unit supplementary enhancement information message.
[0199] In an embodiment, if the video data stream may, for example, include scalable nested supplemental enhancement information messages including timing information, then the scalable nested supplemental enhancement information messages including timing information occur before one or more non-scalable nested supplemental enhancement information messages including timing information within an access unit of the bitstream.
[0200] According to an embodiment, the timing information is timing information for a hypothetical reference decoder.
[0201] In an embodiment, a sub-bitstream is dependent on a set of output layers and / or on a sub-layer, and / or on a sub-picture, and / or on a subset of decoding units.
[0202] Furthermore, according to an embodiment, a video encoder for encoding video into a video data stream is provided, such that the video data stream has the video encoded therein. The video encoder is configured to generate the video data stream such that an indication within the video data stream indicates whether timing information for a sub-bitstream is to be obtained from one or more non-scalable nested picture timing supplemental enhancement information messages of the video data stream.
[0203] According to an embodiment, if the indication indicates that timing information for the sub-bitstream is not to be obtained from one or more non-scalable nested picture timing supplemental enhancement information messages, this indicates that the one or more non-scalable nested picture timing supplemental enhancement information messages are to be replaced by one or more scalable nested picture timing supplemental enhancement information messages.
[0204] In an embodiment, the indication is a flag. The video encoder may, for example, be configured to generate the video data stream such that one of the one or more scalable nested supplemental enhancement information messages may, for example, include a flag indicating whether timing information for a sub-bitstream may be obtained, for example, from one or more non-scalable nested picture timing supplemental enhancement information messages.
[0205] According to an embodiment, the flag is use_orig_pic_timing_flag.
[0206] In an embodiment, a video data stream may, for example, include a video parameter set. The indication may, for example, be a flag. The video encoder may, for example, be configured to generate the video data stream such that the video parameter set includes a flag indicating whether timing information for a sub-bitstream may be obtained, for example, from one or more non-scalable nested picture timing supplemental enhancement information messages.
[0207] According to an embodiment, the flag is same_pic_timing_within_ols_flag.
[0208] In an embodiment, the flag is general_same_pic_timing_in_all_ols_flag.
[0209] According to an embodiment, an indication is given that each of one or more non-scalable nested picture timing supplemental enhancement information messages in each access unit of one or more access units applies to an access unit for any output layer set in a video data stream and that no scalable nested picture timing supplemental enhancement information message exists, or an indication is given that the non-scalable nested picture timing supplemental enhancement information message in each access unit of one or more access units may or may not apply to an access unit for any output layer set in a video data stream and that a scalable nested picture timing supplemental enhancement information message may exist.
[0210] In an embodiment, if the indication indication can be configured, for example, to obtain timing information for a sub-bitstream from one or more non-scalable nested picture timing supplemental enhancement information messages, then within an access unit of the bitstream, at least one of the one or more scalable nested picture timing supplemental enhancement information messages appears before the one or more non-scalable nested picture timing supplemental enhancement information messages.
[0211] According to an embodiment, the indication is a constraint flag. The video encoder is configured to generate a video data stream such that the one or more non-scalable nested picture timing supplemental enhancement information messages include a constraint flag indicating whether the one or more non-scalable nested picture timing supplemental enhancement information messages apply to at least one sub-bitstream of the one or more sub-bitstreams.
[0212] In an embodiment, the indication is a first indication. A second indication within the video data stream indicates whether to further obtain timing information for the sub-bitstream from one or more non-scalable nested buffering period supplemental enhancement information messages of the video data stream; and / or a third indication within the video data stream indicates whether to further obtain timing information for the sub-bitstream from one or more non-scalable nested decoding unit supplemental enhancement information messages of the video data stream.
[0213] According to an embodiment, the indication within the video data stream further indicates whether to further obtain timing information for the sub-bitstream from one or more non-scalable nested buffering period supplemental enhancement information messages of the video data stream; and / or the indication within the video data stream further indicates whether to further obtain timing information for the sub-bitstream from one or more non-scalable nested decoding unit supplemental enhancement information messages of the video data stream.
[0214] Furthermore, according to an embodiment, a video encoder for encoding video into a video data stream is provided, such that the video data stream has the video encoded therein. A first indication within the video data stream indicates whether timing information for a sub-bitstream is to be obtained from one or more non-scalable nested buffering period supplemental enhancement information messages of the video data stream. And / or a second indication within the video data stream indicates whether timing information for the sub-bitstream is to be obtained from one or more non-scalable nested decoding unit supplemental enhancement information messages of the video data stream.
[0215] In addition, according to an embodiment, a video encoder for encoding a video into a video data stream is provided, so that the video data stream has the video encoded therein. The video encoder is used to generate a video data stream so that the video data stream includes one or more non-scalable nested supplementary enhancement information messages including timing information. If the video data stream includes a scalable nested supplementary enhancement information message including timing information, this indicates that, depending on the scalable nested supplementary enhancement information message, all of the one or more non-scalable nested timing information supplementary enhancement information messages are to be replaced with scalable nested supplementary enhancement information messages including timing information (for example, the timing information is picture timing information or buffering period information or decoding unit information). Or: a subset of at least one of the one or more non-scalable nested timing information supplementary enhancement information messages is to be replaced with a scalable nested supplementary enhancement information message including timing information (for example, the timing information is picture timing information or buffering period information or decoding unit information).
[0216] According to an embodiment, a video encoder may be configured, for example, to generate a video data stream such that one or more non-scalable nested supplementary enhancement information messages including timing information are one or more non-scalable nested picture timing supplementary enhancement information messages, and a scalable nested supplementary enhancement information message including timing information is a scalable nested picture timing supplementary enhancement information message. Alternatively, a video encoder may be configured, for example, to generate a video data stream such that one or more non-scalable nested supplementary enhancement information messages including timing information are one or more non-scalable nested buffering period supplementary enhancement information messages, and a scalable nested supplementary enhancement information message including timing information is a scalable nested buffering period supplementary enhancement information message. Alternatively, a video encoder may be configured, for example, to generate a video data stream such that one or more non-scalable nested supplementary enhancement information messages including timing information are one or more non-scalable nested decoding unit supplementary enhancement information messages, and a scalable nested supplementary enhancement information message including timing information is a scalable nested decoding unit supplementary enhancement information message.
[0217] In an embodiment, if a video data stream includes a scalable nested supplemental enhancement information message including timing information, the video encoder is configured to generate the video data stream such that the scalable nested supplemental enhancement information message including timing information occurs before one or more non-scalable nested supplemental enhancement information messages including timing information within an access unit of the bitstream.
[0218] According to an embodiment, the timing information is timing information for a hypothetical reference decoder.
[0219] In an embodiment, a sub-bitstream is dependent on a set of output layers and / or on a sub-layer, and / or on a sub-picture, and / or on a subset of decoding units.
[0220] Furthermore, according to an embodiment, an apparatus is provided for receiving a video data stream as an input bitstream, wherein the video data stream has video encoded therein. The apparatus is configured to process the input bitstream to obtain a sub-bitstream. An indication within the video data stream indicates whether timing information for the sub-bitstream is to be obtained from one or more non-scalable nested picture timing supplemental enhancement information messages of the video data stream.
[0221] According to an embodiment, if the indication indicates that timing information for the sub-bitstream is not to be obtained from one or more non-scalable nested picture timing supplemental enhancement information messages, the device can, for example, be configured to replace one or more non-scalable nested picture timing supplemental enhancement information messages with one or more scalable nested picture timing supplemental enhancement information messages.
[0222] According to an embodiment, the indication may be a flag. One of the one or more scalable nested supplemental enhancement information messages may, for example, include a flag indicating whether timing information for a sub-bitstream may be configured to be obtained from one or more non-scalable nested picture timing supplemental enhancement information messages.
[0223] In an embodiment, the flag is use_orig_pic_timing_flag.
[0224] According to an embodiment, the video data stream may include a video parameter set, for example. The indication may be, for example, a flag. The video parameter set may include, for example, a flag indicating whether timing information for a sub-bitstream may be obtained from, for example, one or more non-scalable nested picture timing supplemental enhancement information messages.
[0225] In an embodiment, the flag is same_pic_timing_within_ols_flag.
[0226] According to an embodiment, the flag is general_same_pic_timing_in_all_ols_flag.
[0227] In an embodiment, the indication indicates that each of the one or more non-scalable nested picture timing supplemental enhancement information messages in each access unit of one or more access units applies to the access unit for any output layer set in the video data stream and no scalable nested picture timing supplemental enhancement information message exists, or indicates that the non-scalable nested picture timing supplemental enhancement information message in each access unit of the one or more access units may or may not apply to the access unit for any output layer set in the video data stream and a scalable nested picture timing supplemental enhancement information message may exist.
[0228] According to an embodiment, if an indication indicates that a non-scalable nested picture timing supplemental enhancement information message in each access unit of one or more access units may or may not apply to access units for any set of output layers in a video data stream, and a scalable nested picture timing supplemental enhancement information message may be present, the apparatus is configured to remove from the input bitstream or from the sub-bitstream all supplemental enhancement information network abstraction layer units including the non-scalable nested supplemental enhancement information message having picture timing content.
[0229] In an embodiment, if the indication indicates that the timing information for the sub-bitstream can be, for example, configured to be obtained from one or more non-scalable nested picture timing supplemental enhancement information messages, then before the device can, for example, be configured to process the one or more non-scalable nested picture timing supplemental enhancement information messages within an access unit, the device can, for example, be configured to process at least one of the one or more scalable nested picture timing supplemental enhancement information messages, the one or more scalable nested picture timing supplemental enhancement information messages occurring before the one or more non-scalable nested picture timing supplemental enhancement information messages within an access unit of the bitstream.
[0230] According to an embodiment, the indication is a constraint flag. The apparatus may, for example, be configured to process one or more non-scalable nested picture timing supplemental enhancement information messages, the non-scalable nested picture timing supplemental enhancement information messages including a constraint flag indicating whether the one or more non-scalable nested picture timing supplemental enhancement information messages apply to at least one sub-bitstream of the one or more sub-bitstreams.
[0231] In an embodiment, the indication is a first indication. A second indication within the video data stream indicates whether to further obtain timing information for the sub-bitstream from one or more non-scalable nested buffering period supplemental enhancement information messages of the video data stream; and / or a third indication within the video data stream indicates whether to further obtain timing information for the sub-bitstream from one or more non-scalable nested decoding unit supplemental enhancement information messages of the video data stream.
[0232] According to an embodiment, the indication within the video data stream further indicates whether to further obtain timing information for the sub-bitstream from one or more non-scalable nested buffering period supplemental enhancement information messages of the video data stream; and / or the indication within the video data stream further indicates whether to further obtain timing information for the sub-bitstream from one or more non-scalable nested decoding unit supplemental enhancement information messages of the video data stream.
[0233] Furthermore, according to an embodiment, an apparatus is provided for receiving a video data stream as an input bitstream, wherein the video data stream has video encoded therein. The apparatus is configured to process the input bitstream to obtain a sub-bitstream. A first indication within the video data stream indicates whether timing information for the sub-bitstream is to be obtained from one or more non-scalable nested buffering period supplemental enhancement information messages of the video data stream. And / or a second indication within the video data stream indicates whether timing information for the sub-bitstream is to be obtained from one or more non-scalable nested decoding unit supplemental enhancement information messages of the video data stream.
[0234] In addition, according to an embodiment, an apparatus for receiving a video data stream as an input bitstream is provided, wherein the video data stream has video encoded therein. The apparatus is configured to process the input bitstream to obtain a sub-bitstream. The video data stream includes one or more non-scalable nested supplementary enhancement information messages including timing information. If the video data stream includes a scalable nested supplementary enhancement information message including timing information, the apparatus relies on the scalable nested supplementary enhancement information message to: replace all one or more non-scalable nested timing information supplementary enhancement information messages with scalable nested supplementary enhancement information messages including timing information (e.g., the timing information is picture timing information or buffering period information or decoding unit information). Or, replace a subset of at least one of the one or more non-scalable nested timing information supplementary enhancement information messages with a scalable nested supplementary enhancement information message including timing information (e.g., the timing information is picture timing information or buffering period information or decoding unit information).
[0235] According to an embodiment, the one or more non-scalable nested supplementary enhancement information messages including timing information are one or more non-scalable nested picture timing supplementary enhancement information messages, while the scalable nested supplementary enhancement information message including timing information is a scalable nested picture timing supplementary enhancement information message. Alternatively, the one or more non-scalable nested supplementary enhancement information messages including timing information are one or more non-scalable nested buffering period supplementary enhancement information messages, while the scalable nested supplementary enhancement information message including timing information is a scalable nested buffering period supplementary enhancement information message. Alternatively, the one or more non-scalable nested supplementary enhancement information messages including timing information are one or more non-scalable nested decoding unit supplementary enhancement information messages, while the scalable nested supplementary enhancement information message including timing information is a scalable nested decoding unit supplementary enhancement information message.
[0236] In an embodiment, if the video data stream may, for example, include a scalable nested supplemental enhancement information message including timing information, then within an access unit of the bitstream, the scalable nested supplemental enhancement information message including timing information occurs before one or more non-scalable nested supplemental enhancement information messages including timing information. The apparatus may, for example, be configured to process the scalable nested supplemental enhancement information message including timing information before the one or more non-scalable nested supplemental enhancement information messages including timing information.
[0237] According to an embodiment, the timing information is timing information for a hypothetical reference decoder.
[0238] In an embodiment, a sub-bitstream is dependent on a set of output layers and / or on a sub-layer, and / or on a sub-picture, and / or on a subset of decoding units.
[0239] According to an embodiment, the apparatus may, for example, be configured to decode a sub-bitstream to decode a video.
[0240] Furthermore, according to an embodiment, a system for encoding a video into a video data stream and decoding the video is provided. The system includes a video encoder as described above and an apparatus as described above. The video encoder can, for example, be configured to encode the video into a video data stream such that the video data stream has the video encoded therein. The apparatus can, for example, be configured to receive the video data stream as an input bitstream. Furthermore, the apparatus can, for example, be configured to process the input bitstream to obtain a sub-bitstream. Furthermore, the apparatus can, for example, be configured to decode the sub-bitstream to decode the video.
[0241] The VVC draft specification contains SEI messages that control the HRD timing behavior, namely, the buffering period SEI message and the picture timing SEI message when the bitstream is decoded.
[0242] Currently, the VVC draft specification already includes all picture timing SEI messages that apply to multiple target temporal IDs, i.e., temporal scalability does not necessarily require scalability nesting of additional timing information. Therefore, when extracting sub-bitstreams (e.g., in layered scenarios via OLS or spatially via sub-picture extraction), in some cases it may not be necessary to modify / exchange picture timing SEI messages in this way.
[0243] In the following, output layer set extraction is described.
[0244] In one embodiment, signaling is added indicating that the picture timing SEI messages in the bitstream apply to any sub-bitstream (as defined / corresponding to a certain output layer set (OLS) A) of the bitstream (as defined / corresponding to some OLSB), and only the buffering period SEI messages will be replaced by their scalable nested counterparts. Example syntax:
[0245]
[0246] Until now, the scalable nesting SEI message followed the HRD SEI message within an access unit. Therefore, as part of the above embodiment, when use_orig_pic_timing_flag is equal to 1, the scalable nesting SEI message containing the OLS-specific HRD SEI message must appear before the corresponding PT SEI message (the message to be retained during extraction) within the access unit in bitstream order.
[0247] Or, in an alternative embodiment, the indication is in the VPS as a constraint flag:
[0248]
[0249] Alternatively, in an alternative embodiment, the indication is provided as a constraint flag in the picture timing SEI message.
[0250] This has changed the extraction process.
[0251] In the following, the sub-bitstream extraction process is described.
[0252] The input to this process is the bitstream inBitstream, the target OLS index targetOlsIdx and the target highest TemporalId value tIdTarget.
[0253] The output of this process is the sub-bitstream outBitstream.
[0254] Bitstream conformance requirements for input bitstreams. Any output sub-bitstream (i.e., the output of the process and bitstream specified in this clause) with targetOlsIdx equal to the index into the list of OLSs specified by the VPS and tIdTarget equal to any value in the range 0 to 6 (inclusive) as input, and meeting the following conditions, shall be a conforming bitstream:
[0255] - The output sub-bitstream contains at least one VCL NAL unit whose nuh_layer_id is equal to each of the nuh_layer_id values in LayerIdInOls[targetOlsIdx].
[0256] - The output sub-bitstream contains at least one VCL NAL unit whose TemporalId is equal to tIdTarget.
[0257] NOTE - A conforming bitstream contains one or more coded slice NAL units with TemporalId equal to 0, but does not necessarily contain a coded slice NAL unit with nuh_layer_id equal to 0.
[0258] The output sub-bitstream OutBitstream is exported as follows:
[0259] - The bitstream outBitstream is set to be the same as the bitstream INBITSTREAM.
[0260] - Remove from outBitstream all NAL units with a TemporalId greater than tIdTarget.
[0261] - Remove from outBitstream all NAL units whose nal_unit_type is not equal to any of VPS_NUT, DPS_NUT and EOB_NUT and whose nuh_layer_id is not included in the LayerIdInOls[targetOlsIdx] list.
[0262] - Remove from outBitstream all SEINAL units containing scalable nesting SEI messages with nesting_ols_flag equal to 1 and no value of i in the range of 0 to nesting_num_olss_minus1 (inclusive) such that NestingOlsIdx[i] is equal to targetOlsIdx.
[0263] - When targetOlsIdx is greater than 0, remove from outBitstream all SEI NAL units containing non-scalable nested SEI messages with payloadType equal to 0 (buffering period), 1 (picture timing), or 130 (decoding unit information).
[0264] - When targetOlsIdx is greater than 0 and use_orig_pic_timing_flag / same_pic_timing_within_ols_flag is equal to 0, remove from outBitstream all SEINAL units containing non-scalable nested SEI messages with payloadType equal to 1 (picture timing).
[0265] - When outBitstream contains a SEINAL unit that contains an extensible nesting SEI message and is applicable to outBitstream (NestingOlsIdx[i] is equal to targetOlsIdx), do the following:
[0266] - When use_orig_pic_timing_flag / same_pic_timing_within_ols_flag is equal to 0, extract the appropriate non-scalable nested SEI messages with payloadType equal to 0 (buffering period), 1 (picture timing), or 130 (decoding unit information) from the scalable nested SEI messages and put these messages into outBitstream.
[0267] Otherwise, (when use_orig_pic_timing_flag / same_pic_timing_within_ols_flag is equal to 1), extract the appropriate non-scalable nesting SEI messages with payloadType equal to 0 (buffering period) or 130 (decoding unit information) from the scalable nesting SEI messages at presentation time, which are applicable to outBitstream (NestingOlsIdx[i] equal to targetOlsIdx), and put these messages into outBitstream.
[0268] - Remove all SEI NAL units containing scalable nested SEI messages from outBitstream
[0269] For example, an indication such as a flag such as general_same_pic_timing_in_all_ols_flag or such as same_pic_timing_within_ols_flag may be utilized.
[0270] general_same_pic_timing_in_all_ols_flag (or same_pic_timing_within_ols_flag) is equal to a first value (e.g., equal to 1), specifying that the non-scalable nested PT SEI message in each AU applies to the AU of any OLS in the bitstream, and there is no scalable nested PT SEI message. general_same_pic_timing_in_all_ols_flag is equal to a second value (e.g., 0), specifying that the non-scalable nested PT SEI message in each AU may or may not apply to the AU of any OLS in the bitstream, and there may be a scalable nested PT SEI message.
[0271] For example, when general_same_pic_timing_in_all_ols_flag is equal to the second value (eg, 0), all SEI NAL units including non-scalable nested SEI messages with payloadType equal to 1 (PT) are removed from the input bitstream or output bitstream (eg, outBitstream / sub-bitstream).
[0272] In an embodiment, similarly, the buffering period SEI message may not need to be replaced in some cases by a scalable nested variant, whereby additional indications can be carried in the parameter set or the buffering period SEI message itself, so that the indicated timing also applies to the extracted sub-bitstream, and the extraction process will be further modified to maintain the original BP and PT SEI messages in dependence on the corresponding indications.
[0273] In the following, sub-image extraction is described.
[0274] In another embodiment that considers the case of sub-picture sub-bitstream extraction, corresponding signaling is added to the bitstream (e.g., in the syntax of the sub-picture nesting SEI message, the picture timing SEI message, the parameter set), which indicates that the picture timing SEI message in the bitstream applies to any sub-bitstream defined by a sub-picture or sub-picture set of the bitstream consisting of a combination of all sub-pictures contained in the bitstream, and only the buffering period SEI messages are used to be replaced by their sub-picture nesting counterparts.
[0275] Alternatively, also in some cases, the buffering period SEI message may not need to be replaced by a variant of sub-picture nesting, so that additional indications can be carried in the parameter set or the buffering period SEI message itself, so that the indicated timing also applies to the extracted sub-bitstream, and the extraction process will be further modified to rely on the corresponding indication to keep the original BP and PT SEI messages.
[0276] Hereinafter, removal of a non-scalable nested HRD SEI message based on the presence of a scalable nested HRD SEI message is described.
[0277] In another embodiment of the present invention, the scope of non-scalable nested HRD SEI messages (e.g., BP, PT, DUI SEI messages), i.e., whether some (e.g., only PT SEI messages) or all of them apply to the extractable sub-bitstream (e.g., OLS), is indicated by a corresponding replacement SEI message in the form of a scalable nested HRD SEI message, wherein the absence of such a message indicates that the corresponding non-scalable HRD SEI message applies to the OLS. As a result of this scope indication, the removal of the non-scalable nested SEI message during the sub-bitstream extraction process depends on the presence of a scalable nested SEI message that can be used as a replacement for the removed message, and in the absence of such a scalable nested SEI message, the non-scalable nested SEI message remains in the extracted sub-bitstream.
[0278] In another embodiment, in an access unit, the corresponding scalable nested HRD SEI message is placed before the non-scalable nested HRD SEI message in the bitstream order to simplify sequential processing of the bitstream during extraction.
[0279] Hereinafter, simplified expandable nesting will be described.
[0280] According to an embodiment, a video data stream having video encoded therein is provided. The video data stream includes a plurality of access units. For each access unit in the plurality of access units, if the access unit includes two or more scalable nested supplementary enhancement information messages, and the two or more scalable nested supplementary enhancement information messages include buffering period information and / or picture timing information for an output layer set, then the buffering period information and / or picture timing information for the output layer set in all two or more scalable nested supplementary enhancement information messages in the access unit are equal.
[0281] Furthermore, according to an embodiment, a video data stream having video encoded therein is provided. The video data stream includes a plurality of access units. For each access unit in the plurality of access units, if the access unit includes two or more scalable nested supplementary enhancement information messages, and the two or more scalable nested supplementary enhancement information messages include buffering period information and / or picture timing information for an output layer set, the video data stream includes an indication indicating whether the buffering period information and / or picture timing information for the output layer set in all two or more scalable nested supplementary enhancement information messages for the access unit are equal.
[0282] In addition, according to an embodiment, a video data stream having a video encoded therein is provided. The video data stream includes a plurality of access units. For each access unit of the plurality of access units, if the access unit includes three or more scalable nested supplementary enhancement information messages, two or more scalable nested supplementary enhancement information messages include buffering period information and / or picture timing information for an output layer set, then the buffering period information and / or picture timing information for the output layer set appears only in one scalable nested supplementary enhancement information message among the three or more scalable nested supplementary enhancement information messages, and in another scalable nested supplementary enhancement information message among the three or more scalable nested supplementary enhancement information messages that immediately follows the one scalable nested supplementary enhancement information message among the three or more scalable nested supplementary enhancement information messages.
[0283] In addition, according to an embodiment, a video data stream having a video encoded therein is provided. The video data stream includes a plurality of access units. For each access unit of the plurality of access units, if the access unit includes three or more scalable nested supplementary enhancement information messages, two or more scalable nested supplementary enhancement information messages include buffering period information and / or picture timing information for an output layer set, then the video data stream includes an indication indicating whether the buffering period information and / or picture timing information for the output layer set appears only in one of the three or more scalable nested supplementary enhancement information messages, and in another scalable nested supplementary enhancement information message of the three or more scalable nested supplementary enhancement information messages that immediately follows the one of the three or more scalable nested supplementary enhancement information messages.
[0284] In an embodiment, all scalable nested supplementary enhancement information messages having the same value of an identifier identifying an output layer set in an access unit may, for example, carry the same buffering period information and / or the same picture timing information.
[0285] According to an embodiment, for a specific access unit among a plurality of access units, any picture timing supplemental enhancement information messages applied to the layer set and sub-layer set of the output layer set may, for example, carry the same picture timing information. And / or, for a specific access unit among a plurality of access units, any buffering period supplemental enhancement information messages applied to the layer set and sub-layer set of the output layer set may, for example, carry the same buffering period information. And / or, for a specific access unit among a plurality of access units, any decoding unit supplemental enhancement information messages applied to the layer set and sub-layer set of the output layer set may, for example, carry the same decoding unit information.
[0286] In an embodiment, two Scalable Nested Supplemental Enhancement Information messages for a specific payload type in an access unit with the same value of an identifier identifying an output layer set may, for example, carry the same content.
[0287] In addition, according to an embodiment, a video encoder for encoding a video into a video data stream is provided, so that the video data stream has the video encoded therein. The video encoder is used to generate the video data stream so that the video data stream includes multiple access units. For each access unit in the multiple access units. If the access unit includes two or more scalable nested supplementary enhancement information messages, and the two or more scalable nested supplementary enhancement information messages include buffering period information and / or picture timing information for the output layer set, the video encoder is used to generate the video data stream so that the buffering period information and / or picture timing information for the output layer set in all two or more scalable nested supplementary enhancement information messages of the access unit are equal.
[0288] Furthermore, according to an embodiment, a video encoder for encoding a video into a video data stream is provided, such that the video data stream has the video encoded therein. The video encoder is configured to generate the video data stream such that the video data stream includes a plurality of access units. For each access unit of the plurality of access units, if the access unit includes two or more scalable nested supplementary enhancement information messages, and the two or more scalable nested supplementary enhancement information messages include buffering period information and / or picture timing information for an output layer set, the video encoder is configured to generate the video data stream such that the video data stream includes an indication of whether the buffering period information and / or picture timing information for the output layer set in all two or more scalable nested supplementary enhancement information messages of the access unit are equal.
[0289] In addition, according to an embodiment, a video encoder for encoding a video into a video data stream is provided, so that the video data stream has the video encoded therein. The video encoder is used to generate the video data stream so that the video data stream includes multiple access units. For each access unit of the multiple access units, if the access unit includes three or more scalable nested supplementary enhancement information messages, two or more scalable nested supplementary enhancement information messages include buffering period information and / or picture timing information for an output layer set, then the video encoder is used to generate the video data stream so that the buffering period information and / or picture timing information for the output layer set appears only in one scalable nested supplementary enhancement information message among the three or more scalable nested supplementary enhancement information messages, and in another scalable nested supplementary enhancement information message among the three or more scalable nested supplementary enhancement information messages that immediately follows the one scalable nested supplementary enhancement information message among the three or more scalable nested supplementary enhancement information messages.
[0290] In addition, according to an embodiment, a video encoder for encoding a video into a video data stream is provided, so that the video data stream has the video encoded therein. The video encoder is used to generate the video data stream so that the video data stream includes a plurality of access units. For each access unit of the plurality of access units, if the access unit includes three or more scalable nested supplementary enhancement information messages, two or more scalable nested supplementary enhancement information messages include buffering period information and / or picture timing information for an output layer set, then the video encoder is used to generate the video data stream so that the video data stream includes an indication indicating whether the buffering period information and / or picture timing information for the output layer set appears only in one scalable nested supplementary enhancement information message among the three or more scalable nested supplementary enhancement information messages and in another scalable nested supplementary enhancement information message among the three or more scalable nested supplementary enhancement information messages that immediately follows the one scalable nested supplementary enhancement information message among the three or more scalable nested supplementary enhancement information messages.
[0291] In an embodiment, all scalable nested supplementary enhancement information messages having the same value of an identifier identifying an output layer set in an access unit may, for example, carry the same buffering period information and / or the same picture timing information.
[0292] According to an embodiment, for a specific access unit among a plurality of access units, any picture timing supplemental enhancement information messages applied to the layer set and sub-layer set of the output layer set may, for example, carry the same picture timing information. And / or, for a specific access unit among a plurality of access units, any buffering period supplemental enhancement information messages applied to the layer set and sub-layer set of the output layer set may, for example, carry the same buffering period information. And / or, for a specific access unit among a plurality of access units, any decoding unit supplemental enhancement information messages applied to the layer set and sub-layer set of the output layer set may, for example, carry the same decoding unit information.
[0293] In an embodiment, two Scalable Nested Supplemental Enhancement Information messages for a specific payload type in an access unit with the same value of an identifier identifying an output layer set carry the same content.
[0294] Furthermore, according to an embodiment, an apparatus is provided for receiving a video data stream as an input bitstream, wherein the video data stream has video encoded therein. The input bitstream includes a plurality of access units. The apparatus is configured to process the access units of the input bitstream to obtain sub-bitstreams. For each access unit of the plurality of access units, if the access unit includes two or more scalable nested supplementary enhancement information messages, the two or more scalable nested supplementary enhancement information messages including buffering period information and / or picture timing information for an output layer set, then the buffering period information and / or picture timing information for the output layer set in all two or more scalable nested supplementary enhancement information messages of the access unit are equal.
[0295] Furthermore, according to an embodiment, an apparatus for receiving a video data stream as an input bitstream is provided, wherein the video data stream has video encoded therein. The input bitstream includes a plurality of access units. The apparatus is configured to process an access unit of the input bitstream to obtain a sub-bitstream. For each access unit of the plurality of access units, if the access unit includes two or more scalable nested supplementary enhancement information messages, the two or more scalable nested supplementary enhancement information messages including buffering period information and / or picture timing information for an output layer set, the video data stream includes an indication of whether the buffering period information and / or picture timing information for the output layer set in all two or more scalable nested supplementary enhancement information messages of the access unit are equal.
[0296] In addition, according to an embodiment, an apparatus for receiving a video data stream as an input bitstream is provided, wherein the video data stream has video encoded therein. The input bitstream includes a plurality of access units, and the apparatus is configured to process the access units of the input bitstream to obtain sub-bitstreams. For each access unit of the plurality of access units, if the access unit includes three or more scalable nested supplementary enhancement information messages, two or more scalable nested supplementary enhancement information messages include buffering period information and / or picture timing information for an output layer set, then the buffering period information and / or picture timing information for the output layer set appears only in one scalable nested supplementary enhancement information message among the three or more scalable nested supplementary enhancement information messages, and in another scalable nested supplementary enhancement information message among the three or more scalable nested supplementary enhancement information messages that immediately follows the one scalable nested supplementary enhancement information message among the three or more scalable nested supplementary enhancement information messages.
[0297] In addition, according to an embodiment, an apparatus for receiving a video data stream as an input bitstream is provided, wherein the video data stream has video encoded therein. The input bitstream includes a plurality of access units. The apparatus is configured to process the access units of the input bitstream to obtain a sub-bitstream. For each access unit of the plurality of access units, if the access unit includes three or more scalable nested supplementary enhancement information messages, two or more scalable nested supplementary enhancement information messages include buffering period information and / or picture timing information for an output layer set, then the video data stream includes an indication of whether the buffering period information and / or picture timing information for the output layer set appears only in one of the three or more scalable nested supplementary enhancement information messages, and in another scalable nested supplementary enhancement information message of the three or more scalable nested supplementary enhancement information messages that immediately follows the one of the three or more scalable nested supplementary enhancement information messages.
[0298] In an embodiment, all scalable nested supplementary enhancement information messages having the same value of an identifier identifying an output layer set in an access unit may, for example, carry the same buffering period information and / or the same picture timing information.
[0299] According to an embodiment, for a specific access unit among a plurality of access units, any picture timing supplemental enhancement information messages applied to the layer set and sub-layer set of the output layer set may, for example, carry the same picture timing information. And / or, for a specific access unit among a plurality of access units, any buffering period supplemental enhancement information messages applied to the layer set and sub-layer set of the output layer set may, for example, carry the same buffering period information. And / or, for a specific access unit among a plurality of access units, any decoding unit supplemental enhancement information messages applied to the layer set and sub-layer set of the output layer set may, for example, carry the same decoding unit information.
[0300] In an embodiment, two Scalable Nested Supplemental Enhancement Information messages for a specific payload type in an access unit with the same value of an identifier identifying an output layer set carry the same content.
[0301] In an embodiment, if the device has found a buffering period scalable nested supplemental enhancement information message for an output layer set, the device may be configured, for example, to use the content of the buffering period scalable nested supplemental enhancement information for the output layer set instead of searching for further buffering period scalable nested supplemental enhancement information messages for the output layer set. And / or, if the device has found a picture timing scalable nested supplemental enhancement information message for an output layer set, the device may be configured, for example, to use the content of the picture timing scalable nested supplemental enhancement information for the output layer set instead of searching for other picture timing scalable nested supplemental enhancement information messages for the output layer set.
[0302] According to an embodiment, the apparatus may, for example, be configured to decode a sub-bitstream to decode a video.
[0303] Furthermore, according to an embodiment, a system for encoding a video into a video data stream and decoding the video is provided. The system includes a video encoder as described above and an apparatus as described above. The video encoder can, for example, be configured to encode the video into a video data stream such that the video data stream has the video encoded therein. The apparatus can, for example, be configured to receive the video data stream as an input bitstream. Furthermore, the apparatus can, for example, be configured to process the input bitstream to obtain a sub-bitstream. Furthermore, the apparatus can, for example, be configured to decode the sub-bitstream to decode the video.
[0304] When using scalable nesting SEI messages to enable replacement of BP and PT SEI messages for all OLs in a bitstream, the current state of the art allows the encoder to spread BP and PT across multiple scalable nesting SEI messages. An extractor processing such a bitstream would likely need to scan all scalable nesting SEI messages in an access unit (which could be a lot, given the number of OLs and repetitions) until it finds the applicable scalable nesting BP and PT SEI messages.
[0305] If the encoder places a first scalable nested BP or PT SEI message into an access unit of a bitstream, and during encoding of a picture of the access unit, the parameters of the BP or PT SEI message are updated by writing another BP or PT SEI message into the access unit of the bitstream, the extractor will then be burdened with ensuring that it uses the most recent or latest BP or PT SEI message from the access unit that was placed into the bitstream when extracting the access unit.
[0306] The present invention simplifies the operation of the extractor by imposing constraints that eliminate the unwise choices mentioned above.
[0307] In an embodiment, for example, a requirement for bitstream consistency may be that all scalable SEI messages in an access unit with the same value of nesting_ols_idx_delta_minus1 (i.e., an identifier / index identifying the OLS) carry the same content (e.g., the same buffering period information and / or the same picture / timing information). Therefore, the BP and PT SEI messages for the OLS can be found in one scalable nesting SEI message, and the extractor can ensure that once it finds the scalable nesting SEI message with the target OLS, it has all the required information.
[0308] In an embodiment, for example, a requirement for bitstream consistency may be that all scalable nested SEI messages in an access unit with the same value of an identifier identifying an OLS (e.g., an index) carry the same content (e.g., the same buffering period information and / or the same picture / timing information). For example, for a particular access unit, the payload of any PT SEI message applied to a set of layers and sublayers (e.g., for a particular OLS) needs to be the same (e.g., for two PT SEI messages included in two separate scalable nested SEI messages). The same applies to BP SEI messages and DUI SEI messages. Therefore, the BP and PT SEI messages for an OLS can be found in one scalable nested SEI message, and the extractor can ensure that once it finds the scalable nested SEI message with the target OLS, it has all the required information.
[0309] According to an embodiment, for example, a requirement for bitstream consistency may be that scalable nested SEI messages of a particular payload type in access units having the same value of an identifier (e.g., index) identifying an OLS carry the same content (e.g., the same buffering period information and / or the same picture / timing information). For example, for a particular access unit in a bitstream carrying two scalable nested PT SEI messages, for example, that apply to the same OLS, the payloads of the two scalable nested PT SEI messages should be equal, and the extractor can ensure that once it finds the scalable nested SEI message with the target OLS, it has all the required information.
[0310] In an embodiment, for example, a requirement for bitstream consistency may be that all scalable nesting SEI messages in an access unit with the same value of an identifier (e.g., index) identifying an OLS carry the same content (e.g., the same buffering period information and / or the same picture / timing information). For example, the BP and PT SEI messages for an OLS may be found in one scalable nesting SEI message, and the extractor may ensure that once it finds the scalable nesting SEI message with the target OLS, it has all the required information.
[0311] Depending on the embodiment, for example, a bitstream conformance requirement may be that the BP and PT SEI messages applicable to OLS are placed back to back within two scalable nested SEI messages without any other scalable nested SEI message NAL units in between that are not applicable to OLS in the bitstream order.
[0312] In the following, temporal scalability for low latency and DU timing is described.
[0313] According to an embodiment, a video data stream having video encoded therein is provided. The video data stream includes a plurality of access units. The video data stream includes a spreading factor that depends on the number of sub-bitstreams of the video data stream or a clock sub-tick value that depends on a highest sub-bitstream of the sub-bitstreams of the video data stream.
[0314] According to an embodiment, each of the sub-bitstreams is dependent on an output layer set and / or dependent on a sub-layer and / or dependent on a sub-picture.
[0315] In an embodiment, the decoding unit removal time is dependent on the access unit removal time and on said spreading factor.
[0316] According to an embodiment, the video data stream may, for example, comprise a temporal distance, wherein the temporal distance may, for example, be configured to be multiplied by a derived clock sub-tick value, wherein the derived clock sub-tick value is derived using the spreading factor.
[0317] In an embodiment, the spreading factor is one of a plurality of spreading factors. Each of the plurality of spreading factors is assigned to a sub-layer of a plurality of sub-layers. The video data stream may, for example, include the plurality of spreading factors.
[0318] According to an embodiment, the clock sub-tick value depends on a clock tick and further depends on the diffusion factor.
[0319] In an embodiment, the clock sub-tick value is defined according to the following steps:
[0320] ClockSubTick=
[0321] -ClockTick+(tick_divisor_minus2+2)*(tick_divisor_factor_minus1[HTid]+1)
[0322] wherein ClockSubTick is the clock subtick, wherein ClockTick is the clock tick, wherein tick_divisor_minus2 is an additional tick divisor, and wherein tick_divisor_factor_minus1[HTid] indicates the diffusion factor of the video data stream.
[0323] Furthermore, according to an embodiment, a video encoder is provided for encoding a video into a video data stream, such that the video data stream has the video encoded therein. The video encoder is configured to generate the video data stream such that the video data stream includes a plurality of access units. Furthermore, the video encoder is configured to generate the video data stream such that the video data stream includes a spreading factor that depends on the number of sub-bitstreams of the video data stream, or such that the video data stream includes a clock sub-tick value that depends on a highest sub-bitstream among the sub-bitstreams of the video data stream.
[0324] According to an embodiment, each of the sub-bitstreams is dependent on an output layer set and / or dependent on a sub-layer and / or dependent on a sub-picture.
[0325] In an embodiment, the video encoder may, for example, be configured to generate the video data stream such that a decoding unit removal time is dependent on an access unit removal time and on the spreading factor.
[0326] According to an embodiment, the video encoder may, for example, be configured to generate the video data stream such that the video data stream may, for example, include a temporal distance, wherein the temporal distance may, for example, be configured to be multiplied by a derived clock sub-beat value, wherein the derived clock sub-beat value may be derived using the diffusion factor.
[0327] In an embodiment, each of the plurality of spreading factors is assigned to a sub-layer of the plurality of sub-layers.The video encoder may, for example, be configured to generate the video data stream such that the video data stream may, for example, comprise the plurality of spreading factors.
[0328] According to an embodiment, the clock sub-tick value depends on the clock tick and further depends on the diffusion factor.
[0329] In an embodiment, the clock sub-tick value is defined according to the following steps:
[0330] ClockSubTick=
[0331] -ClockTick+(tick_divisor_minus2+2)*(tick_divisor_factor_minus1[HTid]+1)
[0332] wherein ClockSubTick is the clock subtick value, wherein ClockTick is the clock tick, wherein tick_divisor_minus2 is an additional tick divisor, and wherein tick_divisor_factor_minus1[HTid] indicates the diffusion factor of the video data stream.
[0333] Furthermore, according to an embodiment, a video encoder for encoding a video into a video data stream is provided, such that the video data stream has the video encoded therein. The video encoder is configured to generate the video data stream such that the video data stream includes sub-layer specific frame rate information for a sub-layer; and / or the video encoder is configured to generate the video data stream such that the video data stream includes sub-layer specific frame display duration information for the sub-layer.
[0334] Furthermore, according to an embodiment, a video decoder is provided for receiving a video data stream as an input bitstream, wherein the video data stream has video encoded therein. The video data stream includes a plurality of access units. The video decoder is configured to decode the video data stream to decode the video. The video data stream includes a spreading factor that depends on a number of sub-bitstreams of the video data stream, wherein the video decoder is configured to use the spreading factor to decode the video; or the video data stream includes a clock sub-tick value that depends on a highest sub-bitstream among the sub-bitstreams of the video data stream, wherein the video decoder is configured to use the clock sub-tick value to decode the video.
[0335] According to an embodiment, each of the sub-bitstreams is dependent on an output layer set and / or dependent on a sub-layer and / or dependent on a sub-picture.
[0336] In an embodiment, the decoding unit removal time is dependent on the access unit removal time and on said spreading factor.
[0337] According to an embodiment, the video data stream may, for example, include a temporal distance. The video decoder may be configured to multiply the temporal distance by a derived clock sub-tick value. The video decoder may, for example, be configured to derive the derived clock sub-tick value using the spreading factor.
[0338] In an embodiment, the spreading factor is one of a plurality of spreading factors. Each of the plurality of spreading factors may, for example, be assigned to a sublayer of a plurality of sublayers. The video data stream may, for example, include the plurality of spreading factors. The video decoder may, for example, be configured to decode the video using the plurality of spreading factors.
[0339] According to an embodiment, the clock sub-tick value depends on the clock tick and further depends on the diffusion factor.
[0340] In an embodiment, the clock sub-tick value is defined according to the following steps:
[0341] ClockSubTick=
[0342] -ClockTick+(tick_divisor_minus2+2)*(tick_divisor_factor_minus1[HTid]+1)
[0343] wherein ClockSubTick is the clock subtick value, wherein ClockTick is the clock tick, wherein tick_divisor_minus2 is an additional tick divisor, and wherein tick_divisor_factor_minus1[HTid] indicates the diffusion factor of the video data stream.
[0344] Furthermore, according to an embodiment, a video decoder is provided for receiving a video data stream as an input bitstream, wherein the video data stream has video encoded therein. The video decoder is configured to decode the video data stream to decode the video. The video data stream includes sub-layer-specific frame rate information for a sub-layer, and / or wherein the video data stream includes sub-layer-specific frame display duration information for the sub-layer. The decoder is configured to determine a spreading factor using the sub-layer-specific frame rate information for the sub-layer and / or using the sub-layer-specific frame display duration information.
[0345] Furthermore, according to an embodiment, a system for encoding a video into a video data stream and decoding the video is provided. The system includes a video encoder as described above and a video decoder as described above. The video encoder is configured to encode the video into the video data stream so that the video data stream has the video encoded therein. The video decoder is configured to receive the video data stream and decode the video data stream to decode the video.
[0346] Removal of a temporal sub-layer changes the removal times of the remaining access units of lower temporal sub-layers. This is particularly true when there is frame reordering. However, in low-latency configurations, and more particularly when DU timing is provided, the removal times change in a very structured manner, i.e., they are modified to extend or spread the decoding time over time. Figure 4 An explanation is provided of the problem when there are 3 DUs per access unit.
[0347] Note that since the final decoding time of an AU is different when decoding the entire bitstream than when decoding only its substream, the ultra-low latency property is lost. Typically, this is a result of different decoding capabilities, i.e., a decoder capable of decoding 60fps can process a frame in 1 / 60 second, while a decoder capable of decoding 30fps can only process a frame in 1 / 30 second.
[0348] Note that the DU removal time is indicated as a delta to the AU removal time. The deltaTime is signaled in the picture timing SEI message or the decoding unit information SEI message in the number of ClockSubTicks. This deltaTime is indicating the removal time of the DU compared to the AU removal time. Instead of using a loop over multiple temporal sub-layers to indicate the DU removal time from the CPB (as a deltaTime compared to the AU removal time), a spreading factor can be derived. In one embodiment, the spreading factor is derived from sub-layer specific frame rate information (e.g. sub-layer T0 30fps, sub-layer L1 60fps) or sub-layer specific frame display duration (FrameTimeInterval explained below in Section 6, e.g. sub-layer L01 / 30s and sub-layer L11 / 60s), for example depending on the ratio of such information for two such sub-layers.
[0349] Alternatively, such a spreading factor may be indicated. This embodiment is used to add information indicating the DU removal time of the spreading factor to calculate the corresponding DU time as an increment of the AU removal time.
[0350] Or as indicated below:
[0351]
[0352] Or, as an alternative embodiment, as follows:
[0353]
[0354] The export of the associated variable ClockSubTick in the HRD specification will change as follows:
[0355] The variable ClockSubTick is derived as follows and is called the clock subtick:
[0356] ClockSubTick-ClockTick+(tick_divisor_minus2+2)*(tick_divisor_factor_minus1[HTid]+1)(C-2)
[0357] The deltaTime is then using ClockSubTick which is larger or smaller depending on the highest sublayer present in the bitstream.
[0358] Alternatively, the syntax element tick_divisor_minus2 can be replicated per sublayer indicating the correct beat divisor when the highest temporal sublayer (HTid) is set equal to the corresponding sublayer. In this case, there will be no spreading factor, but multiple ClockSubTicks will be signaled, each with a different value for the highest sublayer present in the bitstream, so that when the sublayer is removed from the original bitstream, the ClockSubTick used is a different ClockSubTick.
[0359] Hereinafter, correlation extracted for OLS when using common DU timing is explained.
[0360] In the case where the bitstream contains extractable sub-bitstreams specific to OLS and common DU timing (containing non-scalable nested PT SEI messages for common DU timing), it is desirable from the perspective of bitrate overhead and processing complexity to reuse those PT SEI messages in a manner comparable to the above indications (use_orig_pic_timing_flag and same_pic_timing_within_ols_flag) rather than providing scalable nested PT SEI messages as an alternative.
[0361] Figure 5 is an illustration where two layers are depicted in the top and removal times of AUs and DUs of the two layers are shown in the bottom for a complete bitstream (eg 0th OLS) and a single layer extracted with layer ID L0 (eg 1st OLS).
[0362] In this case, the common timing needs to be extended in the same way as above by using tick_divisor_factor_minus1[] or an absolute indication of the adjusted tick divisor (see the spread of DU times in the figure). Therefore, in one embodiment, each OLS-specific HRD parameter syntax structure in the VPS carries the absolute value of the relative factor or tick divisor to extend the common DU removal timing.
[0363] However, in this usage scenario, the AU contains both L0 and L1 pictures, and therefore the number of DUs per AU changes with extraction. This means that it also requires the aspects described later in aspect 6 to derive the correct number of remaining DUs after extraction.
[0364] In the following, the correlation for sub-picture extraction when using a common DU timing is explained.
[0365] In the case where the bitstream contains extractable sub-pictures (for sub-picture boundary handling with motion compensated prediction enabled) and common DU timing (non-scalable nested PT SEI messages containing common timing), it is also desirable from the perspective of bitrate overhead and processing complexity to reuse those SEI messages in a manner comparable to the above indications (use_orig_pic_timing_flag and same_pic_timing_within_ols_flag) instead of providing scalable nested PT SEI messages as an alternative.
[0366] In this case, the common timing needs to be extended in the same way as above by tick_divisor_factor_minus1[] or an absolute indication of the adjusted tick divisor in order to derive the correct number of remaining DUs after extraction. Therefore, in one embodiment, the sub-picture specific HRD parameter syntax structure in the VPS / SPS carries the absolute value of the relative factor or tick divisor in order to extend the common DU removal timing.
[0367] Alternatively, a SEI message (eg, a sub-picture level information SEI message) carries extended information or absolute values so that the HRD parameters in the extracted sub-bitstream can be properly derived.
[0368] In the following, CPB / rate size derivation for sub-pictures is described.
[0369] According to an embodiment, a video decoder is provided for receiving a video data stream as an input bitstream, wherein the video data stream has video encoded therein. The video decoder is configured to decode the video data stream to decode the video. To decode the video, the video decoder estimates a decoded picture buffer size for a sub-picture based on information within the video data stream indicating a picture buffer size for a current decoding region.
[0370] According to an embodiment, the video data stream may, for example, comprise a signaled coded picture buffer size for a reference level.
[0371] In an embodiment, if the current level is equal to the reference level, the video decoder may, for example, be configured to determine the decoded picture buffer size using the decoded picture buffer size for the reference level.
[0372] According to an embodiment, the video decoder may, for example, be configured to estimate the coded picture buffer size relying on the syntax element cpb_size_value_minus1[i][j] of the video data stream.
[0373] In an embodiment, the video decoder may, for example, be configured to rely on the syntax element cpb_size_scale of the video data stream to estimate the coded picture buffer size.
[0374] According to an embodiment, the video decoder may be configured, for example, to estimate a coded picture buffer size for a sub-picture as a video coding layer coded picture buffer size, and may be configured, for example, to estimate another coded picture buffer size for a sub-picture as a network abstraction layer coded picture buffer size.
[0375] In an embodiment, the video decoder may be configured to estimate the video coding layer coded picture buffer size and / or the network abstraction layer coded picture buffer size, for example, depending on a reference level fraction value.
[0376] According to an embodiment, the video decoder may be configured to estimate the video coding layer coded picture buffer size according to:
[0377] SubPicCbpSizeVcl[s]=
[0378] =Floor((cpb_size_value_minus1[i][j]+1)
[0379] *2 (4+cpb_size_scale) *RefLevelFraction[i][j]+256)
[0380] The device may be configured, for example, to estimate the network abstraction layer decoding picture buffer size according to the following:
[0381] SubPicCbpSizeNal[s]=
[0382] =Floor((cpb_size_value_minus1[i][j]+1)
[0383] *2 (4+cpb_size_scale) *RefLevelFraction[i][j]+256),
[0384] Where RefLevelFraction is the reference level fraction value.
[0385] In an embodiment, the video decoder may be configured to estimate the video coding layer coded picture buffer size according to:
[0386] SubpicCpbSizeVcl[i][j][k]=
[0387] Floor(CpbVclFactor*MaxCPB*OlsRefLevelFraction[i][j][k]+256)
[0388] SubpicCpbSizeNal[i][j][k]=
[0389] Floor(CpbNalFactor*MaxCPB*OlsRefLevelFraction[i][j][k]+256)
[0390] where i, j and k are indices and OlsRefLevelFraction[i][j][k] are real numbers.
[0391] Furthermore, according to an embodiment, a video decoder is provided for receiving a video data stream as an input bitstream, wherein the video data stream has video encoded therein. The video decoder is configured to decode the video data stream to decode the video. To decode the video, the video decoder is configured to estimate a bitrate for a sub-picture based on information within the video data stream indicating bitrate information of a currently decoded video sequence.
[0392] According to an embodiment, the video data stream may include, for example, an indication indicating whether the bit rate for the sub-picture should be estimated using the bit rate information of the currently decoded video sequence. If the indication of the video data stream indicates that the bit rate for the sub-picture should be estimated using the bit rate information of the currently decoded video sequence, the video decoder estimates the bit rate using the bit rate information of the currently decoded video sequence. If the indication of the video data stream indicates that the bit rate for the sub-picture should be estimated without using the bit rate information of the currently decoded video sequence, the video decoder estimates the bit rate using a predetermined value or a worst-case value without using the bit rate information of the currently decoded video sequence.
[0393] In an embodiment, the bit rate information of the currently decoded video sequence is the signaled bit rate for the reference level. The video data stream may, for example, include the signaled bit rate for the reference level. If the current level is equal to the reference level, the video decoder may, for example, be configured to use the signaled bit rate for the reference level to determine the bit rate for the sub-picture.
[0394] According to an embodiment, a video decoder may, for example, be configured to estimate the bit rate for a sub-picture by relying on the syntax element bit_rate_value_minus1[i][j] of the video data stream.
[0395] In an embodiment, the video decoder may be configured to estimate the bit rate for a sub-picture, for example, relying on a syntax element bit_rate_scale of the video data stream.
[0396] According to an embodiment, the video decoder may be configured, for example, to estimate the bit rate for the sub-picture as the video decoding layer bit rate for the sub-picture, and may be configured, for example, to estimate another decoding picture buffer size for the sub-picture as the network abstraction layer bit rate for the sub-picture.
[0397] In an embodiment, the video decoder may, for example, be configured to estimate the video coding layer coded picture buffer size and / or the network abstraction layer coded picture buffer size depending on the reference level score value.
[0398] According to an embodiment, the video decoder may be configured to estimate the video coding layer bit rate for a sub-picture according to:
[0399]
[0400] The apparatus may be configured, for example, to estimate the network abstraction layer bit rate for the sub-picture according to:
[0401]
[0402] Where RefLevelFraction is the reference level fraction value.
[0403] In an embodiment, the video decoder may be configured to estimate the video coding layer rate for a sub-picture according to:
[0404] SubpicBitRateVcl[i][j][k]=
[0405] Floor(CpbVclFactor*ValBR*OlsRefLevelFraction[0][j][k]+256)
[0406] SubpicBitRateNal[i][j][k]=
[0407] Floor(CpbNalFactor*ValBR*OlsRefLevelFraction[0][j][k]+256)
[0408] where i, j and k are indices and OlsRefLevelFraction[0][j][k] are real numbers.
[0409] In an embodiment, i may, for example, indicate an index of a particular indicated reference level, j may, for example, indicate an index of a particular sub-picture of a picture of an access unit in a video data stream, and k may, for example, indicate an index of a maximum temporal sub-layer that the video data stream includes and / or on which the video decoder operates.
[0410] According to an embodiment, OlsRefLevelFraction[i][j][k] may, for example, depend on a variable sli_non_subpic_layers_fraction[i][k] indicating the i-th fraction of the bitstream level limit associated with layers in targetCvss with sps_num_subpics_minus1 equal to 0 when Htid equals k.
[0411] In an embodiment,
[0412] If vps_max_layers_minus1 is equal to 0, or when there is no layer in the bitstream and sps_num_subpics_minus1 is equal to 0, then for example sli_non_subpic_layers_fraction[i][k]=0, and
[0413] If k is less than sli_max_sublayers_minus1 and sli_non_subpic_layers_fraction[i][k] does not exist, then, for example, sli_non_subpic_layers_fraction[i][k]=sli_non_subpic_layers_fraction[i][k+1].
[0414] According to an embodiment,
[0415] OlsRefLevelFraction[i][j][k]=
[0416] =sli_non_subpic_layers_fraction[i][k]+(n-sli_non_subpic_layers_fraction[i][k])+n.
[0417] *(sli_ref_level_fraction_minus1[i][j][k]+1).
[0418] Here, n indicates a positive integer.
[0419] According to an embodiment, for example, n=256; or n=128; or n=512; or n=1024; or n=2048; or n=4096.
[0420] According to an embodiment, i, j and k are defined depending on sli_ref_level_fraction_minus1, where sli_ref_level_fraction_minus1[i][j][k] plus 1 specifies the i-th fraction of the level limit associated with sli_ref_level_idc[i][k] for the sub-picture, when Htid as the sub-layer index considered is equal to k, and the sub-picture has a sub-picture index equal to j in the layer in targetCvss with sps_num_subpics_minus1 greater than 0.
[0421] Furthermore, according to an embodiment, a video decoder is provided for receiving a video data stream as an input bitstream, wherein the video data stream has video encoded therein. The video decoder is configured to decode the video data stream to decode the video. To decode the video, the video decoder is configured to receive a decoded picture buffer size for a sub-picture encoded within the video data stream and to decode the video using the decoded picture buffer size for the sub-picture; and / or to decode the video, the video decoder is configured to receive a bitrate for a sub-picture encoded within the video data stream and to decode the video using the bitrate for the sub-picture.
[0422] According to an embodiment, the video data stream may include, for example, an indication indicating whether the decoded picture buffer size for the sub-picture is encoded within the video data stream or whether the decoded picture buffer size for the sub-picture should be estimated. If the indication of the video data stream indicates that the decoded picture buffer size for the sub-picture should be estimated, the video decoder estimates the decoded picture buffer size for the sub-picture. If the indication of the video data stream indicates that the decoded picture buffer size for the sub-picture is encoded within the video data stream, the video decoder uses the decoded picture buffer size for the sub-picture encoded within the video data stream.
[0423] In an embodiment, the video data stream may include, for example, an indication indicating whether the bit rate for the sub-picture is encoded within the video data stream or whether the bit rate for the sub-picture should be estimated. If the indication of the video data stream indicates that the bit rate for the sub-picture should be estimated, the video decoder estimates the bit rate for the sub-picture. If the indication of the video data stream indicates that the internal bit rate for the sub-picture is encoded in the video data stream, the video decoder uses the decoded picture buffer size for the sub-picture encoded within the video data stream.
[0424] According to an embodiment, each of a plurality of extractable sub-bitstreams is specific to an output layer set, wherein a sub-picture is assigned to at least one of the plurality of extractable sub-bitstreams. If the video data stream may, for example, include common decoding unit removal timing information and a plurality of extractable sub-bitstreams, then a hypothetical reference decoder parameter syntax structure specific to each output layer set in a video parameter set or sequence parameter set or in a supplemental enhancement information message of the video data stream may, for example, include a spreading factor or an absolute value of a beat divisor for extending the common decoding unit removal timing. The video decoder may, for example, be configured to process the absolute value of the spreading factor or the beat divisor.
[0425] Furthermore, according to an embodiment, a video data stream having video encoded therein is provided. The video data stream may, for example, include the syntax element cpb_size_value_minus1[i][j] and the syntax element cpb_size_scale. Or the video data stream may, for example, include the syntax element bit_rate_value_minus1[i][j] and the syntax element bit_rate_scale.
[0426] Furthermore, according to an embodiment, a video data stream having video encoded therein is provided, wherein the video data stream includes an indication indicating whether currently decoded picture buffer size information should be used to estimate a decoded picture buffer size of a sub-picture, and / or the video data stream includes an indication indicating whether currently decoded video sequence bit rate information should be used to estimate a bit rate for the sub-picture.
[0427] According to an embodiment, the currently decoded picture buffer size information is the decoded picture buffer size for signaling of a reference level, wherein the video data stream may, for example, include the decoded picture buffer size for signaling of a reference level; and / or the currently decoded video sequence bit rate information is the signaling bit rate for a reference level, wherein the video data stream may, for example, include the signaling bit rate for a reference level.
[0428] Furthermore, according to an embodiment, a video data stream having a video encoded therein is provided, wherein the video data stream includes an indication indicating whether a decoded picture buffer size of a sub-picture is encoded in the video data stream or whether the decoded picture buffer size of the sub-picture should be estimated, and / or the video data stream includes an indication indicating whether a bit rate of the sub-picture is encoded within the video data stream or whether the bit rate of the sub-picture should be estimated.
[0429] According to an embodiment, each of the plurality of extractable sub-bitstreams is specific to an output layer set, wherein a sub-picture is assigned to at least one of the plurality of extractable sub-bitstreams. If the video data stream may, for example, include common decoding unit removal timing information and a plurality of extractable sub-bitstreams, then a hypothetical reference decoder parameter syntax structure specific to each output layer set in a video parameter set or a sequence parameter set or in a supplemental enhancement information message of the video data stream may, for example, include a spreading factor or an absolute value of a beat divisor for extending the common decoding unit removal timing.
[0430] Furthermore, according to an embodiment, a video encoder for encoding a video into a video data stream is provided, such that the video data stream has the video encoded therein. The video encoder is configured to generate the video data stream such that the video data stream may, for example, include the syntax element cpb_size_value_minus1[i][j] and the syntax element cpb_size_scale. Alternatively, the video encoder may generate the video data stream such that the video data stream may, for example, include the syntax element bit_rate_value_minus1[i][j] and the syntax element bit_rate_scale.
[0431] Furthermore, according to an embodiment, a video encoder for encoding a video into a video data stream is provided, such that the video data stream has the video encoded therein. The video encoder generates the video data stream such that the video data stream includes an indication indicating whether the decoded picture buffer size information of the sub-picture should be used to estimate the decoded picture buffer size of the sub-picture, and / or the video encoder generates the video data stream such that the video data stream includes an indication indicating whether the bit rate information of the currently decoded video sequence should be used to estimate the bit rate of the sub-picture.
[0432] According to an embodiment, the video encoder may, for example, be configured to generate a video data stream such that the currently decoded picture buffer size information is the signaled decoded picture buffer size for the reference level, wherein the video data stream may, for example, include the signaled decoded picture buffer size for the reference level; and / or the video encoder may, for example, be configured to generate a video data stream such that the currently decoded video sequence bitrate information is the signaled bitrate for the reference level, wherein the video data stream may, for example, include the signaled bitrate for the reference level.
[0433] Furthermore, according to an embodiment, a video encoder for encoding a video into a video data stream is provided, such that the video data stream has the video encoded therein. The video encoder is configured to generate the video data stream such that the video data stream includes an indication indicating whether a decoded picture buffer size for a sub-picture is encoded within the video data stream or whether the decoded picture buffer size for the sub-picture should be estimated, and / or the video encoder is configured to generate the video data stream such that the video data stream includes an indication indicating whether a bit rate for the sub-picture is encoded within the video data stream or whether the bit rate for the sub-picture should be estimated.
[0434] Furthermore, according to an embodiment, a video encoder for encoding video into a video data stream is provided, such that the video data stream has the video encoded therein. The video encoder is configured to generate the video data stream, wherein, if the video data stream includes common decoding unit removal timing information and a plurality of extractable sub-bitstreams, each of the plurality of sub-bitstreams being specific to an output layer set, a hypothetical reference decoder parameter syntax structure specific to each output layer set in a video parameter set or a sequence parameter set or a supplemental enhancement information message of the video data stream includes an absolute value of a beat divisor or a spreading factor for extending the common decoding unit removal timing.
[0435] Furthermore, according to an embodiment, a system for encoding a video into a video data stream and for decoding the video is provided. The system includes a video encoder as described above and a video decoder as described above. The video encoder can, for example, be configured to encode a video into a video data stream so that the video data stream has the video encoded therein. The video decoder can, for example, be configured to receive a video data stream and decode the video data stream to decode the video.
[0436] The current VVC draft specification contains SEI messages that indicate level information for sub-pictures and additional information to help estimate the level of a set of sub-pictures. The way this is achieved is by signaling the fraction that a sub-picture contributes to a given signaled reference level. Based on the level score of each sub-picture, the variables CPBsize and bitrate and subsequently the level of the sub-picture bitstream are approximated. Furthermore, the accumulated level scores are used to derive the CPB size and bitrate of a bitstream consisting of a set of sub-pictures and also to approximate their levels. However, in all these derivations the MaxCPBSize and / or MaxBitrate of the reference level are used, which is a problem because most bitstreams may not fully occupy the CPB and bitrate budget up to the MaxCPBSize and / or MaxBitrate for a given level. When using these values to approximate the level of a merged bitstream of a sub-picture bitstream or a set of sub-picture bitstreams, the (cumulative) CPB size and bitrate are most likely to be over-provisioned.
[0437] The VVC draft specification includes the derivation of the variables SubPicCpbSizeVcl[i][j] and SubPicCpbSizeNal[i][j] as follows:
[0438] SubPicCpbSizeVcl[i][j]=Floor(CpbVclFactor*MaxCPB*RefLevelFraction[i][j]+256)
[0439] SubPicCpbSizeNal[i][j]=Floor(CpbNalFactor*MaxCPB*RefLevelFraction[i][j]+256)
[0440] Because the original bitstream containing all sub-pictures may already carry more accurate information about the exact CPB size and bitrate (cpb_size_value_minus1 and cpb_size_scale) of the bitstream, rather than just referring to the corresponding maximum values derived from the bitstream level. Therefore, the purpose of the present invention is to derive and use more accurate values of the corresponding values of the CBP size and bitrate for each sub-picture in the approximation at the sub-picture set level. In another embodiment, the sub-picture CPB is derived as:
[0441]
[0442] Note that the syntax elements cpb_size_value_minus1[i][j] are submitted separately for the Vcl and Nal HRD parameters and are used separately in the derivation above. Therefore, the values of SubPicCbpSizeVcl[s] and SubPicCbpSizeVcl[s] may be derived to different values.
[0443] Since the CPB size is signaled for a given level, the above derivation can only be performed if that level is included as a reference level.
[0444] For bitrate, change the corresponding derivation from using the maximum bitrate as reference (Br[Vcl / Nal]Factor*MaxBR) to the following
[0445] SubPicBitRateVcl[s]=Floor(BrVclFactor*MaxBR*RefLevelFraction[i][j]+256)
[0446] SubPicBitRatcNal[s]=Floor(BrNalFactor*MaxBR*RcfLevelFraction[i][j]+256)
[0447] The actual bit rate of the used bit stream is signaled as ((bit_rate_value_minus1[i][j]+1)*2 (6+bit_rate_scale ),as follows
[0448]
[0449] Note that the same applies here as above for the CPB size, and the value of the syntax element bit_rate_value_minus1[i][j] depends on whether Nal or Vcl HRD is considered, and thus the values of SubBitRateVcl[s] and SubPicBitrateVcl[s] may be derived to different values.
[0450] In an embodiment, the variables SubpicCpbSizeVcl[i][j][k] and SubpicCpbSizeNal[i][j][k] are derived as follows:
[0451] SubpicCpbSizeVcl[i][j][k]=Floor(CpbVclFactor*MaxCPB*OlsRcfLevelFraction[i][j][k]+256)
[0452] SubpicCpbSizeVcl[i][j][k]=Floor(CpbVclFactor*MaxCPB*OlsRcfLevelFraction[i][j][k]+256)
[0453] where MaxCPB is derived from sli_ref_level_idc[i][k]
[0454] In an embodiment, the variables SubpicBitRateVcl[i][j][k] and SubpicBitRateNal[i][j][k] are derived as follows:
[0455] SubpicBitRateVcl[i][j][k]=Floor(CpbVclFactor*ValBR*OlsRefLevelFraction[i][j][k]+256)
[0456] SubpicBitRateVcl[i][j][k]=Floor(CpbNalFactor*ValBR*OlsRefLevelFraction[0][j][k]+256)
[0457] For example, the variable OlsRefLevelFraction[i][j][k] is a number, such as a real number.
[0458] For example, the variable OlsRefLevelFraction[i][j][k] may depend on
[0459] sli_non_subpic_layers_fraction[i][k]+(n-sli_non_subpic_layers_fraction[i][k])+n
[0460] *(sli_ref_level_fraction_minus1[i][j][k]+1).
[0461] wherein n indicates a positive integer, for example, n=256; or, for example, n=128; or, for example, n=512; or, for example, n=1024; or, for example, n=2048; or, for example, n=4096;
[0462] So, for example:
[0463] OlsRefLevelFraction[i][j][k]=
[0464] =sli_non_subpic_layers_fraction[i][k]+(256-sli_non_subpic_layers_fraction[i][k])
[0465] +256*(sli_ref_level_fraction_minus1[i][j][k]+1).
[0466] For example, sli_non_subpic_layers_fraction[i][k] may indicate the i-th fraction of the bitstream level limit associated with layers in targetCvss with sps_num_subpics_minus1 equal to 0 when Htid is equal to k. When vps_max_layers_minus1 is equal to 0 or when there are no layers in the bitstream with sps_num_subpics_minus1 equal to 0, then sli_non_subpic_layers_fraction[i][k] shall be equal to 0. When k is less than sli_max_sublayers_minus1 and sli_non_subpic_layers_fraction[i][k] is not present, it is inferred to be equal to sli_non_subpic_layers_fraction[i][k+1], and sli_ref_level_fraction_minus1[i][j][k] plus 1 specifies the i-th fraction of the level limit associated with sli_ref_level_idc[i][k] for the sub-picture with sub-picture index equal to j in the layer in targetCvss for which sps_num_subpics_minus1 is greater than 0 when Htid is equal to k. When k is less than sli_max_sublayers_minus1 and sli_ref_level_fraction_minus1[i][j][k] is not present, it is inferred to be equal to sli_ref_level_fraction_minus1[i][j][k+1].
[0467] Alternatively, in another embodiment, the CPB size and / or code rate of each sub-picture can be directly signaled instead of being derived. Or further, there can be a gating flag indicating whether the value can be derived or explicitly signaled.
[0468] Hereinafter, DU timing signaling in picture timing SEI is described.
[0469] According to an embodiment, a video data stream having video encoded therein is provided. The video data stream includes a plurality of access units. The video data stream further includes delta time information for each of two or more decoding units of an access unit in the plurality of access units, wherein a decoding unit removal time for each of the two or more decoding units of the access unit is dependent on the access unit removal time for the access unit and on the delta time information for the decoding unit.
[0470] In an embodiment, the video data stream may, for example, include picture timing supplementary enhancement information.The picture timing supplementary enhancement information may, for example, include delta time information for two or more decoding units of the access unit.
[0471] In an embodiment, the delta time information indicates a removal time difference between two decoding units of the two or more decoding units of the access unit.
[0472] According to an embodiment, a last decoding unit of the two or more decoding units of the access unit has a removal time equal to the removal time of the access unit.
[0473] In an embodiment, the access unit may, for example, include three or more decoding units.For each pair of two consecutive decoding units in the three or more decoding units of the access unit, the removal time differences are equal.
[0474] According to an embodiment, the picture timing supplemental enhancement information is applied to a sub-bitstream derived from a video data stream, wherein the number of decoding units remains constant.
[0475] In an embodiment, the frame time interval is signaled in a parameter set of the video data stream, in an HRD parameter in a sequence parameter set.
[0476] According to an embodiment, the frame time interval may be derived as the difference in removal times of two consecutive access units at the highest temporal level.
[0477] In an embodiment, each of the two or more decoding units in the access unit may, for example, comprise a video coding layer network abstraction layer unit.
[0478] According to an embodiment, picture timing supplemental enhancement information is applied to sub-bitstreams derived from a video data stream, where there are different numbers of decoding units.
[0479] In an embodiment, the frame time interval may be derived as (elemental_duration_in_tc_minus1[maxTiD]+1) multiplied by ClockTicks.
[0480] According to an embodiment, the video data stream may, for example, include an indication indicating whether the number of decoding units is variable for the video data stream.
[0481] In an embodiment, the video data stream does not comprise an indication indicating whether the video data stream is variable in the instance of the decoding unit.
[0482] According to an embodiment, the number of decoding units within an access unit depends on the frame time interval and on a common delay increase.
[0483] In an embodiment, the video data stream may include, for example, a decoding unit information supplementary enhancement information message for a decoding unit of the two or more decoding units of the access unit. The decoding unit information supplementary enhancement information message for the decoding unit may, for example, include delta time information for the decoding unit.
[0484] According to an embodiment, the video data stream may, for example, include a minimum duration flag of a video parameter set or a sequence parameter set of the video data stream, wherein the minimum picture duration flag indicates whether frame time interval information exists when there is no constant frame rate.
[0485] Furthermore, according to an embodiment, a video encoder for encoding a video into a video data stream is provided, such that the video data stream has the video encoded therein. The video encoder is configured to generate the video data stream such that the video data stream includes a plurality of access units. Furthermore, the video encoder is configured to generate the video data stream such that the video data stream includes delta time information for each of two or more decoding units of the access unit in the plurality of access units, wherein a decoding unit removal time for each of the two or more decoding units of the access unit is dependent on the access unit removal time for the access unit and on the delta time information for the decoding unit.
[0486] According to an embodiment, the video encoder may be configured to generate a video data stream, such that the video data stream may include picture timing supplementary enhancement information. The video encoder may be configured to generate a video data stream, such that the picture timing supplementary enhancement information may include delta time information for two or more decoding units of the access unit.
[0487] In an embodiment, the video encoder may, for example, be configured to generate the video data stream such that the delta time information indicates a removal time difference between two decoding units of the two or more decoding units of the access unit.
[0488] According to an embodiment, a last decoding unit of the two or more decoding units of the access unit has a removal time equal to the removal time of the access unit.
[0489] In an embodiment, the video encoder may be configured to generate a video data stream such that the access unit may include three or more decoding units. For each pair of two consecutive decoding units in the three or more decoding units of the access unit, the removal time difference is equal.
[0490] According to an embodiment, picture timing supplemental enhancement information is applied to a sub-bitstream derived from a video data stream, wherein the number of decoding units remains constant.
[0491] In an embodiment, the video encoder may, for example, be configured to generate the video data stream such that the frame time interval is signaled in a parameter set of the video data stream, in an HRD parameter in a sequence parameter set.
[0492] According to an embodiment, the video encoder may, for example, be configured to generate the video data stream such that the frame time interval is derivable as the difference in removal times of two consecutive access units at the highest temporal level.
[0493] In an embodiment, the video encoder may, for example, be configured to generate the video data stream such that each of the two or more decoding units of the access unit may, for example, comprise one video coding layer network abstraction layer unit.
[0494] According to an embodiment, the video encoder may, for example, be configured to generate a video data stream such that picture timing supplemental enhancement information is applied to sub-bitstreams derived from the video data stream, wherein there are different numbers of decoding units.
[0495] In an embodiment, the frame time interval may be derived as (elemental_duration_in_tc_minus1[maxTiD]+1) multiplied by ClockTicks.
[0496] According to an embodiment, the video encoder may, for example, be configured to generate the video data stream such that the video data stream may, for example, include an indication indicating whether the video data stream is variable in the number of decoding units.
[0497] In an embodiment, the video encoder may, for example, be configured to generate the video data stream such that the video data stream does not include an indication indicating whether the video data stream is variable in the number of decoding units.
[0498] According to an embodiment, the video encoder may, for example, be configured to generate the video data stream such that the number of decoding units within an access unit depends on the frame time interval and on a common delay increment.
[0499] In an embodiment, the video encoder may be configured to generate a video data stream, for example, such that the video data stream may include a decoding unit information supplemental enhancement information message for a decoding unit of two or more decoding units of the access unit. The video encoder may generate the video data stream, for example, such that the decoding unit information supplemental enhancement information message for the decoding unit may include delta time information for the decoding unit.
[0500] According to an embodiment, the video data stream may, for example, include a minimum duration flag of a video parameter set or a sequence parameter set of the video data stream, wherein the minimum picture duration flag indicates whether frame time interval information exists when there is no constant frame rate.
[0501] Furthermore, according to an embodiment, a video decoder for receiving a video data stream as an input bitstream is provided, wherein the video data stream has video encoded therein. The video data stream includes a plurality of access units. The video decoder is configured to decode the video data stream to decode the video. Furthermore, the video data stream includes delta time information for each of two or more decoding units of an access unit in the plurality of access units, wherein a decoding unit removal time for each of the two or more decoding units of the access unit depends on the access unit removal time for the access unit and on the delta time information for the decoding unit, wherein the video decoder is configured to decode the video data stream using the delta time information for each of the two or more decoding units of the access unit.
[0502] According to an embodiment, the video data stream may include, for example, picture timing supplementary enhancement information. The picture timing supplementary enhancement information may include, for example, delta time information for two or more decoding units of the access unit.
[0503] In an embodiment, the delta time information indicates a removal time difference between two decoding units of the two or more decoding units of the access unit, wherein the video decoder may, for example, be configured to use the removal time difference between the two decoding units to decode the video data stream.
[0504] According to an embodiment, a last decoding unit of the two or more decoding units of the access unit has a removal time equal to a removal time of the access unit, wherein the video decoder may eg be configured to use the removal time of the access unit to decode the video data stream.
[0505] In an embodiment, the access unit may, for example, include three or more decoding units.For each pair of two consecutive decoding units in the three or more decoding units of the access unit, the removal time differences are equal.
[0506] According to an embodiment, picture timing supplemental enhancement information is applied to a sub-bitstream derived from a video data stream, wherein the number of decoding units remains constant.
[0507] In an embodiment, the frame time interval is signaled in a parameter set of the video data stream, in an HRD parameter in a sequence parameter set, wherein a video decoder may eg be configured to use the frame time interval for decoding the video data stream.
[0508] According to an embodiment, the frame time interval may be derived as the difference in removal times of two consecutive access units at the highest temporal level, wherein a video decoder may eg be configured to use the frame time interval for decoding a video data stream.
[0509] In an embodiment, each of the two or more decoding units in the access unit may, for example, comprise a video coding layer network abstraction layer unit.
[0510] According to an embodiment, picture timing supplemental enhancement information is applied to sub-bitstreams derived from a video data stream, where there are different numbers of decoding units.
[0511] In an embodiment, the video decoder may be configured to derive the frame time interval according to:
[0512] (elemental_duration_in_tc_minus1[maxTiD]+1) multiplied by ClockTicks.
[0513] According to an embodiment, the video data stream may, for example, comprise an indication indicating whether the number of decoding units is variable for the video data stream, wherein the video decoder may, for example, be configured to decode the video data stream by processing the indication.
[0514] In an embodiment, the video data stream does not comprise an indication indicating whether the video data stream is variable in the instance of the decoding unit.
[0515] According to an embodiment, the number of decoding units within an access unit depends on the frame time interval and on a common delay increase.
[0516] In an embodiment, the video data stream may include, for example, a decoding unit information supplemental enhancement information message for a decoding unit of two or more decoding units of the access unit. The decoding unit information supplemental enhancement information message for a decoding unit may, for example, include delta time information for the decoding unit, wherein the video decoder may, for example, be configured to use the delta time information for the decoding unit to decode the video data stream.
[0517] According to an embodiment, the video data stream may, for example, include a minimum picture duration flag of a video parameter set or a sequence parameter set of the video data stream, wherein the minimum picture duration flag indicates whether frame time interval information exists when a constant frame rate does not exist.
[0518] Furthermore, according to an embodiment, a system for encoding a video into a video data stream and for decoding the video is provided. The system includes a video encoder as described above and a video decoder as described above. The video encoder can, for example, be configured to encode a video into a video data stream so that the video data stream has the video encoded therein. The video decoder can, for example, be configured to receive a video data stream and decode the video data stream to decode the video.
[0519] As mentioned above, DU timing is given as an increment of AU timing. More specifically, the DU removal time is indicated by giving a delta time relative to the removal time of the AU containing a specific DU in a picture timing SEI message or a decoding unit information SEI message.
[0520] If this information is included in the picture timing SEI message, the information signaled is the removal time difference between the two DUs. The last DU in an AU has a removal time equal to the AU removal time, and any other DUs are signaled as the difference in removal time to the next DU. There are two options to represent this (highlighted in different colors):
[0521]
[0522] The first way to signal it is for the case where the DUs have the same removal time difference that is common to all DUs. The second case is where the removal time difference between DUs within an AU is not the same.
[0523] The current syntax in the picture timing SEI message prevents the applicability of aspect 2 discussed within this disclosure when DU timing is present, because the number of DUs may change when extracting sub-bitstreams from the bitstream.
[0524] In principle, when the common removal time difference between CUs is the same, the picture timing SEI message can still be applied even if the number of DUs changes. In one embodiment, there is a mode in the PT SEI message where the PT SEI message applies to sub-bitstreams with different numbers of DUs. The number of DUs is derived from other syntax elements. The PT SEI message changes as follows:
[0525] If du_not_constraint_flag is equal to 0, the value of du_common_cpb_removal_delay_flag is inferred to be 1. The value of num_decoding_units_minus1 is inferred to be equal to FrameTimeInterval divided by (du_common_cpb_removal_delay_increment_minus1+1)*ClockSubTicks minus 1.
[0526] The FrameTimeInterval may be signaled in a parameter set, in an HRD parameter in an SPS, or derived as the difference of the removal times of two consecutive access units at the highest temporal level.
[0527] In this case, there is an additional constraint that each DU should contain one VCL NAL unit.
[0528] Explicitly signaling the FrameTimeInterval can be done as follows:
[0529]
[0530] Among them, min_pic_duration_within_cvs_present_flag is a flag in SPS or VPS that indicates the presence of FrameTimeInterval when there is no constant frame rate. This allows FrameTimeInterval to be indicated in both constant frame rate and non-constant frame rate cases.
[0531] In the second case, du_not_constraint_flag can be set to 1 only when there is a constant frame rate. In this case, the value of FrameTimeInterval is derived as (elemental_duration_in_tc_minus1[maxTiD]+1) multiplied by ClockTicks.
[0532] Alternatively, the additional signaling flag can be omitted by merging the corresponding indication with the common DU timing mode signaling, as follows:
[0533]
[0534]
[0535] Note that the derivation of the number of DUs within an AU described above is based on the FrameTimeInterval and the common delay increment, and secondly, is given in terms of the number of clock ticks. Also note that in Aspect 4, the clock ticks vary depending on the temporal sub-layers present in the bitstream. The same result can be achieved by deriving the number of DUs using the clock ticks of the highest temporal layer, or by also considering the spreading factor discussed in Aspect 4 for the FrameTimeInterval.
[0536] Although some aspects have been described in the context of an apparatus, it is clear that these aspects also represent descriptions of corresponding methods, wherein a block or device corresponds to a method step or a feature of a method step. Similarly, aspects described in the context of a method step also represent descriptions of corresponding blocks or items or features of a corresponding apparatus. Some or all of the method steps can be performed by (or using) hardware devices such as, for example, a microprocessor, a programmable computer, or an electronic circuit. In some embodiments, one or more of the most important method steps can be performed by such a device.
[0537] Depending on certain implementation requirements, embodiments of the present invention can be implemented in hardware or in software or at least partially in hardware or at least partially in software. Implementations can be performed using a digital storage medium (e.g., a floppy disk, DVD, Blu-ray, CD, ROM, PROM, EPROM, EEPROM, or flash memory) having electronically readable control signals stored thereon that cooperate (or are capable of cooperating) with a programmable computer system to cause the corresponding method to be performed. Thus, the digital storage medium can be computer-readable.
[0538] Some embodiments according to the invention comprise a data carrier having electronically readable control signals, which are capable of cooperating with a programmable computer system, such that one of the methods described herein is performed.
[0539] Generally speaking, embodiments of the present invention can be implemented as a computer program product with a program code, the program code being operative to perform one of the methods when the computer program product runs on a computer. The program code may, for example, be stored on a machine-readable carrier.
[0540] Other embodiments comprise the computer program for performing one of the methods described herein, stored on a machine readable carrier.
[0541] In other words, an embodiment of the inventive method is, therefore, a computer program having a program code for performing one of the methods described herein, when the computer program runs on a computer.
[0542] A further embodiment of the inventive method is, therefore, a data carrier (or a digital storage medium or a computer-readable medium) comprising, recorded thereon, the computer program for carrying out one of the methods described herein. The data carrier, the digital storage medium or the recorded medium are typically tangible and / or non-transitory.
[0543] A further embodiment of the inventive method is, therefore, a data stream or a sequence of signals representing the computer program for carrying out one of the methods described herein. This data stream or signal sequence can, for example, be configured to be transferred via a data communication connection (for example, via the Internet).
[0544] Further embodiments comprise a processing means, for example a computer or a programmable logic device, configured to or adapted to carry out one of the methods described herein.
[0545] A further embodiment comprises a computer having installed thereon the computer program for performing one of the methods described herein.
[0546] Another embodiment according to the present invention comprises an apparatus or system configured to transfer (e.g., electronically or optically) a computer program for carrying out one of the methods described herein to a receiver. The receiver may be, for example, a computer, a mobile device, a memory device, etc. The apparatus or system may, for example, comprise a file server for transferring the computer program to the receiver.
[0547] In some embodiments, a programmable logic device (e.g., a field programmable gate array) can be used to implement some or all of the functionality of the methods described herein. In some embodiments, the field programmable gate array can cooperate with a microprocessor to implement one of the methods described herein. In general, the method is preferably implemented by any hardware device.
[0548] The devices described herein may be implemented using a hardware device or using a computer or using a combination of a hardware device and a computer.
[0549] The methods described herein may be implemented using a hardware device or using a computer or using a combination of a hardware device and a computer.
[0550] The above embodiments are merely illustrative of the principles of the present invention. It is understood that modifications and variations of the arrangements and details described herein will be readily apparent to those skilled in the art. It is therefore intended that the present invention be limited only by the scope of the appended claims and not by the specific details presented by way of description and explanation of the embodiments herein.
[0551] References
[0552] [1] ISO / IEC, ITU-T. High-Efficiency Video Coding. ITU-T Recommendation H.265 | ISO / IEC 23008-10 (HEVC), Version 1, 2013; Version 2, 2014.
Claims
1. A method for decoding a picture from a bitstream, the method comprising the following steps: decoding a first syntax element indicating a bit rate value and a second syntax element indicating a bit rate level; The code rate of the bit stream is calculated by the following formula: (first syntax element + 1) * 2 (6+第二语法元素) ; decoding a sub-picture level information (SLI) supplemental enhancement information (SEI) message indicating a score of a reference level corresponding to a sub-picture of the picture; as well as A rate for the sub-picture bitstream is calculated based in part on multiplying the fraction by the rate of the bitstream.
2. The decoding method according to claim 1 further comprises the following steps: decoding a third syntax element from the SLI SEI message, the third syntax element indicating a level of confirmation of the sub-picture bitstream; as well as A coded picture buffer size is determined based on the fraction and the third syntax element.
3. The method according to claim 1, wherein: The first and second syntax elements are associated with a bit rate of the picture; The sub-picture includes multiple slices; and The value of the first syntax element depends on the network abstraction layer (NAL) or the video coding layer (VCL) of the sub-picture.
4. The method of claim 1 , wherein determining the score comprises: decoding a third syntax element indicating a reference level score value from the SLI SEI message; as well as The score is determined based on the third syntax element.
5. The decoding method according to claim 1, wherein the bit rate of the sub-picture is expressed by the following formula: Floor(bitrate of the bitstream * the fraction ÷ 256).
6. The method of claim 1 , wherein determining the bit rate of the sub-picture bitstream comprises: The bit rate of the sub-picture bit stream is determined based on multiplying the fraction by the bit rate of the bit stream and right shifting.
7. A video decoder for decoding a picture from a bitstream, the video decoder comprising: A processor configured to: decoding a first syntax element indicating a bit rate value and a second syntax element indicating a bit rate level; The code rate of the bit stream is determined by the following formula: (first syntax element + 1) * 2 (6+第二语法元素) ; decoding a sub-picture level information (SLI) supplemental enhancement information (SEI) message indicating a score of a reference level corresponding to a sub-picture of the picture; as well as A rate for a sub-picture bitstream is determined based in part on multiplying the fraction by the rate of the bitstream.
8. The video decoder according to claim 7, configured to perform the method according to any one of claims 2 to 6.
9. A method for encoding a picture, the method comprising the following steps: A first syntax element indicating a rate value of the picture and a second syntax element indicating a rate level of the picture are encoded, wherein the first syntax element and the second syntax element indicate a rate of a bit stream represented by the following equation: (first syntax element+1)*2 (6+第二语法元素) ; determining a score corresponding to a reference level of a sub-picture of the picture; determining a bitrate for a sub-picture bitstream based in part on multiplying the fraction by the bitrate; as well as The fraction is encoded using a Sub-Picture Level Information (SLI) Supplemental Enhancement Information (SEI) message indicating a code rate of the sub-picture bitstream.
10. A video encoder for encoding a video into pictures, the video encoder comprising: A processor configured to: A first syntax element indicating a rate value of the picture and a second syntax element indicating a rate level of the picture are encoded, wherein the first syntax element and the second syntax element indicate a rate of a bit stream represented by the following equation: (first syntax element+1)*2 (6+第二语法元素) ; determining a score corresponding to a reference level of a sub-picture of the picture; determining a bitrate for a sub-picture bitstream based in part on multiplying the fraction by the bitrate; as well as The fraction is encoded using a Sub-Picture Level Information (SLI) Supplemental Enhancement Information (SEI) message indicating a code rate of the sub-picture bitstream.
11. A computer program product comprising a computer program including instructions which, when executed by a processor of a video decoder for decoding a picture from a bitstream, cause the video decoder to perform the method of any one of claims 1 to 6.
12. A computer program product comprising a computer program including instructions which, when executed by a processor of a video encoder for encoding a picture from a bitstream, cause the video encoder to perform the method of claim 9.