Method and apparatus for encoding and decoding a video data stream and technique for controlling sub-picture bitrate and buffer size
By incorporating scalable nested SEI messages for timing information in video data streams, the method addresses inefficiencies in HEVC's parallel processing, enhancing sub-picture and bit rate encoding and buffer management for improved video encoding and decoding efficiency.
Patent Information
- Application Number
- JP2025200815
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2019-12-20
- Filing Date
- 2025-11-20
- Publication Date
- 2026-02-18
AI Technical Summary
Existing video coding standards like HEVC lack efficient mechanisms for supporting parallel processing capabilities in video encoders and decoders, particularly in handling sub-pictures and bit rates, leading to inefficiencies in buffer management and network transmission.
The implementation of scalable nested Supplemental Enhancement Information (SEI) messages within video data streams to provide timing information for output layer sets, allowing for efficient sub-bitstream extraction and improved buffer management, thereby enhancing parallel processing capabilities.
This approach enables more efficient sub-picture and bit rate encoding, improving parallel processing and buffer management, ensuring HRD conformance, and optimizing network transmission.
Smart Images

Figure 2026027513000009 
Figure 2026027513000010 
Figure 2026027513000011
Abstract
Description
[Technical Field]
[0001] The present invention relates to video encoding and decoding, and in particular to methods and apparatus for encoding and decoding video data streams, and for controlling sub-picture bit rates and buffer sizes.
[0002] With H.265 / HEVC (HEVC = High Efficiency Video Coding), video codecs already provide tools to improve or further enable parallel processing in encoders and / or decoders. For example, HEVC supports the subdivision of a picture into an array of tiles that are coded independently of each other. Another concept supported by HEVC relates to WPP, according to which CTU rows or CTU lines of a picture can be processed in parallel from left to right, e.g., in stripes (CTU = coding tree unit), if some minimum CTU offset is respected during the processing of consecutive CTU lines. However, it would be advantageous to have a video codec available that more efficiently supports parallel processing capabilities in video encoders and / or video decoders.
[0003] Usually, in video coding, the coding process of picture samples requires smaller partitions, where the samples are divided into several rectangular regions for joint processing such as predictive coding or transform coding. Therefore, a picture is partitioned into blocks of a certain size that remains constant during the coding of a video sequence. In the H.264 / AVC standard, fixed-size blocks of 16x16 samples, so-called macroblocks, are used (AVC = Advanced Video Coding).
[0004] In the latest HEVC standard (see [1]), there are coding tree blocks (CTBs) or coding tree units (CTUs) of maximum size 64x64 samples. In further descriptions of HEVC, the more general term CTU is used for such kind of block.
[0005] The CTUs are processed in raster scan order, starting with the top left CTU and processing the CTUs in the picture line-wise to the bottom right CTU.
[0006] The coded CTU data is organized into a type of container called a slice. Originally, in previous video coding standards, a slice meant a segment containing one or more consecutive CTUs of a picture. Slices are used to segment coded data. From another perspective, a complete picture can also be defined as one large segment, so historically, the term slice continues to apply. In addition to the coded picture samples, slices also contain additional information related to the coding process of the slice itself, which is located in a so-called slice header.
[0007] According to the state of the art, the VCL (Video Coding Layer) also includes techniques for fragmentation and spatial partitioning. Such partitioning can be applied to video coding for various reasons, such as load balancing in parallel processing, CTU size matching in network transmission, error mitigation, etc.
[0008] Bitstreams specified in video coding standards have HRD conformance-related information. This conformance consists of a hypothetical reference decoder (HRD) that includes a buffer model that assumes that NAL units enter a coded picture buffer (CPB) before the decoder and are removed from there at specific times that ensure that the CPB size is not exceeded (buffer overrun) or that NAL units do not arrive later than they should be removed (buffer underrun). Furthermore, this model also includes a decoded picture buffer (DPB) from which decoded pictures are output when they are no longer needed for prediction, and whose size is similarly constrained in many implementations. HRD timing information is conveyed in the bitstream by so-called SEI messages, in particular a buffering period SEI message that defines specific timing information for a buffering period (BP) (a number of access units or AUs), a picture timing (PT) SEI message that conveys timing information for a single associated AU, and a decoding unit information (DUI) SEI message that conveys timing information for an associated subset of AUs, i.e., a decoding unit or DU.
[0009] Bitstreams specified in video coding standards contain information related to Hypothetical Reference Decoder (HRD) conformance. This conformance consists of a hypothetical buffer model that assumes that NAL units enter the Coded Picture Buffer (CPB) and are removed from it at specific times, ensuring that the CPB size is not exceeded (buffer overrun) or that NAL units do not arrive later than they need to be removed (buffer underrun).
[0010] If the bitstream is a scalable bitstream, pruning can be performed to obtain sub-bitstreams that are also conforming bitstreams. For example, if there is an output layer set (OLS) that includes three layers with resolution scalability (e.g., a 480p base layer, a first enhancement layer at 720p, and a second enhancement layer at 1080p), hereafter referred to as B3, two sub-bitstreams can be obtained: one sub-bitstream with two layers (480p and 720p) B2 and one sub-bitstream with one layer (480p) B1. Similarly, OLS can be used for temporal scalability, in which case B3, B2, and B1 have the same resolution but different frame rates.
[0011] Obviously, such bitstreams B3, B2 and B1 have different HRD conformance since the required CPB size, bit rate and timing information may differ.
[0012] The various CPB sizes and bit rates are indicated in the VPS as characteristics of the defined output layer set (3 in the example described). Various timing information is provided by so-called nesting SEI messages. The nested SEI messages can contain nested buffering period SEI and picture timing SEI messages that are applied to sub-bitstreams that can be obtained by bitstream pruning (bitstream extraction). Then, during this operation (extraction or pruning is performed), for example, the buffering period SEI message and picture timing SEI message of input bitstream B3 are removed from the bitstream, and also the NAL units belonging to the second enhancement layer are removed. Furthermore, the buffering period SEI message and picture timing SEI message corresponding to bitstream B2, carried in a nested SEI message, are placed in the bitstream from the nested SEI message bitstream, thus replacing the removed ones.
[0013] SUMMARY OF THE INVENTION It is an object of the present invention to provide an improved method and apparatus for sub-picture and bit rate encoding.
[0014] The object of the present invention is solved by the subject matter of the independent claims.
[0015] Preferred embodiments are provided in the dependent claims.
[0016] According to one embodiment, there is provided a video data stream having video encoded therein, the video data stream including an indication of whether one or more scalable nested Supplemental Enhancement Information (SEI) messages containing timing information for each of one or more output layer sets are present within the video data stream.
[0017] Further, according to one embodiment, there is provided a video encoder for encoding video into a video data stream, such that the video data stream has video encoded therein, the video encoder for generating the video data stream such that the video data stream includes an indication of whether one or more scalable nested supplemental enhancement information messages containing timing information for each of one or more output layer sets are present within the video data stream.
[0018] Additionally, according to one embodiment, there is provided an apparatus for receiving a video data stream as an input bitstream, the video data stream having video encoded therein, the apparatus being for processing the input bitstream to obtain sub-bitstreams, and an indication indicating whether one or more scalable nested supplemental enhancement information messages containing timing information for each of one or more output layer sets are present in the video data stream.
[0019] Further, according to an embodiment, there is provided a method for encoding video into a video data stream, such that the video data stream has video encoded therein, the method including generating a video data stream such that the video data stream includes an indication of whether one or more scalable nested supplemental enhancement information messages containing timing information for each of one or more output layer sets are present within the video data stream.
[0020] Further, according to one embodiment, there is provided a method for receiving a video data stream as an input bitstream, the video data stream having video encoded therein. The method includes processing the input bitstream to obtain sub-bitstreams. An indication indicates whether one or more scalable nested supplemental enhancement information messages containing timing information for each of one or more output layer sets are present in the video data stream.
[0021] Furthermore, a computer program is provided for implementing one of the above methods when run on a computer or signal processor.
[0022] Further, according to one embodiment, there is provided a video data stream having video encoded therein, wherein an indication within the video data stream indicates whether timing information for the sub-bitstreams is obtained from one or more non-scalable nested picture timing supplemental enhancement information messages of the video data stream.
[0023] Further, according to an embodiment, there is provided a video data stream having video encoded therein, wherein a first indication in the video data stream indicates whether timing information of a sub-bitstream is obtained from one or more non-scalable nested buffering period supplemental enhancement information messages of the video data stream, and / or a second indication in the video data stream indicates whether the timing information of a sub-bitstream is obtained from one or more non-scalable nested decoding unit supplemental enhancement information messages of the video data stream.
[0024] Furthermore, according to an embodiment, the video data stream includes one or more non-scalable nested supplemental enhancement information messages containing timing information. If the video data stream includes a scalable nested supplemental enhancement information message containing timing information, this indicates that, depending on the scalable nested supplemental enhancement information message, all of the one or more non-scalable nested timing information extension information messages are substituted by the scalable nested supplemental enhancement information message containing timing information (e.g., the timing information is picture timing information, buffering period information, or decoding unit information). Or, a subset including at least one of the one or more non-scalable nested timing information extension information messages is substituted by the non-scalable nested supplemental enhancement information message containing said timing information (e.g., the timing information is picture timing information, buffering period information, or decoding unit information).
[0025] Further, according to an embodiment, there is provided a video encoder for encoding video into a video data stream, such that the video data stream has video encoded therein, wherein the video encoder generates the video data stream such that an indication in the video data stream indicates whether timing information for a sub-bitstream is obtained from one or more non-scalable nested picture timing supplemental enhancement information messages of the video data stream.
[0026] Further, according to an embodiment, there is provided a video encoder for encoding video into a video data stream, such that the video data stream has video encoded therein, wherein a first indication in the video data stream indicates whether timing information of sub-bitstreams is obtained from the one or more non-scalable nested buffering period supplemental enhancement information messages of the video data stream, and / or a second indication in the video data stream indicates whether timing information of sub-bitstreams is obtained from one or more non-scalable nested decoding unit supplemental enhancement information messages of the video data stream.
[0027] Further, according to an embodiment, a video encoder is provided for encoding video into a video data stream, such that the video data stream has video encoded therein. The video encoder generates a video data stream such that the video data stream includes one or more non-scalable nested supplemental enhancement information messages including timing information. If the video data stream includes a scalable nested supplemental enhancement information message including timing information, this indicates that, in response to the scalable nested supplemental enhancement information message, all of the one or more non-scalable nested timing information enhancement information messages are substituted by a scalable nested supplemental enhancement information message including timing information (e.g., the timing information is picture timing information, buffering period information, or decoding unit information). Alternatively, a subset including at least one of the one or more non-scalable nested timing information supplemental enhancement information messages is substituted by the non-scalable nested supplemental enhancement information message including the timing information (e.g., the timing information is picture timing information, buffering period information, or decoding unit information).
[0028] Further, according to an embodiment, there is provided an apparatus for receiving a video data stream as an input bitstream, the video data stream having video encoded therein, the apparatus being for processing the input bitstream to obtain sub-bitstreams, and an indication in the video data stream indicating whether timing information for the sub-bitstreams is obtained from one or more non-scalable nested picture timing supplemental enhancement information messages of the video data stream.
[0029] Further, according to an embodiment, there is provided an apparatus for receiving a video data stream as an input bitstream, the video data stream having video encoded therein, the apparatus being configured to process the input bitstream to obtain sub-bitstreams, wherein a first indication in the video datastream indicates whether timing information for the sub-bitstreams is obtained from one or more non-scalable nested buffering period supplemental enhancement information messages of the video datastream, and / or a second indication in the video datastream indicates whether timing information for the sub-bitstreams is obtained from one or more non-scalable nested decoding unit supplemental enhancement information messages of the video datastream.
[0030] Further, according to an embodiment, there is provided an apparatus for receiving a video data stream as an input bitstream, the video data stream having video encoded therein. The apparatus is for processing the input bitstream to obtain sub-bitstreams. The video data stream includes one or more non-scalable nested supplemental enhancement information messages including timing information. If the video data stream includes a scalable nested supplemental enhancement information message including timing information, the apparatus is responsive to the scalable nested supplemental enhancement information message to substitute all of the one or more non-scalable nested timing information enhancement information messages with scalable nested supplemental enhancement information messages including timing information (e.g., the timing information is picture timing information, buffering period information, or decoding unit information). Alternatively, a subset including at least one of the one or more non-scalable nested timing information supplemental enhancement information messages is substituted with a non-scalable nested supplemental enhancement information message including timing information (e.g., the timing information is picture timing information, buffering period information, or decoding unit information).
[0031] Further, according to an embodiment, there is provided a method for encoding video into a video data stream, such that the video data stream has video encoded therein, the method including generating a video data stream such that indications within the video data stream are of sub-bitstream timing information obtained from one or more non-scalable nested picture timing supplemental enhancement information messages of the video data stream.
[0032] Further, according to an embodiment, there is provided a method for receiving a video data stream as an input bitstream, the video data stream having video encoded therein, the method including processing the input bitstream to obtain sub-bitstreams, wherein an indication in the video data stream indicates whether timing information for the sub-bitstreams is obtained from one or more non-scalable nested picture timing supplemental enhancement information messages of the video data stream.
[0033] Further, according to an embodiment, there is provided a method for encoding video into a video data stream, such that the video data stream has video encoded therein, wherein a first indication in the video data stream indicates whether timing information of the sub-bitstreams is obtained from one or more non-scalable nested buffering period supplemental enhancement information messages of the video data stream, and / or a second indication in the video data stream indicates whether timing information of the sub-bitstreams is obtained from one or more non-scalable nested decoding unit supplemental enhancement information messages of the video data stream.
[0034] Further, according to an embodiment, there is provided a method for receiving a video data stream as an input bitstream, the video data stream having video encoded therein. The method includes processing the input bitstream to obtain sub-bitstreams. A first indication in the video datastream indicates whether timing information for the sub-bitstreams is obtained from one or more non-scalable nested buffering period supplemental enhancement information messages of the video datastream. And / or a second indication in the video datastream indicates whether timing information for the sub-bitstreams is obtained from one or more non-scalable nested decoding unit supplemental enhancement information messages of the video datastream.
[0035] Further, according to an embodiment, there is provided a method for encoding video into a video data stream, such that the video data stream has video encoded therein. The method includes generating a video data stream such that the video data stream includes one or more non-scalable nested supplemental enhancement information messages that include timing information. If the video data stream includes a scalable nested supplemental enhancement information message that includes timing information, this indicates that at least one of the one or more non-scalable nested picture timing supplemental enhancement information messages is substituted by a scalable nested supplemental enhancement information message that includes timing information.
[0036] Further, according to an embodiment, there is provided a method for receiving a video data stream as an input bitstream, the video data stream having video encoded therein. The method includes processing the input bitstream to obtain sub-bitstreams. The video data stream includes one or more non-scalable nested supplemental enhancement information messages that include timing information. If the video data stream includes a scalable nested supplemental enhancement information message that includes timing information, the method includes substituting at least one of the one or more non-scalable nested picture timing supplemental enhancement information messages with a scalable nested supplemental enhancement information message that includes timing information.
[0037] Furthermore, a computer program is provided for implementing one of the above methods when run on a computer or signal processor.
[0038] Further, according to an embodiment, there is provided a video data stream having video encoded therein, the video data stream including a plurality of access units, wherein for each access unit of the plurality of access units, if the access unit includes two or more scalable nested supplemental enhancement information messages that include buffer period information and / or picture timing information of an output layer set, the buffer period information and / or picture timing information of the output layer set are equal in all of the two or more scalable nested supplemental enhancement information messages of the access unit.
[0039] Further, according to an embodiment, there is provided a video data stream having video encoded therein, the video data stream including a plurality of access units, and for each access unit of the plurality of access units, if the access unit includes two or more scalable nested supplemental enhancement information messages that include buffer period information and / or picture timing information of an output layer set, the video data stream including an indication of whether the buffer period information and / or picture timing information of the output layer set are equal in all two or more scalable nested supplemental enhancement information messages of the access unit.
[0040] Further, according to an embodiment, there is provided a video data stream having video encoded therein. The video data stream includes a plurality of access units. For each access unit of the plurality of access units, if the access unit includes three or more scalable nested supplemental enhancement information messages that include buffer period information and / or picture timing information of an output layer set, the buffer period information and / or picture timing information of the output layer set appears only in one of the three or more scalable nested supplemental enhancement information messages and in another of the three or more scalable nested supplemental enhancement information messages that immediately follows said one of the three or more scalable nested supplemental enhancement information messages.
[0041] Further, according to an embodiment, there is provided a video data stream having video encoded therein. The video data stream includes a plurality of access units. For each access unit of the plurality of access units, if the access unit includes three or more scalable nested supplemental enhancement information messages that include buffer period information and / or picture timing information for an output layer set, the video data stream includes an indication of whether the buffer period information and / or picture timing information for the output layer set appears only in one of the three or more scalable nested supplemental enhancement information messages and in another of the three or more scalable nested supplemental enhancement information messages that immediately follows the one of the three or more scalable nested supplemental enhancement information messages.
[0042] Further, according to an embodiment, there is provided a video encoder for encoding video into a video data stream, such that the video data stream has video encoded therein. The video encoder generates a video data stream such that the video data stream includes a plurality of access units. For each access unit of the plurality of access units, if the access unit includes two or more scalable nested supplemental enhancement information messages that include buffer period information and / or picture timing information of an output layer set, the video encoder generates the video data stream such that the video data stream includes an indication of whether the buffer period information and / or picture timing information of the output layer set are equal in all two or more scalable nested supplemental enhancement information messages of the access unit.
[0043] Further, according to an embodiment, there is provided a video encoder for encoding video into a video data stream, such that the video data stream has video encoded therein. The video encoder generates a video data stream, such that the video data stream includes a plurality of access units. For each access unit of the plurality of access units, if the access unit includes three or more scalable nested supplemental enhancement information messages that include buffer period information and / or picture timing information of an output layer set, the buffer period information and / or picture timing information of the output layer set appears only in one of the three or more scalable nested supplemental enhancement information messages and in another of the three or more scalable nested supplemental enhancement information messages that immediately follows the one of the three or more scalable nested supplemental enhancement information messages.
[0044] Further, according to an embodiment, there is provided a video encoder for encoding video into a video data stream, such that the video data stream has video encoded therein. The video encoder generates a video data stream such that the video data stream includes a plurality of access units. For each access unit of the plurality of access units, if the access unit includes three or more scalable nested supplemental enhancement information messages that include buffer period information and / or picture timing information of an output layer set, the video encoder generates the video data stream such that the video data stream includes an indication of whether the buffer period information and / or picture timing information of the output layer set appears only in one of the three or more scalable nested supplemental enhancement information messages and in another of the three or more scalable nested supplemental enhancement information messages that immediately follows the one of the three or more scalable nested supplemental enhancement information messages.
[0045] Further, according to an embodiment, there is provided an apparatus for receiving a video data stream as an input bitstream, the video data stream having video encoded therein. The input bitstream includes a plurality of access units. The apparatus processes the access units of the input bitstream to obtain sub-bitstreams. For each access unit of the plurality of access units, if the access unit includes two or more scalable nested supplemental enhancement information messages that include buffer period information and / or picture timing information of an output layer set, the buffer period information and / or picture timing information of the output layer set are equal in all of the two or more scalable nested supplemental enhancement information messages of the access unit.
[0046] Further, according to an embodiment, there is provided an apparatus for receiving a video data stream as an input bitstream, the video data stream having video encoded therein. The input bitstream includes a plurality of access units. The apparatus processes the access units of the input bitstream to obtain sub-bitstreams. For each access unit of the plurality of access units, if the access unit includes two or more scalable nested supplemental enhancement information messages including buffer period information and / or picture timing information of an output layer set, the video data stream includes an indication of whether the buffer period information and / or picture timing information of the output layer set are equal in all of the two or more scalable nested supplemental enhancement information messages of the access unit.
[0047] Further, according to an embodiment, there is provided an apparatus for receiving a video data stream as an input bitstream, the video data stream having video encoded therein. The input bitstream includes a plurality of access units, and the apparatus processes the access units of the input bitstream to obtain sub-bitstreams. For each access unit of the plurality of access units, if the access unit includes three or more scalable nested supplemental enhancement information messages that include buffer period information and / or picture timing information of an output layer set, the buffer period information and / or picture timing information of the output layer set appears only in one of the three or more scalable nested supplemental enhancement information messages and in another of the three or more scalable nested supplemental enhancement information messages that immediately follows said one of the three or more scalable nested supplemental enhancement information messages.
[0048] Further according to an embodiment, there is provided an apparatus for receiving a video data stream as an input bitstream, the video data stream having video encoded therein. The input bitstream includes a plurality of access units. The apparatus processes the access units of the input bitstream to obtain sub-bitstreams. For each access unit of the plurality of access units, if the access unit includes three or more scalable nested supplemental enhancement information messages including buffer period information and / or picture timing information of an output layer set, the video datastream includes an indication of whether the buffer period information and / or picture timing information of the output layer set appears only in one of the three or more scalable nested supplemental enhancement information messages and in another one of the three or more scalable nested supplemental enhancement information messages that immediately follows the one of the three or more scalable nested supplemental enhancement information messages.
[0049] Further, according to an embodiment, there is provided a method for encoding video into a video data stream, such that the video data stream has video encoded therein. The method includes generating a video data stream such that the video data stream includes a plurality of access units. For each access unit of the plurality of access units, if the access unit includes two or more scalable nested supplemental enhancement information messages that include buffer period information and / or picture timing information of an output layer set, the method includes generating the video data stream such that the buffer period information and / or picture timing information of the output layer set are equal in all two or more scalable nested supplemental enhancement information messages of the access unit.
[0050] Further, according to an embodiment, there is provided a method for encoding video into a video data stream, such that the video data stream has video encoded therein. The method includes generating a video data stream such that the video data stream includes a plurality of access units. For each access unit of the plurality of access units, if the access unit includes two or more scalable nested supplemental enhancement information messages that include buffer period information and / or picture timing information of an output layer set, the method includes generating the video data stream such that the video data stream includes an indication of whether the buffer period information and / or picture timing information of the output layer set are equal in all two or more scalable nested supplemental enhancement information messages of the access unit.
[0051] Further, according to an embodiment, there is provided a method for encoding video into a video data stream, such that the video data stream has video encoded therein. The method includes generating a video data stream, such that the video data stream includes a plurality of access units. For each access unit of the plurality of access units, if the access unit includes three or more scalable nested supplemental enhancement information messages that include buffer period information and / or picture timing information of an output layer set, the buffer period information and / or picture timing information of the output layer set appears only in one of the three or more scalable nested supplemental enhancement information messages and in another of the three or more scalable nested supplemental enhancement information messages that immediately follows the one of the three or more scalable nested supplemental enhancement information messages.
[0052] Further according to an embodiment, there is provided a method for encoding video into a video data stream, such that the video data stream has video encoded therein. The method includes generating a video data stream such that the video data stream includes a plurality of access units. For each access unit of the plurality of access units, if the access unit includes three or more scalable nested supplemental enhancement information messages that include buffer period information and / or picture timing information of an output layer set, the method includes generating the video data stream to include an indication of whether the buffer period information and / or picture timing information of the output layer set appears only in one of the three or more scalable nested supplemental enhancement information messages and in another one of the three or more scalable nested supplemental enhancement information messages that immediately follows the one of the three or more scalable nested supplemental enhancement information messages.
[0053] Further, according to an embodiment, there is provided a method for receiving a video data stream as an input bitstream, the video data stream having video encoded therein. The input bitstream includes a plurality of access units. The method includes processing the access units of the input bitstream to obtain sub-bitstreams. For each access unit of the plurality of access units, if the access unit includes two or more scalable nested supplemental enhancement information messages that include buffer period information and / or picture timing information of an output layer set, the buffer period information and / or picture timing information of the output layer set are equal in all of the two or more scalable nested supplemental enhancement information messages of the access unit.
[0054] Further, according to an embodiment, there is provided a method for receiving a video data stream as an input bitstream, the video data stream having video encoded therein. The input bitstream includes a plurality of access units. The method includes processing the access units of the input bitstream to obtain sub-bitstreams. For each access unit of the plurality of access units, if the access unit includes two or more scalable nested supplemental enhancement information messages that include buffer period information and / or picture timing information of an output layer set, the video data stream includes an indication of whether the buffer period information and / or picture timing information of the output layer set are equal in all two or more scalable nested supplemental enhancement information messages of the access unit.
[0055] Further, according to an embodiment, there is provided a method for receiving a video data stream as an input bitstream, the video data stream having video encoded therein. The input bitstream includes a plurality of access units. The method includes processing the access units of the input bitstream to obtain sub-bitstreams. For each access unit of the plurality of access units, if the access unit includes three or more scalable nested supplemental enhancement information messages that include buffer period information and / or picture timing information of an output layer set, the buffer period information and / or picture timing information of the output layer set appears only in one of the three or more scalable nested supplemental enhancement information messages and in another of the three or more scalable nested supplemental enhancement information messages that immediately follows the one of the three or more scalable nested supplemental enhancement information messages.
[0056] Further, according to an embodiment, there is provided a method for receiving a video data stream as an input bitstream, the video data stream having video encoded therein. The input bitstream includes a plurality of access units. The method includes processing the access units of the input bitstream to obtain sub-bitstreams. For each access unit of the plurality of access units, if the access unit includes three or more scalable nested supplemental enhancement information messages that include buffer period information and / or picture timing information of an output layer set, the video data stream includes an indication of whether the buffer period information and / or picture timing information of the output layer set appears only in one of the three or more scalable nested supplemental enhancement information messages and in another of the three or more scalable nested supplemental enhancement information messages that immediately follows the one of the three or more scalable nested supplemental enhancement information messages.
[0057] Furthermore, a computer program is provided for implementing one of the above methods when run on a computer or signal processor.
[0058] Further, according to an embodiment, there is provided a video data stream having video encoded therein, the video data stream including a plurality of access units, the video data stream including a spreading factor that depends on the number of sub-bitstreams of the video data stream, or the video data stream including a clock sub-tick value that depends on the highest sub-bitstream among the sub-bitstreams of the video data stream.
[0059] Further, according to an embodiment, there is provided a video encoder for encoding video into a video data stream, such that the video data stream has video encoded therein. The video encoder generates the video data stream, such that the video data stream includes a plurality of access units. Further, the video encoder generates the video data stream, such that the video data stream includes a spreading factor that depends on the number of sub-bitstreams of the video data stream, or the video encoder generates the video data stream, such that the video data stream includes a clock sub-tick value that depends on the highest sub-bitstream among the sub-bitstreams of the video data stream.
[0060] Further, according to an embodiment, there is provided a video encoder for encoding video into a video data stream, such that the video data stream has video encoded therein, wherein the video encoder generates the video data stream such that the video data stream includes sub-layer-specific frame rate information for the sub-layer, and / or the video encoder generates the video data stream such that the video data stream includes sub-layer-specific frame display time information for the sub-layer.
[0061] Further, according to an embodiment, there is provided a video decoder for receiving a video data stream as an input bitstream, the video data stream having video encoded therein, the video data stream including a plurality of access units, the video decoder being configured to decode the video data stream to decode the video, the video data stream including a spreading factor that depends on the number of sub-bitstreams of the video data stream, the video decoder being configured to decode the video using the spreading factor, or the video data stream including a clock sub-tick value that depends on the highest sub-bitstream of the video data stream, the video decoder being configured to decode the video using the clock sub-tick value.
[0062] Further, according to an embodiment, there is provided a method for receiving a video data stream as an input bitstream, the video data stream having video encoded therein. A video decoder is for decoding the video data stream to decode the video. The video data stream includes sub-layer-specific frame rate information for a sub-layer and / or the video data stream includes sub-layer-specific frame display period information for the sub-layer. The decoder is for determining a spreading coefficient using the sub-layer-specific frame rate information for the sub-layer and / or using the sub-layer-specific frame display period information.
[0063] Further, according to an embodiment, there is provided a method for encoding video into a video data stream, such that the video data stream has video encoded therein, the method including generating a video data stream such that the video data stream includes a plurality of access units, the method including generating the video data stream such that the video data stream includes a spreading factor that depends on the number of sub-bitstreams of the video data stream, or the method including generating the video data stream such that the video data stream includes a clock sub-tick value that depends on the highest sub-bitstream among the sub-bitstreams of the video data stream.
[0064] Further, according to an embodiment, there is provided a method for receiving a video data stream as an input bitstream, the video data stream having video encoded therein. The video data stream includes a plurality of access units. The method includes decoding the video data stream to decode the video. The video data stream includes a spreading factor that depends on a number of sub-bitstreams of the video data stream, and the method includes decoding the video using the spreading factor. Alternatively, the video data stream includes a clock sub-tick value that depends on a highest sub-bitstream of the sub-bitstreams of the video data stream, and the method includes decoding the video using the clock sub-tick value.
[0065] Further according to an embodiment, there is provided a method for receiving a video data stream as an input bitstream, the video data stream having video encoded therein. The method includes decoding the video data stream to decode the video. The video data stream includes sub-layer-specific frame rate information for a sub-layer and / or the video data stream includes sub-layer-specific frame display duration information for a sub-layer. The method includes determining a spreading coefficient using the sub-layer-specific frame rate information for a sub-layer and / or using the sub-layer-specific frame display duration information.
[0066] Furthermore, a computer program is provided for implementing one of the above methods when run on a computer or signal processor.
[0067] According to an embodiment, a video decoder is provided for receiving a video data stream as an input bitstream, the video data stream having video encoded therein. The video decoder is configured to decode the video data stream to decode the video. To decode the video, the video decoder estimates a coded picture buffer size for a sub-picture depending on information in the video data stream indicating current coded picture buffer size information.
[0068] Further, according to an embodiment, there is provided a video decoder for receiving a video data stream as an input bitstream, the video data stream having video encoded therein, the video decoder being configured to decode the video data stream to decode said video, and for decoding the video, the video decoder being configured to estimate a bitrate of a sub-picture as a function of information in the video data stream indicating bitrate information of a video sequence currently being coded.
[0069] Further, according to an embodiment, there is provided a method for receiving a video data stream as an input bitstream, the video data stream having video encoded therein, and a video decoder for decoding the video data stream to decode the video, wherein to decode the video, the video decoder receives a coded picture buffer size of a sub-picture coded in the video data stream and uses the coded picture buffer size of the sub-picture to decode the video, and / or the video decoder receives a bitrate of the sub-picture coded in the video data stream and decodes the video using the bitrate of the sub-picture.
[0070] Further, according to an embodiment, there is provided a video data stream having video encoded therein, the video data stream including, for example, the syntax element cpb_size_value_minus1[i][j] and the syntax element cpb_size_scale, or the video data stream including the syntax element bit_rate_value_minus1[i][j] and the syntax element bit_rate_scale.
[0071] Further, according to an embodiment, a video data stream having video encoded therein is provided, the video data stream including an indication indicating whether the coded picture buffer size of a sub-picture should be estimated using current coded picture buffer size information and / or the video data stream including an indication indicating whether the bit rate of a sub-picture is estimated using bit rate information of the video sequence currently being coded.
[0072] Further, according to an embodiment, a video data stream having video encoded therein is provided, wherein the video data stream includes an indication indicating whether a coded picture buffer size for a sub-picture has been encoded in the video data stream or whether the coded picture buffer size for a sub-picture should be estimated, and / or the video data stream includes an indication indicating whether a bit rate for a sub-picture has been encoded in the video data stream or whether the bit rate for a sub-picture needs to be estimated.
[0073] Further, according to an embodiment, there is provided a video encoder for encoding video into a video data stream, such that the video data stream has video encoded therein. The video encoder generates the video data stream, such that the video data stream includes, for example, the syntax element cpb_size_value_minus1[i][j] and the syntax element cpb_size_scale. Alternatively, the video encoder generates the video data stream, such that the video data stream includes the syntax element bit_rate_value_minus1[i][j] and the syntax element bit_rate_scale.
[0074] Further, according to an embodiment, there is provided a video encoder for encoding video into a video data stream, such that the video data stream has video encoded therein, wherein the video encoder generates the video data stream such that the video data stream includes an indication indicating whether a coded picture buffer size of a sub-picture should be estimated using current coded picture buffer size information, and / or the video encoder generates the video data stream such that the video data stream includes generating an indication indicating whether a bitrate of a sub-picture is estimated using currently coded video sequence bitrate information.
[0075] Further, according to an embodiment, there is provided a video encoder for encoding video into a video data stream, such that the video data stream has video encoded therein, wherein the video encoder generates the video data stream such that the video data stream includes an indication indicating whether a coded picture buffer size for a sub-picture has been encoded in the video data stream or whether the coded picture buffer size for a sub-picture should be estimated, and / or the video encoder generates the video data stream such that the video data stream includes an indication indicating whether a bit rate for a sub-picture has been encoded in the video data stream or whether the bit rate for a sub-picture needs to be estimated.
[0076] Further, according to an embodiment, there is provided a video encoder for encoding video into a video data stream, such that the video data stream has video encoded therein. The video encoder generates the video data stream such that, when the video data stream includes common decoding unit removal timing information and multiple extractable sub-bitstreams, each of the multiple sub-bitstreams is specific to an output layer set, and parsing each output layer set-specific hypothetical reference decoder parameter in a video parameter set or a sequence parameter set or a supplemental enhancement information message of the video data stream includes either a spreading factor or an absolute value of a tick divisor in scaling the removal timing of the common decoding unit.
[0077] Further, according to an embodiment, there is provided a method for receiving a video data stream as an input bitstream, the video data stream having video encoded therein, the method including decoding the video data stream to decode the video, the method including estimating a coded picture buffer size for a sub-picture as a function of information in the video data stream indicating current coded picture buffer size information.
[0078] Further, according to an embodiment, there is provided a method for receiving a video data stream as an input bitstream, the video data stream having video encoded therein, the method comprising: decoding the video data stream to decode the video, and for decoding the video, the method comprises estimating a sub-picture bitrate as a function of information in the video data stream indicating currently coded video sequence bitrate information.
[0079] Further, according to an embodiment, there is provided a method for receiving a video data stream as an input bitstream, the video data stream having video encoded therein. The method includes decoding the video data stream to decode the video. To decode the video, the method includes receiving a coded picture buffer size of a sub-picture coded in the video data stream and using the coded picture buffer size of the sub-picture to decode the video, and / or the method includes receiving a bitrate of the sub-picture coded in the video data stream and decoding the video using the bitrate of the sub-picture.
[0080] Further, according to an embodiment, there is provided a method for encoding video into a video data stream, such that the video data stream has video encoded therein, the method comprising generating the video data stream such that the video data stream includes syntax elements cpb_siges_veret_minus1[i][j] and cpb_zesic_class, or the method comprises generating the video data stream such that the video data stream includes syntax elements bit_rate_value_minus1[i][j] and bit_rate_scale.
[0081] Further, according to an embodiment, there is provided a method for encoding video into a video data stream, such that the video data stream has video encoded therein, the method including generating the video data stream such that the video data stream includes an indication indicating whether a coded picture buffer size of a sub-picture should be estimated using current coded picture buffer size information, and / or the method including generating the video data stream such that the video data stream includes an indication indicating whether a bit rate of a sub-picture is estimated using bit rate information of a video sequence currently being coded.
[0082] Further, according to an embodiment, there is provided a method for encoding video into a video data stream, such that the video data stream has video encoded therein, the method comprising generating the video data stream such that the video data stream includes an indication indicating whether a coded picture buffer size for a sub-picture has been encoded within the video data stream or whether the coded picture buffer size for a sub-picture should be estimated, and / or the method comprises generating the video data stream such that the video data stream includes an indication indicating whether a bit rate for a sub-picture has been encoded within the video data stream or whether the bit rate for a sub-picture needs to be estimated.
[0083] Furthermore, a computer program is provided for implementing one of the above methods when run on a computer or signal processor.
[0084] According to an embodiment, a video data stream having video encoded therein is provided, the video data stream including a plurality of access units, and further including delta time information for each of two or more decoding units of the plurality of access units, wherein a decoding unit removal time for each of the two or more decoding units of the access units is dependent on the access unit removal time of the access unit and is dependent on the delta time information for the decoding unit.
[0085] Further, according to an embodiment, there is provided a video encoder for encoding video into a video data stream, such that the video data stream has video encoded therein. The video encoder generates the video data stream, such that the video data stream includes a plurality of access units. Furthermore, the video encoder generates the video data stream, such that the video data stream includes delta time information for each of two or more decoding units of the plurality of access units, wherein a decoding unit removal time of each decoding unit of the two or more decoding units of the access unit depends on the access unit removal time of the access unit and depends on the delta time information of the decoding unit.
[0086] Further, according to an embodiment, there is provided a method for receiving a video data stream as an input bitstream, the video data stream having video encoded therein. The video data stream includes a plurality of access units. A video decoder is configured to decode the video data stream to decode the video. Furthermore, the video data stream includes delta time information for each of two or more decoding units of the plurality of access units, wherein a decoding unit removal time of each decoding unit of the two or more decoding units of the access units depends on the access unit removal time of the access unit and depends on the delta time information of the decoding unit, and the video decoder decodes the video data stream using the delta time information for each of the two or more decoding units of the access unit.
[0087] Further, according to an embodiment, there is provided a method for encoding video into a video data stream, such that the video data stream has video encoded therein, the method including generating the video data stream such that the video data stream includes a plurality of access units, the method further including generating the video data stream such that the video data stream includes delta time information for each of two or more decoding units of the plurality of access units, and a decoding unit removal time for each decoding unit of the two or more decoding units of the access unit depends on the access unit removal time of the access unit and depends on the delta time information of the decoding unit.
[0088] Further, according to an embodiment, there is provided a method for receiving a video data stream as an input bitstream, the video data stream having video encoded therein. The video data stream includes a plurality of access units. The method includes decoding the video data stream to decode the video. Furthermore, the video data stream includes delta time information for each of two or more decoding units of the plurality of access units, and a decoding unit removal time for each decoding unit of the two or more decoding units of the access unit depends on the access unit removal time of the access unit and depends on the delta time information for the decoding unit, and the method includes decoding the video data stream using the delta time information for each of the two or more decoding units of the access unit.
[0089] Furthermore, a computer program is provided for implementing one of the above methods when run on a computer or signal processor. [Brief explanation of the drawings]
[0090] [Figure 1] 1 illustrates a video encoder for encoding video into a video data stream according to an embodiment; [Figure 2]1 illustrates an apparatus for receiving an input video data stream according to an embodiment. [Figure 3] 1 illustrates a video decoder for receiving a video data stream having video stored therein, according to an embodiment; [Figure 4] The variation of removal time when there are three decode units per access unit is shown. [Figure 5] Two layers and the removal times for access and decode units for both layers are shown. [Figure 6] 1 shows a video encoder. [Figure 7] 1 shows a video decoder. [Figure 8] It shows the relationship between, on the one hand, a reconstructed signal, e.g. a reconstructed picture, and, on the other hand, a combination of a prediction residual signal and a prediction signal signaled in a data stream. DETAILED DESCRIPTION OF THE INVENTION
[0091] The following description of the figures begins with presenting a description of an encoder and decoder of a block-based predictive codec for encoding pictures of video, to form an example of an encoding framework in which embodiments of the present invention may be incorporated. Respective encoders and decoders are described with reference to Figures 6 through 8. Below, descriptions of embodiments of the inventive concepts are presented along with an explanation of how such concepts may be incorporated into the respective encoders and decoders of Figures 6 and 7, although the embodiments described in Figures 1 through 3 and thereafter may also be used to form encoders and decoders that do not operate according to the underlying encoding framework of the encoders and decoders of Figures 6 and 7.
[0092] FIG. 6 illustrates a video encoder, illustratively an apparatus for predictively encoding picture 12 into data stream 14 using transform-based residual coding. The apparatus, or encoder, is indicated using the reference symbol 10. FIG. 7 illustrates a corresponding video decoder 20, e.g., apparatus 20, also configured to predictively decode picture 12′ from data stream 14 using transform-based residual decoding; an apostrophe is used to indicate that picture 12′ reconstructed by decoder 20 deviates from picture 12 originally encoded by apparatus 10 in terms of coding loss introduced by quantization of the prediction residual signal. While FIGS. 6 and 7 illustratively use transform-based predictive residual coding, embodiments of the present application are not limited to this type of predictive residual coding. This also applies to other details described with respect to FIGS. 6 and 7, as outlined below.
[0093] The encoder 10 is configured to subject the prediction residual signal to a spatial-to-spectral transform and encode the prediction residual signal into a data stream 14, thereby obtaining the prediction residual signal. Similarly, the decoder 20 is configured to decode the prediction residual signal from the data stream 14 and subject the prediction residual signal to a spectral-to-spatial transform, thereby obtaining the prediction residual signal.
[0094] Internally, the encoder 10 may include a prediction residual signal former 22, which generates a prediction residual 24 to measure the deviation of a prediction signal 26 from an original signal, e.g., picture 12. The prediction residual signal former 22 may be, for example, a subtractor that subtracts the prediction signal from the original signal, e.g., picture 12. The encoder 10 then further includes a transformer 28, which subjects the prediction residual signal 24 to a spatial-to-spectral transform to obtain a spectral-domain prediction residual signal 24'. The spectral-domain prediction residual signal is then quantized by a quantizer 32 and included by the encoder 10. The quantized prediction residual signal 24" is therefore coded into the bitstream 14. To this end, the encoder 10 may optionally include an entropy coder 34, which entropy codes the transformed and quantized prediction residual signal into the data stream 14. A prediction signal 26 is generated by a prediction stage 36 of the encoder 10 based on the prediction residual signal 24", which is coded into the data stream 14 and is decodable therefrom. To this end, the prediction stage 36 may internally include, as shown in FIG. 6, an inverse quantizer 38 that inversely quantizes the prediction residual signal 24" to obtain a spectral-domain prediction residual signal 24'" that corresponds to the signal 24' except for quantization losses, and then an inverse transformer 40 that subjects the latter prediction residual signal 24'" to an inverse transform, for example a spectral-to-spatial transform, to obtain a prediction residual signal 24'''' that corresponds to the original prediction residual signal 24 except for quantization losses. A combiner 42 of the prediction stage 36 then recombines the prediction signal 26 and the prediction residual signal 24'''', for example by addition, to obtain a reconstructed signal 46, for example a reconstruction of the original signal 12. The reconstructed signal 46 may correspond to the signal 12'. A prediction module 44 of the prediction stage 36 then generates the prediction signal 26 based on the signal 46, for example by using spatial prediction (such as intra-picture prediction) and / or temporal prediction (such as inter-picture prediction).
[0095] Similarly, decoder 20 may be internally constructed from components corresponding to prediction stage 36 and interconnected in a manner corresponding to the prediction stage, as shown in Figure 7. In particular, entropy decoder 50 of decoder 20 may entropy decode a quantized spectral domain prediction residual signal 24" from the data stream, where inverse quantizer 52, inverse transformer 54, synthesizer 56 and prediction module 58 are interconnected and cooperate in the manner described above with respect to the modules of prediction stage 36 to recover a reconstructed signal based on prediction residual signal 24" such that the output of synthesizer 56 results in a reconstructed signal, i.e., picture 12', as shown in Figure 7.
[0096] Although not specifically described above, it will be readily apparent that the encoder 10 may set some coding parameters, including, for example, prediction modes, motion parameters, etc., according to some optimization scheme, such as a method for optimizing some rate- and distortion-related criteria, such as coding cost. For example, the encoder 10 and decoder 20 and corresponding modules 44, 58 may each support different prediction modes, such as intra-coding and inter-coding modes. The granularity at which the encoder and decoder switch between their prediction mode types may correspond to the subdivision of the pictures 12 and 12′, respectively, into coding segments or coding blocks. In units of these coding segments, for example, a picture may be subdivided into intra-coded blocks and inter-coded blocks. The intra-coded blocks are predicted based on their spatial, already coded / decoded neighbors, as outlined in more detail below. Several intra-coding modes may be present and selected for each intra-coded segment, including directional or angular intra-coding modes, according to which each segment is filled by extrapolating neighboring sample values along a specific direction specific to each directional intra-coding mode into the respective intra-coded segment. These intra-coding modes may also include one or more additional modes, such as a DC coding mode, according to which prediction of each intra-coded block assigns a DC value to all samples in the respective intra-coded segment, and / or a planar intra-coding mode, according to which prediction of each block is approximated and determined to be a spatial distribution of sample values described by a two-dimensional linear function over the sample positions of each intra-coded block, as determined by the slope and offset of the plane defined by the two-dimensional linear function based on neighboring samples. In contrast, inter-coded blocks may be predicted, for example, temporally.For inter-coded blocks, motion vectors may be signaled within the data stream, indicating the spatial displacement of portions of previously coded pictures of the video to which picture 12 belongs, where the previously coded / decoded pictures are sampled to obtain a prediction signal for each inter-coded block. This means that in addition to the coding of the residual signal that data stream 14 contains, such as entropy-coded transform coefficient levels representing the quantized spectral-domain prediction residual signal 24", data stream 14 may also include coded therein coding mode parameters for assigning coding modes to various blocks, prediction parameters for some of the blocks, such as motion parameters for inter-coded segments, and optional further parameters, such as parameters for controlling and signaling the subdivision of pictures 12 and 12' into their respective segments. Decoder 20 uses these parameters to subdivide the picture in the same way as the encoder did, assign the same prediction modes to the segments, and perform the same prediction, resulting in the same prediction signal.
[0097] 8 shows the relationship between, on the one hand, a reconstructed signal, e.g., a reconstructed picture 12′, and, on the other hand, a combination of a prediction residual signal 24″″ signaled in the data stream 14 and a prediction signal 26. As already indicated above, this combination may be additive. The prediction signal 26 is shown in FIG. 8 as a subdivision of the picture region into intra-coded blocks, exemplarily shown using hatching, and inter-coded blocks, exemplarily shown without hatching. This subdivision may be any subdivision, such as a regular subdivision of the picture region into rows and columns of square or non-square blocks, or a multi-tree subdivision of the picture 12 from a tree root block into multiple leaf blocks of various sizes (e.g., quad-tree subdivision, etc.), a combination of which is shown in FIG. 8, in which the picture region is first subdivided into rows and columns of a tree root block, and then further subdivided into one or more leaf blocks according to a recursive multi-tree subdivision.
[0098] Again, data stream 14 may have an intra-coding mode coded therein for intra-coded blocks 80, such that one of several supported intra-coding modes is assigned to each intra-coded block 80. In the case of inter-coded blocks 82, data stream 14 may have one or more motion parameters coded therein. In general, inter-coded blocks 82 are not restricted to being temporally coded. Instead, inter-coded blocks 82 may be any blocks that are predicted from previously coded portions beyond current picture 12 itself, for example, from a previously coded picture of the video to which picture 12 belongs, or from a picture of another view or a hierarchically lower layer if the encoder and decoder are scalable encoder and decoder, respectively.
[0099] In FIG. 8 , the prediction residual signal 24″″ is also shown as a subdivision of the picture region into blocks 84. These blocks may be referred to as transform blocks to distinguish them from the coding blocks 80 and 82. In fact, FIG. 8 shows that the encoder 10 and the decoder 20 may use two different subdivisions of the picture 12 and the picture 12′ into blocks, respectively: one subdivision into the respective coding blocks 80 and 82, and another subdivision into the transform blocks 84. While both subdivisions may be the same, e.g., the coding blocks 80 and 82, respectively, and may simultaneously form the transform blocks 84, FIG. 8 also shows the case where, for example, the subdivision into the transform blocks 84 forms an extension of the subdivision into the coding blocks 80, 82, so that any boundary between the two blocks 80 and 82 overlaps with the boundary between the two blocks 84, or, in other words, each block 80, 82 coincides with one of the transform blocks 84 or with a cluster of transform blocks 84. However, these subdivisions may also be determined or selected independently of one another such that transformation blocks 84 may alternately cross the block boundaries between blocks 80 and 82. Thus, as far as the subdivision into transformation blocks 84 is concerned, similar statements are true as made with regard to the subdivision into blocks 80, 82; for example, blocks 84 may be the result of a regular subdivision of a picture region into blocks (with or without arrangement into rows and columns), the result of a recursive multi-tree subdivision of a picture region, or a combination thereof or any other type of block result. As an aside, it is noted that blocks 80, 82 and 84 are not limited to being square, rectangular or of any other shape.
[0100] 8 further illustrates that the combination of prediction signal 26 and prediction residual signal 24"" directly results in reconstructed signal 12'. Note, however, that in alternative embodiments, more than one prediction signal 26 may be combined with prediction residual signal 24"" to result in picture 12'.
[0101] In Fig. 8, the transform blocks 84 have the following significance: the transformer 28 and the inverse transformer 54 perform their transforms in units of their transform blocks 84. For example, many codecs use some kind of DST or DCT for all transform blocks 84. Some codecs allow to skip transforms, so that for some transform blocks 84, the prediction residual signal is directly coded in the spatial domain. However, according to the embodiment below, the encoder 10 and the decoder 20 are configured in such a way that they support several transforms. For example, the transforms supported by the encoder 10 and the decoder 20 are: DCT-II (or DCT-III) (DCT stands for Discrete Cosine Transform) o DST-IV (DST stands for Discrete Sine Transform) DCT-IV DST-VII o Identity Transformation (IT) may include
[0102] Naturally, the transformer 28 supports all of the forward transform versions of those transforms, but the decoder 20 or inverse transformer 54 supports only their corresponding backward or inverse versions, o Inverse DCT-II (or Inverse DCT-III), o Reverse DST-IV o Inverse DCT-IV o Reverse DST-VII o Identity Transformation (IT) Support.
[0103] Subsequent descriptions provide further details on which transforms may be supported by the encoder 10 and decoder 20. Note that in any case, the supported transform set may include only one transform, such as one spectral-to-spatial or spatial-to-spectral transform.
[0104] As already outlined above, Figures 6 to 8 are presented as an example in which the inventive concepts described further below may be implemented to form specific examples of an encoder and decoder according to the present application. Thus far, the encoder and decoder, respectively, of Figures 6 and 7 may represent possible implementations of the encoder and decoder described below in this specification. However, Figures 6 and 7 are merely examples. However, an encoder according to an embodiment of the present application may use concepts outlined in more detail below to perform block-based encoding of picture 12 that differs from the encoder of Figure 6 in that it is not a video encoder but is still a picture encoder, that it does not support inter-prediction, or that the subdivision into blocks 80 is performed differently from the way illustrated in Figure 8. Similarly, a decoder according to an embodiment of the present application may perform block-based decoding of picture 12′ from data stream 14 using the coding concepts further outlined below, but may differ from decoder 20 of FIG. 7 in that, for example, it is not a video decoder but is still a picture decoder, that it does not support intra prediction, or that it subdivides picture 12′ into blocks in a different way than described with respect to FIG. 8, and / or that it derives prediction residuals from data stream 14 in the spatial domain rather than the transform domain, for example.
[0105] 1 shows a video encoder 100 for encoding video into a video data stream according to one embodiment. The video encoder 100 is configured to generate a video data stream.
[0106] 2 shows an apparatus 200 for receiving an input video data stream having video encoded therein, according to one embodiment, and configured to generate an output video data stream from the input video data stream.
[0107] 3 shows a video decoder 300 for receiving a video data stream having video stored therein, according to one embodiment. The video decoder 300 is configured to decode video from the video data stream.
[0108] Furthermore, a system according to an embodiment is presented, which includes the device of Figure 2 and the video decoder of Figure 3. The video decoder (300) of Figure 3 is configured to receive the output video data stream of the device (200) of Figure 2. The video decoder 300 of Figure 3 is configured to decode video from the output video data stream of the device 200 of Figure 2.
[0109] In one embodiment, the system may further include, for example, the video encoder 100 of Figure 1. The apparatus 200 of Figure 2 may, for example, be configured to receive the video data stream from the video encoder 100 of Figure 1 as an input video data stream.
[0110] Intermediate device 210 (optional) of apparatus 200 may be configured, for example, to receive a video data stream from video encoder 100 as an input video data stream and to generate an output video data stream from the input video data stream. For example, the intermediate device may be configured, for example, to modify information (header / metadata information) of the input video data stream and / or may be configured, for example, to remove pictures from the input video data stream and / or may be configured to mix / splice the input video data stream with an additional second bitstream having a second video encoded therein.
[0111] The video decoder 221 (optional) may be configured, for example, to decode video from the output video data stream.
[0112] The hypothetical reference decoder 222 (optional) may be configured, for example, to determine timing information of the video according to the output video data stream, or may be configured, for example, to determine buffer information of a buffer in which the video or a portion of the video is stored.
[0113] The system includes a video encoder 101 of FIG. 1 and a video decoder 151 of FIG.
[0114] The video encoder 101 is configured to generate an encoded video signal. The video decoder 151 is configured to decode the encoded video signal to reconstruct video pictures.
[0115] Specific embodiments are described below.
[0116] In HEVC, a note in the Extraction Process specification explains the processing of nested SEI messages as follows:
[0117] A "smart" bitstream extractor can include appropriate non-scalable nested buffering picture SEI messages, non-scalable nested picture timing SEI messages, and non-scalable nested decoding unit information SEI messages in the extracted sub-bitstreams, provided that the messages applicable to the sub-bitstreams were present as scalably nested SEI messages in the original bitstream.
[0118] In VVC, the assumed design is properly set up so that the behavior is normatively specified, e.g., for the extraction process defined in JVET-P2001-vC with the embodiments of the present invention added, as follows:
[0119] Sub-bitstream Extraction Process The inputs to this process are the bitstream inBitstream, the target OLS index targetOlsIdx, and the target's highest TemporalId value tIdTarget.
[0120] The output of this process is the sub-bitstream outBitstream. For a bitstream, it is a requirement of the input bitstream that any output sub-bitstream that is the output of a process specified in this clause shall have targetOlsIdx equal to an index in the list of OLSs specified in the VPS, and tIdTarget equal to any value in the range 0 to 6, inclusive, and that satisfying the following conditions shall be a conforming bitstream: The output sub-bitstream contains at least one VCL NAL unit with a nuh_layer_id value equal to each of the LayerIdInOls[targetOlsIdx]. · The output sub-bitstream contains at least one VCNLAL unit whose TemporalId is equal to tIdTarget. NOTE − A compliant bitstream contains one or more coded slice NAL units with TemporalId equal to 0, but is not required to contain any coded slice NAL units with nuh_layer_id equal to 0. The output sub-bitstream OutBitstream is derived as follows. The bitstream outBitstream is set to be the same as the bitstream inBitstream. Remove all NAL units with TemporalId greater than tIdTarget from outBitstream. Remove all NAL units from outBitstream whose nal_unit_type is not equal to any of VPS_NUT, DPS_NUT, and EOB_NUT and whose nuh_layer_id is not included in the list LayerIdInOls[targetOlsIdx]. Remove from outBitstream all SEI NAL units containing scalable nesting SEI messages where nesting_ols_flag is 1 and there are no values of i in the range from 0 to nesting_num_olss_minus1, inclusive, e.g., NestingOlsIdx[i] is the same as targetOlsIdx. If targetOlsIdx is greater than 0, remove all SEI NAL units from outBitstream that contain non-scalable nested SEI messages with payloadType equal to 0 (buffering period), 1 (picture timing), or 130 (decode unit information).
[0121] According to a particular embodiment: · If outBitstream contains a scalable nesting SEI message with nesting_ols_flag equal to 1 and contains SEI NAL units applicable to outBitstream (NestingOlsIdx[i] equals targetOlsIdx), then do the following: Extract the appropriate non-scalable nested SEI messages from the scalable nested SEI message where payloadType equals 0 (buffering period), 1 (picture timing), or 130 (decode information) and put those messages into outBitstream. Remove all SEI NAL units, including scalable nested SEI messages, from outBitstream.
[0122] The following describes the presence of OLS scalable nested SEI messages in the bitstream.
[0123] According to an embodiment, there is provided a video data stream having video encoded therein, the video data stream including an indication of whether one or more scalable nested supplemental enhancement information messages containing timing information for each of one or more sets of output layers are present within the video data stream.
[0124] In an embodiment, the indication is a parameter set flag, e.g., the video data stream includes a parameter set flag that indicates whether one or more scalable nested supplemental enhancement information messages containing timing information for each of one or more output layer sets are present in the video data stream.
[0125] According to an embodiment, in a video data stream as claimed in claim 2, the sequence parameter sets of the video data stream may for example comprise a parameter set flag.
[0126] In an embodiment, the parameter setting flag is sps_ols_nest_timing_present_flag.
[0127] According to an embodiment, the video data stream may include further supplemental enhancement information messages, which may include, for example, parameter set flags indicating whether one or more scalable nested supplemental enhancement information messages for each of one or more output layer sets are present in the video data stream.
[0128] In an embodiment, the timing information may include, for example, at least one of picture timing information, buffering period information, and decoding unit information.
[0129] In an embodiment, the timing information is timing information of a hypothetical reference decoder.
[0130] Further, according to an embodiment, there is provided a video encoder for encoding video into a video data stream, such that the video data stream has video encoded therein, the video encoder for generating the video data stream such that the video data stream includes an indication of whether one or more scalable nested supplemental enhancement information messages containing timing information for each of one or more output layer sets are present within the video data stream.
[0131] In an embodiment, the indication is a parameter set flag. The video encoder may, for example, be configured to generate the video data stream such that the video data stream may include, for example, a parameter set flag that indicates whether one or more scalable nested supplemental enhancement information messages containing timing information for each of one or more output layer sets are present in the video data stream.
[0132] In an embodiment, the video encoder may be configured to generate a video data stream such that a series of parameter sets for the video data stream may include, for example, a decoding parameter set.
[0133] In an embodiment, the video encoder may for example be configured to generate the video data stream such that the parameter set flag is sps_ols_nest_timing_present_flag.
[0134] In an embodiment, the video encoder is, for example, configured to generate the video data stream such that the video data stream may include, for example, further supplemental enhancement information messages, which may include, for example, parameter set flags. The video encoder may, for example, be configured to generate the video data stream such that the parameter set flags indicate whether one or more scalable nested supplemental enhancement information messages for each of one or more output layer sets are present in the video data stream.
[0135] According to an embodiment, the timing information may include, for example, at least one of picture timing information, buffering period information, and decoding unit information.
[0136] In an embodiment, the timing information is timing information of a hypothetical reference decoder.
[0137] Further, according to an embodiment, there is provided an apparatus for receiving a video data stream as an input bitstream, the video data stream having video encoded therein. The apparatus is for processing the input bitstream to obtain sub-bitstreams. An indication indicates whether one or more scalable nested supplemental enhancement information messages containing timing information for each of one or more output layer sets are present in the video data stream.
[0138] In an embodiment, the indication is a parameter set flag. The device may, for example, be configured to process a video data stream that includes a parameter set flag that indicates whether one or more scalable nested supplemental enhancement information messages containing timing information for each of one or more output layer sets are present in the video data stream.
[0139] In an embodiment, the video data stream of a series of parameter sets may include, for example, a parameter set flag.
[0140] According to an embodiment, the parameter setting flag is sps_ols_nest_timing_present_flag.
[0141] In an embodiment, the video data stream may include a further supplemental enhancement information message, which may include, for example, a parameter set flag. The parameter set flag indicates whether one or more scalable nested supplemental enhancement information messages for each of one or more output layer sets are present in the video data stream. The device may be configured to process the further supplemental enhancement information message, for example.
[0142] According to an embodiment, the timing information may include, for example, at least one of picture timing information, buffering period information, and decoding unit information.
[0143] According to an embodiment, if one or more scalable nested supplemental enhancement information messages include timing information that may include, for example, picture timing information for each of one or more output layer sets present in the video data stream, the device may be configured, for example, to substitute picture timing information for non-scalable nested picture timing supplemental enhancement information messages. If one or more scalable nested supplemental enhancement information messages include timing information that includes buffering period information for each of one or more output layer sets present in the video data stream, the device may be configured, for example, to substitute buffering period information for non-scalable nested buffering period supplemental enhancement information messages. If one or more scalable nested supplemental enhancement information messages include timing information that includes decoding unit information for each of one or more output layer sets present in the video data stream, the device may be configured, for example, to substitute decoding unit information for non-scalable nested decoding unit supplemental enhancement information messages.
[0144] In an embodiment, the timing information is timing information of a hypothetical reference decoder.
[0145] According to an embodiment, the device may be configured to, for example, decode the sub-bitstream to decode the video.
[0146] Further, according to an embodiment, there is provided a system for encoding video into a video data stream and for decoding the video. The system includes a video encoder as described above and an apparatus as described above. The video encoder may, for example, be configured to encode the video into a video data stream, such that the video data stream has the video encoded therein. The apparatus may, for example, be configured to receive the video data stream as an input bitstream. Further, the apparatus may, for example, be configured to process the input bitstream to obtain a sub-bitstream. Further, the apparatus may, for example, be configured to decode the sub-bitstream to decode the video.
[0147] The VVC draft specification includes a definition of OLS within a VPS, which can also be used for HRD-based conformance testing of OLS sub-bitstreams based on the respective HRD SEI messages (BP, PT, DUI) present in the bitstream in a nested format (scalable nested SEI message nesting). When defining an OLS, it is important to ensure that the HRD SEI messages for each of these OLS are present in the bitstream for conformance testing to be valid.
[0148] Therefore, it is part of the invention that an indication is present in the bitstream (e.g., a parameter set flag such as sps_ols_scal_nest_timing_present_flag in SPS or a new SEI message containing such a flag) indicating scalable nested SEI messages to all OLSs must be present in the bitstream.
[0149] The following describes how the HRD SEI is applied to the sub-bitstreams.
[0150] According to an embodiment, a video data stream having video encoded therein is provided, wherein an indication within the video data stream indicates whether timing information for a sub-bitstream is obtained from one or more non-scalable nested picture timing supplemental enhancement information messages of the video data stream.
[0151] In an embodiment, if the indication indicates that sub-bitstream timing information is not obtained from one or more non-scalable nested picture timing supplemental enhancement information messages, this indicates that one or more non-scalable nested picture timing supplemental enhancement information messages are substituted for one or more scalable nested picture timing supplemental enhancement information messages.
[0152] In an embodiment, the indication is a flag. One of the one or more scalable nested supplemental enhancement information messages may, for example, include a flag indicating whether sub-bitstream timing information may, for example, be configured to be obtained from one or more non-scalable nested picture timing supplemental enhancement information messages.
[0153] In an embodiment, the flag is use_orig_pic_timing_flag.
[0154] In an embodiment, the video data stream may include, for example, a video parameter set, the indication being a flag. The video parameter set may include, for example, a flag indicating whether timing information of the sub-bitstream may be configured to be obtained from one or more non-scalable nested picture timing supplemental enhancement information messages.
[0155] In an embodiment, the flag is use_orig_pic_timing_flag.
[0156] In an embodiment, the flag is general_same_pic_timing_in_all_ols_flag.
[0157] In an embodiment, the indication indicates that each of the one or more non-scalable nested picture timing supplemental enhancement information messages in each access unit of the one or more access units applies to an access unit of any output layer configured in the video data stream, and that either there are no scalable nested picture timing supplemental enhancement information messages, or that the non-scalable nested picture timing supplemental enhancement information messages in each access unit of the one or more access units may or may not apply to an access unit of any output layer configured in the video data stream, and that there may be scalable nested picture timing supplemental enhancement information messages.
[0158] In an embodiment, if the indication indicates that the timing information of the sub-bitstream may be configured to be obtained from one or more non-scalable nested picture timing supplemental enhancement information messages, at least one of the one or more scalable nested supplemental enhancement information messages occurs before the one or more non-scalable nested picture timing supplemental enhancement information messages in the access unit of the bitstream.
[0159] In an embodiment, the indication is a constraint flag. The one or more non-scalable nested picture timing supplemental enhancement information messages include a constraint flag that indicates whether the one or more non-scalable nested picture timing supplemental enhancement information messages apply to at least one of the one or more sub-bitstreams.
[0160] In an embodiment, the indication is a first indication, a second indication in the video data stream indicates whether the timing information of the sub-bitstream is further obtained from one or more non-scalable nested buffering period supplemental enhancement information messages of the video data stream, and / or a third indication in the video data stream indicates whether the timing information of the sub-bitstream is further obtained from one or more non-scalable nested decoding unit supplemental enhancement information messages of the video data stream.
[0161] According to an embodiment, the indication in the video data stream further indicates whether the timing information of the sub-bitstreams is further obtained from one or more non-scalable nested buffering period supplemental enhancement information messages of the video data stream, and / or the indication in the video data stream further indicates whether the timing information of the sub-bitstreams is further obtained from one or more non-scalable nested decoding unit supplemental enhancement information messages of the video data stream.
[0162] Further, according to an embodiment, there is provided a video data stream having video encoded therein, wherein a first indication in the video data stream indicates whether timing information for the sub-bitstreams is obtained from one or more non-scalable nested buffering period supplemental enhancement information messages of the video data stream, and / or a second indication in the video data stream indicates whether timing information for the sub-bitstreams is obtained from one or more non-scalable nested decoding unit supplemental enhancement information messages of the video data stream.
[0163] Furthermore, according to an embodiment, the video data stream includes one or more non-scalable nested supplemental enhancement information messages that include timing information. If the video data stream includes a scalable nested supplemental enhancement information message that includes timing information, this indicates that, depending on the scalable nested supplemental enhancement information message, all of the one or more non-scalable nested timing information extension information messages are substituted by the scalable nested supplemental enhancement information message that includes timing information (e.g., the timing information is picture timing information, buffering period information, or decoding unit information). Or, a subset including at least one of the one or more non-scalable nested timing information extension information messages is substituted by the scalable nested supplemental enhancement information message that includes timing information (e.g., the timing information is picture timing information, buffering period information, or decoding unit information).
[0164] According to embodiments, the one or more non-scalable nested supplemental enhancement information messages containing timing information are one or more non-scalable nested picture timing supplemental enhancement information messages, and the scalable nested supplemental enhancement information messages containing timing information are scalable nested picture timing supplemental enhancement information messages, or the one or more non-scalable nested supplemental enhancement information messages containing timing information are one or more non-scalable nested buffering period supplemental enhancement information messages, and the scalable nested supplemental enhancement information messages containing timing information are scalable nested buffering period supplemental enhancement information messages, or the one or more non-scalable nested supplemental enhancement information messages containing timing information are one or more non-scalable nested decoding unit supplemental enhancement information messages, and the scalable nested supplemental enhancement information messages containing timing information are scalable nested decoding unit supplemental enhancement information messages.
[0165] In an embodiment, if the video data stream may, for example, include the scalable nested supplemental enhancement information message containing the timing information, the scalable nested supplemental enhancement information message containing the timing information occurs before one or more non-scalable nested supplemental enhancement information messages containing timing information within an access unit of the bitstream.
[0166] According to an embodiment, the timing information is timing information of a hypothetical reference decoder.
[0167] In an embodiment, the sub-bitstreams are output layer set dependent and / or sub-layer dependent and / or sub-picture dependent and / or decoding unit subset dependent.
[0168] Further, according to an embodiment, there is provided a video encoder for encoding video into a video data stream, such that the video data stream has video encoded therein, wherein the video encoder generates the video data stream such that an indication in the video data stream indicates whether timing information of a sub-bitstream is obtained from one or more non-scalable nested picture timing supplemental enhancement information messages of the video data stream.
[0169] According to an embodiment, if the indication indicates that sub-bitstream timing information is not obtained from one or more non-scalable nested picture timing supplemental enhancement information messages, this indicates that one or more non-scalable nested picture timing supplemental enhancement information messages are substituted for one or more scalable nested picture timing supplemental enhancement information messages.
[0170] In an embodiment, the indication is a flag. The video encoder may, for example, be configured to generate the video data stream such that one of the one or more scalable nested supplemental enhancement information messages may, for example, include a flag indicating whether sub-bitstream timing information may, for example, be configured to be obtained from one or more non-scalable nested picture timing supplemental enhancement information messages.
[0171] In an embodiment, the flag is use_orig_pic_timing_flag.
[0172] In an embodiment, the video data stream may include, for example, a video parameter set. The indication may be, for example, a flag. The video encoder may be configured to generate the video data stream such that the video parameter set includes a flag indicating that timing information of the sub-bitstream may be configured to be obtained from one or more non-scalable nested picture timing supplemental enhancement information messages.
[0173] In an embodiment, the flag is use_orig_pic_timing_flag.
[0174] In an embodiment, the flag is general_same_pic_timing_in_all_ols_flag.
[0175] According to an embodiment, the indication indicates that each of the one or more non-scalable nested picture timing supplemental enhancement information messages in each access unit of the one or more access units applies to an access unit of any output layer configured in the video data stream, and that either there are no scalable nested picture timing supplemental enhancement information messages, or that the non-scalable nested picture timing supplemental enhancement information messages in each access unit of the one or more access units may or may not apply to an access unit of any output layer configured in the video data stream, and that a scalable nested picture timing supplemental enhancement information message may be present.
[0176] In an embodiment, if the indication indicates that the timing information of the sub-bitstream may be configured to be obtained from one or more non-scalable nested picture timing supplemental enhancement information messages, at least one of the one or more scalable nested supplemental enhancement information messages occurs before the one or more non-scalable nested picture timing supplemental enhancement information messages in the access unit of the bitstream.
[0177] According to an embodiment, the indication is a constraint flag. A video encoder may for example be configured to generate the video data stream such that one or more non-scalable nested picture timing supplemental enhancement information messages include a constraint flag that indicates whether the one or more non-scalable nested picture timing supplemental enhancement information messages apply to at least one of the one or more sub-bitstreams.
[0178] In an embodiment, the indication is a first indication, a second indication in the video data stream indicates whether the timing information of the sub-bitstream is further obtained from one or more non-scalable nested buffering period supplemental enhancement information messages of the video data stream, and / or a third indication in the video data stream indicates whether the timing information of the sub-bitstream is further obtained from one or more non-scalable nested decoding unit supplemental enhancement information messages of the video data stream.
[0179] According to an embodiment, the indication in the video data stream further indicates whether the timing information of the sub-bitstreams is further obtained from one or more non-scalable nested buffering period supplemental enhancement information messages of the video data stream, and / or the indication in the video data stream further indicates whether the timing information of the sub-bitstreams is further obtained from one or more non-scalable nested decoding unit supplemental enhancement information messages of the video data stream.
[0180] Further, according to an embodiment, there is provided a video encoder for encoding video into a video data stream, such that the video data stream has video encoded therein, wherein a first indication in the video data stream indicates whether timing information of sub-bitstreams is obtained from the one or more non-scalable nested buffering period supplemental enhancement information messages of the video data stream, and / or a second indication in the video data stream indicates whether timing information of sub-bitstreams is obtained from one or more non-scalable nested decoding unit supplemental enhancement information messages of the video data stream.
[0181] Further, according to an embodiment, a video encoder is provided for encoding video into a video data stream, such that the video data stream has video encoded therein. The video encoder generates a video data stream such that the video data stream includes one or more non-scalable nested supplemental enhancement information messages including timing information. If the video data stream includes a scalable nested supplemental enhancement information message including timing information, this indicates that, in response to the scalable nested supplemental enhancement information message, all of the one or more non-scalable nested timing information enhancement information messages are substituted by a scalable nested supplemental enhancement information message including timing information (e.g., the timing information is picture timing information, buffering period information, or decoding unit information). Alternatively, a subset including at least one of the one or more non-scalable nested timing information supplemental enhancement information messages is substituted by the non-scalable nested supplemental enhancement information message including the timing information (e.g., the timing information is picture timing information, buffering period information, or decoding unit information).
[0182] According to an embodiment, a video encoder may be configured, for example, to generate a video data stream such that one or more non-scalable nested supplemental enhancement information messages containing timing information are one or more non-scalable nested picture timing supplemental enhancement information messages and a scalable nested supplemental enhancement information message containing timing information is a scalable nested picture timing supplemental enhancement information message, or such that one or more non-scalable nested supplemental enhancement information messages containing timing information are one or more non-scalable nested buffering period supplemental enhancement information messages and a scalable nested supplemental enhancement information message containing timing information is a scalable nested buffering period supplemental enhancement information message. Or, the video encoder may be configured, for example, to generate the video data stream such that one or more non-scalable nested supplemental enhancement information messages containing timing information are one or more non-scalable nested decoding unit supplemental enhancement information messages, and the scalable nested supplemental enhancement information messages containing timing information are scalable nested decoding unit supplemental enhancement information messages.
[0183] In an embodiment, if the video data stream may, for example, include a scalable nested supplemental enhancement information message that includes timing information, the video encoder may, for example, be configured to generate the video data stream such that the scalable nested supplemental enhancement information message that includes timing information is generated before one or more non-scalable nested supplemental enhancement information messages that include timing information within an access unit of the bitstream.
[0184] According to an embodiment, the timing information is timing information of a hypothetical reference decoder.
[0185] In an embodiment, the sub-bitstreams are output layer set dependent and / or sub-layer dependent and / or sub-picture dependent and / or decoding unit subset dependent.
[0186] Further, according to an embodiment, there is provided an apparatus for receiving a video data stream as an input bitstream, the video data stream having video encoded therein, the apparatus being for processing the input bitstream to obtain sub-bitstreams, and an indication in the video data stream indicating whether timing information for the sub-bitstreams is obtained from one or more non-scalable nested picture timing supplemental enhancement information messages of the video data stream.
[0187] According to an embodiment, if the indication indicates that sub-bitstream timing information is not obtained from one or more non-scalable nested picture timing supplemental enhancement information messages, the device indicates that the one or more non-scalable nested picture timing supplemental enhancement information messages may, for example, be configured to be substituted for one or more scalable nested picture timing supplemental enhancement information messages.
[0188] According to an embodiment, the indication may be, for example, a constraint flag: one of the one or more scalable nested supplemental enhancement information messages may, for example, include a flag indicating that timing information of the sub-bitstream may, for example, be configured to be obtained from one or more non-scalable nested picture timing supplemental enhancement information messages.
[0189] In an embodiment, the flag is use_orig_pic_timing_flag.
[0190] In an embodiment, the video data stream may include, for example, a video parameter set. The indication may be, for example, a flag. The video parameter set may include, for example, a flag indicating whether timing information of the sub-bitstream may be configured to be obtained from one or more non-scalable nested picture timing supplemental enhancement information messages.
[0191] In an embodiment, the flag is same_pic_timing_within_ols_flag.
[0192] According to an embodiment, the flag is general_same_pic_timing_in_all_ols_flag.
[0193] In an embodiment, the indication indicates that each of the one or more non-scalable nested picture timing supplemental enhancement information messages in each access unit of the one or more access units applies to an access unit of any output layer configured in the video data stream, and that either there are no scalable nested picture timing supplemental enhancement information messages, or that the non-scalable nested picture timing supplemental enhancement information messages in each access unit of the one or more access units may or may not apply to an access unit of any output layer configured in the video data stream, and that there may be scalable nested picture timing supplemental enhancement information messages.
[0194] According to an embodiment, if the indication indicates that non-scalable nested picture timing supplemental enhancement information messages within each access unit of one or more access units may or may not be applied to access units of any output layer configured for the video data stream, and that scalable nested picture timing supplemental enhancement information messages may be present, the device is configured to remove from the input bitstream or sub-bitstream all supplemental enhancement information network abstraction layer units with picture timing content, including non-scalable nested supplemental enhancement information messages.
[0195] In an embodiment, if the indication indicates that the timing information of the sub-bitstream may, for example, be configured to be obtained from one or more non-scalable nested picture timing supplemental enhancement information messages, the device may, for example, be configured to process at least one of the one or more scalable nested supplemental enhancement information messages, which occurs before the device processes one or more non-scalable nested picture timing supplemental enhancement information messages in an access unit of the bitstream, for example, before the device may be configured to process one or more non-scalable nested picture timing supplemental enhancement information messages in the access unit.
[0196] In an embodiment, the indication is a constraint flag. The apparatus may, for example, be configured to process one or more non-scalable nested picture timing supplemental enhancement information messages that include a constraint flag indicating whether the one or more non-scalable nested picture timing supplemental enhancement information messages apply to at least one of the one or more sub-bitstreams.
[0197] In an embodiment, the indication is a first indication, a second indication in the video data stream indicates whether the timing information of the sub-bitstream is further obtained from one or more non-scalable nested buffering period supplemental enhancement information messages of the video data stream, and / or a third indication in the video data stream indicates whether the timing information of the sub-bitstream is further obtained from one or more non-scalable nested decoding unit supplemental enhancement information messages of the video data stream.
[0198] According to an embodiment, the indication in the video data stream further indicates whether the timing information of the sub-bitstreams is further obtained from one or more non-scalable nested buffering period supplemental enhancement information messages of the video data stream, and / or the indication in the video data stream further indicates whether the timing information of the sub-bitstreams is further obtained from one or more non-scalable nested decoding unit supplemental enhancement information messages of the video data stream.
[0199] Further, according to an embodiment, there is provided an apparatus for receiving a video data stream as an input bitstream, the video data stream having video encoded therein, the apparatus being configured to process the input bitstream to obtain sub-bitstreams, wherein a first indication in the video datastream indicates whether timing information for the sub-bitstreams is obtained from one or more non-scalable nested buffering period supplemental enhancement information messages of the video datastream, and / or a second indication in the video datastream indicates whether timing information for the sub-bitstreams is obtained from one or more non-scalable nested decoding unit supplemental enhancement information messages of the video datastream.
[0200] Further, according to an embodiment, there is provided an apparatus for receiving a video data stream as an input bitstream, the video data stream having video encoded therein. The apparatus processes the input bitstream to obtain sub-bitstreams. The video data stream includes one or more non-scalable nested supplemental enhancement information messages including timing information. If the video data stream includes a scalable nested supplemental enhancement information message including timing information, the apparatus, in response to the scalable nested supplemental enhancement information message, adjusts all of the one or more non-scalable nested timing information enhancement information messages to correspond to the scalable nested supplemental enhancement information message including timing information (e.g., the timing information is picture timing information, buffering period information, or decoding unit information). Or, a subset including at least one of the one or more non-scalable nested timing information supplemental enhancement information messages to correspond to the scalable nested supplemental enhancement information message including timing information (e.g., the timing information is picture timing information, buffering period information, or decoding unit information).
[0201] According to embodiments, the one or more non-scalable nested supplemental enhancement information messages containing timing information are one or more non-scalable nested picture timing supplemental enhancement information messages, and the scalable nested supplemental enhancement information messages containing timing information are scalable nested picture timing supplemental enhancement information messages, or the one or more non-scalable nested supplemental enhancement information messages containing timing information are one or more non-scalable nested buffering period supplemental enhancement information messages, and the scalable nested supplemental enhancement information messages containing timing information are scalable nested buffering period supplemental enhancement information messages, or the one or more non-scalable nested supplemental enhancement information messages containing timing information are one or more non-scalable nested decoding unit supplemental enhancement information messages, and the scalable nested supplemental enhancement information messages containing timing information are scalable nested decoding unit supplemental enhancement information messages.
[0202] In an embodiment, a video data stream may, for example, include a scalable nested supplemental enhancement information message that includes timing information, where the scalable nested supplemental enhancement information message that includes timing information occurs before one or more non-scalable nested supplemental enhancement information messages that include timing information within an access unit of the bitstream, and an apparatus may, for example, be configured to process a scalable nested supplemental enhancement information message that includes timing information before one or more non-scalable nested supplemental enhancement information messages that include timing information.
[0203] According to an embodiment, the timing information is timing information of a hypothetical reference decoder.
[0204] In an embodiment, the sub-bitstreams are output layer set dependent and / or sub-layer dependent and / or sub-picture dependent and / or decoding unit subset dependent.
[0205] According to an embodiment, the device may be configured to, for example, decode the sub-bitstream to decode the video.
[0206] Further, according to an embodiment, there is provided a system for encoding video into a video data stream and for decoding the video. The system includes a video encoder as described above and an apparatus as described above. The video encoder may, for example, be configured to encode the video into a video data stream, such that the video data stream has the video encoded therein. The apparatus may, for example, be configured to receive the video data stream as an input bitstream. Further, the apparatus may, for example, be configured to process the input bitstream to obtain a sub-bitstream. Further, the apparatus may, for example, be configured to decode the sub-bitstream to decode the video.
[0207] The VVC draft specification includes SEI messages that control HRD timing operations, namely, the buffering period SEI message and the picture timing SEI message when the bitstream is decoded.
[0208] Currently, the VVC draft specification already includes all Picture Timing SEI messages that apply to some target Temporal IDs. This means that temporal scalability does not necessarily require scalable nesting of additional timing information. Therefore, when extracting sub-bitstreams (e.g., in a layering scenario via OLS or spatially by extracting sub-pictures), in some cases it may not be necessary to modify / exchange Picture Timing SEI messages in such a way.
[0209] In the following, the output layer set abstraction is described.
[0210] In one embodiment, signaling is added to indicate that Picture Timing SEI messages in a bitstream apply to any sub-bitstream (defined / corresponding to some Output Layer Set (OLS) A) of the bitstream (defined / corresponding to some OLS B), and only SEI messages with Buffering periods are replaced with scalably nested pairs. Example syntax: [Table 1]
[0211] Currently, scalable nested SEI messages follow HRD SEI messages in an access unit. Therefore, as part of the above embodiment, if use_orig_pic_timing_flag is equal to 1, scalable nested SEI messages, including OLS-specific HRD SEI messages, must occur before their respective PT SEI messages (messages to keep during extraction) in an access unit in bitstream order.
[0212] Or, in an alternative embodiment, the indication is in the VPS as a constraint flag: [Table 2]
[0213] Or, in an alternative embodiment, the indication is in the Picture Timing SEI as a constraint flag.
[0214] This extraction process has been modified.
[0215] In the following, DPS rewrite during the sub-bitstream extraction process is described.
[0216] The inputs to this process are the bitstream inBitstream, the target OLS index targetOlsIdx, and the target's highest TemporalId value tIdTarget.
[0217] The output of this process is the sub-bitstream outBitstream.
[0218] Any output sub-bitstream that is the output of a process specified in this clause with targetOlsIdx equal to an index in the list of OLSs specified in the VPS as the bitstream, and tIdTarget equal to any value in the range 0 to 6, inclusive, as the input, and that satisfies the following conditions, shall be a conforming bitstream: The output sub-bitstream contains at least one VCL NAL unit, where nuh_layer_id is equal to each of LayerIdInOls[targetOlsIdx]. The output sub-bitstream contains at least one VCL NAL unit whose TemporalId is equal to tIdTarget. NOTE - A compliant bitstream contains one or more coded slice NAL units with TemporalId equal to 0, but is not required to contain any coded slice NAL units with nuh_layer_id equal to 0. The output sub-bitstream OutBitstream is derived as follows. The bitstream outBitstream is set to be the same as the bitstream inBitstream. Remove all NAL units with TemporalId greater than tIdTarget from outBitstream. Remove all NAL units from outBitstream whose nal_unit_type is not equal to any of VPS_NUT, DPS_NUT, and EOB_NUT and whose nuh_layer_id is not included in the list LayerIdInOls[targetOlsIdx]. Remove from outBitstream all SEI NAL units containing scalable nesting SEI messages where nesting_ols_flag is 1 and there are no values of i in the range from 0 to nesting_num_olss_minus1, inclusive, e.g., NestingOlsIdx[i] is the same as targetOlsIdx. If targetOlsIdx is greater than 0, remove all SEI NAL units from outBitstream that contain non-scalable nested SEI messages with payloadType equal to 0 (buffering period), 1 (picture timing), or 130 (decoding unit information). ·If targetOlsIdx is greater than 0 and use_orig_pic_timing_flag / same_pic_timing_within_ols_flag is equal to 0, remove all SEI NAL units containing non-scalable nested SEI messages whose payloadType is equal to 1 (picture timing) from outBitstream. If outBitstream contains a scalable nesting SEI message and outBitstream contains applicable SEI NAL units (NestingOlsIdx[i] equals targetOlsIdx), do the following: Extract the appropriate non-scalable nested SEI messages from the scalable nested SEI message with payloadType equal to 0 (buffering period), 1 (picture timing), or 130 (decoding unit information) and put those messages into outBitstream. Otherwise (when use_orig_pic_timing_flag / same_pic_timing_within_ols_flag is equal to 1), extract appropriate non-scalable nested SEI messages with payloadType equal to 0 (buffering period) or 130 (decoding unit information) from the scalable nested SEI message if one exists that is applicable to outBitstream (NestingOlsIdx[i] is equal to targetOlsIdx) and put those messages into outBitstream. Remove all SEI NAL units, including scalable nested SEI messages, from outBitstream.
[0219] The indication, for example a flag, may be, for example, general_same_pic_timing_in_all_ols_flag, or, for example, same_pic_timing_within_ols_flag.
[0220] When general_same_pic_timing_in_all_ols_flag (or same_pic_timing_within_ols_flag) is equal to its first value, e.g., 1, it specifies that the non-scalable nested PT SEI messages for each AU apply to all OLS AUs in the bitstream and there are no scalable nested PT SEI messages. When general_same_pic_timing_in_all_ols_flag is equal to its second value, e.g., 0, it specifies that the non-scalable nested PT SEI messages for each AU may or may not apply to any OLS AUs in the bitstream and there may be scalable nested PT SEI messages.
[0221] For example, if general_same_pic_timing_in_all_ols_flag is equal to the second value, e.g., 0, the device removes from the input bitstream or output bitstream (e.g., outBitstream / sub-bitstream) all SEI NAL units that constitute non-scalable nested SEI messages with payloadType equal to 1 (PT).
[0222] In an embodiment, the buffering period SEI message may also not require replacement by a scalable nested variant in some cases, so that an additional indication can be carried in the parameter set or buffering period SEI message itself, and the indicated timing is also applied to the extracted sub-bitstream, and the extraction process is further modified to retain the original BP and PT SEI messages according to the respective indication.
[0223] In the following, the extraction of sub-pictures is explained.
[0224] In another embodiment that considers the case of sub-picture sub-bitstream extraction, the respective signals are added to the bitstream (e.g., in the syntax of a sub-picture nested SEI message, a Picture Timing SEI message, or a parameter set), and the Picture Timing SEI message in the bitstream applies to any sub-bitstream defined by a sub-picture or sub-picture set of the bitstream formed by the combination of all sub-pictures contained in the bitstream, with only the buffering period SEI message being replaced by its sub-picture nested counterpart.
[0225] Alternatively, the buffering period SEI message may also in some cases not require replacement by sub-picture nested variants, so that an additional indication can be carried in the parameter set or buffering period SEI message itself, and the indicated timing will also be applied to the extracted sub-bitstreams, with the extraction process being further modified to preserve the original BP and PT SEI messages according to the respective indication.
[0226] The following describes the removal of non-scalable nested HRD SEI messages based on the presence of scalable nested HRD SEI messages.
[0227] In another embodiment of the present invention, the scope of non-scalable nested HRD SEI messages (e.g., BP, PT, and DUI SEI messages), i.e., the sub-bitstream (e.g., OLS) from which some of them (e.g., only PT SEI messages) or all of them can be extracted, is indicated by the presence of a respective substitute SEI message in the form of a scalable nested HRD SEI message, and the absence of such a message indicates that the respective non-scalable HRD SEI message applies to the OLS. A consequence of this scope indication is that the removal of non-scalable nested SEI messages in the sub-bitstream extraction process depends on the presence of a scalable nested SEI message that can serve as a substitute for the removed message; in the absence of such a scalable nested SEI message, the non-scalable nested SEI message remains in the extracted sub-bitstream.
[0228] In a further embodiment, each scalable nested HRD SEI message is placed before the non-scalable nested HRD SEI message in bitstream order in the access unit to facilitate sequential processing of the bitstream during extraction.
[0229] Below we describe a simplified, scalable nest.
[0230] According to an embodiment, a video data stream having video encoded therein is provided, the video data stream including a plurality of access units, and for each access unit of the plurality of access units, if the access unit includes two or more scalable nested supplemental enhancement information messages that include buffer period information and / or picture timing information of an output layer set, the buffer period information and / or picture timing information of the output layer set are equal in all of the two or more scalable nested supplemental enhancement information messages of the access unit.
[0231] Further, according to an embodiment, there is provided a video data stream having video encoded therein, the video data stream including a plurality of access units, and for each access unit of the plurality of access units, if the access unit includes two or more scalable nested supplemental enhancement information messages that include buffer period information and / or picture timing information of an output layer set, the video data stream including an indication of whether the buffer period information and / or picture timing information of the output layer set are equal in all two or more scalable nested supplemental enhancement information messages of the access unit.
[0232] Further, according to an embodiment, there is provided a video data stream having video encoded therein. The video data stream includes a plurality of access units. For each access unit of the plurality of access units, if the access unit includes three or more scalable nested supplemental enhancement information messages that include buffer period information and / or picture timing information of an output layer set, the buffer period information and / or picture timing information of the output layer set appears only in one of the three or more scalable nested supplemental enhancement information messages and in another of the three or more scalable nested supplemental enhancement information messages that immediately follows said one of the three or more scalable nested supplemental enhancement information messages.
[0233] Further, according to an embodiment, there is provided a video data stream having video encoded therein. The video data stream includes a plurality of access units. For each access unit of the plurality of access units, if the access unit includes three or more scalable nested supplemental enhancement information messages that include buffer period information and / or picture timing information for an output layer set, the video data stream includes an indication of whether the buffer period information and / or picture timing information for the output layer set appears only in one of the three or more scalable nested supplemental enhancement information messages and in another of the three or more scalable nested supplemental enhancement information messages that immediately follows the one of the three or more scalable nested supplemental enhancement information messages.
[0234] In an embodiment, all scalable nested supplemental enhancement information messages within an access unit whose identifiers have the same value and identify an output layer set may, for example, carry the same buffer period information and / or the same picture timing information.
[0235] According to an embodiment, for a particular access unit of a plurality of access units, any picture timing supplemental enhancement information messages that apply to a set of layers and sub-layers of the output layer set may, for example, carry the same picture timing information, and / or for a particular access unit of a plurality of access units, any buffer periodicity supplemental enhancement information messages that apply to a set of layers and sub-layers of the output layer set may, for example, carry the same buffer periodicity information, and / or for a particular access unit of a plurality of access units, any decoding unit supplemental enhancement information messages that apply to a set of layers and sub-layers of the output layer set may, for example, carry the same decoding unit information.
[0236] In an embodiment, two scalable nested supplemental enhancement information messages of a particular payload type within an access unit with the same identifier value and identifying an output tier set may, for example, carry the same content.
[0237] Further, according to an embodiment, there is provided a video encoder for encoding video into a video data stream, such that the video data stream has video encoded therein. The video encoder generates the video data stream such that the video data stream includes a plurality of access units, for each access unit of the plurality of access units. If an access unit includes two or more scalable nested supplemental enhancement information messages that include buffer period information and / or picture timing information of an output layer set, the video encoder generates the video data stream such that the buffer period information and / or picture timing information of the output layer set are equal in all of the two or more scalable nested supplemental enhancement information messages of the access unit.
[0238] Further, according to an embodiment, there is provided a video encoder for encoding video into a video data stream, such that the video data stream has video encoded therein. The video encoder generates a video data stream such that the video data stream includes a plurality of access units. For each access unit of the plurality of access units, if the access unit includes two or more scalable nested supplemental enhancement information messages that include buffer period information and / or picture timing information of an output layer set, the video encoder generates the video data stream such that the video data stream includes an indication of whether the buffer period information and / or picture timing information of the output layer set are equal in all two or more scalable nested supplemental enhancement information messages of the access unit.
[0239] Further, according to an embodiment, there is provided a video encoder for encoding video into a video data stream, such that the video data stream has video encoded therein. The video encoder generates the video data stream such that the video data stream includes a plurality of access units. For each access unit of the plurality of access units, if the access unit includes three or more scalable nested supplemental enhancement information messages that include buffer period information and / or picture timing information of an output layer set, the video encoder generates the video data stream such that the buffer period information and / or picture timing information of the output layer set appears only in one of the three or more scalable nested supplemental enhancement information messages and in another of the three or more scalable nested supplemental enhancement information messages that immediately follows the one of the three or more scalable nested supplemental enhancement information messages.
[0240] Further, according to an embodiment, there is provided a video encoder for encoding video into a video data stream, such that the video data stream has video encoded therein. The video encoder generates a video data stream such that the video data stream includes a plurality of access units. For each access unit of the plurality of access units, if the access unit includes three or more scalable nested supplemental enhancement information messages that include buffer period information and / or picture timing information of an output layer set, the video encoder generates the video data stream such that the video data stream includes an indication of whether the buffer period information and / or picture timing information of the output layer set appears only in one of the three or more scalable nested supplemental enhancement information messages and in another of the three or more scalable nested supplemental enhancement information messages that immediately follows the one of the three or more scalable nested supplemental enhancement information messages.
[0241] In an embodiment, all scalable nested supplemental enhancement information messages within an access unit whose identifiers have the same value and identify an output layer set may, for example, carry the same buffer period information and / or the same picture timing information.
[0242] According to an embodiment, for a particular access unit of a plurality of access units, any picture timing supplemental enhancement information messages that apply to a set of layers and sub-layers of the output layer set may, for example, carry the same picture timing information, and / or for a particular access unit of a plurality of access units, any buffer periodicity supplemental enhancement information messages that apply to a set of layers and sub-layers of the output layer set may, for example, carry the same buffer periodicity information, and / or for a particular access unit of a plurality of access units, any decoding unit supplemental enhancement information messages that apply to a set of layers and sub-layers of the output layer set may, for example, carry the same decoding unit information.
[0243] In an embodiment, two scalable nested supplemental enhancement information messages of a particular payload type within an access unit with the same identifier value and identifying an output tier set may, for example, carry the same content.
[0244] Further, according to an embodiment, there is provided an apparatus for receiving a video data stream as an input bitstream, the video data stream having video encoded therein. The input bitstream includes a plurality of access units. The apparatus processes the access units of the input bitstream to obtain sub-bitstreams. For each access unit of the plurality of access units, if the access unit includes two or more scalable nested supplemental enhancement information messages that include buffer period information and / or picture timing information of an output layer set, the buffer period information and / or picture timing information of the output layer set are equal in all of the two or more scalable nested supplemental enhancement information messages of the access unit.
[0245] Further, according to an embodiment, there is provided an apparatus for receiving a video data stream as an input bitstream, the video data stream having video encoded therein. The input bitstream includes a plurality of access units. The apparatus processes the access units of the input bitstream to obtain sub-bitstreams. For each access unit of the plurality of access units, if the access unit includes two or more scalable nested supplemental enhancement information messages including buffer period information and / or picture timing information of an output layer set, the video data stream includes an indication of whether the buffer period information and / or picture timing information of the output layer set are equal in all of the two or more scalable nested supplemental enhancement information messages of the access unit.
[0246] Further, according to an embodiment, there is provided an apparatus for receiving a video data stream as an input bitstream, the video data stream having video encoded therein. The input bitstream includes a plurality of access units, and the apparatus processes the access units of the input bitstream to obtain sub-bitstreams. For each access unit of the plurality of access units, if the access unit includes three or more scalable nested supplemental enhancement information messages that include buffer period information and / or picture timing information of an output layer set, the buffer period information and / or picture timing information of the output layer set appears only in one of the three or more scalable nested supplemental enhancement information messages and in another of the three or more scalable nested supplemental enhancement information messages that immediately follows said one of the three or more scalable nested supplemental enhancement information messages.
[0247] Further according to an embodiment, there is provided an apparatus for receiving a video data stream as an input bitstream, the video data stream having video encoded therein. The input bitstream includes a plurality of access units. The apparatus processes the access units of the input bitstream to obtain sub-bitstreams. For each access unit of the plurality of access units, if the access unit includes three or more scalable nested supplemental enhancement information messages including buffer period information and / or picture timing information of an output layer set, the video datastream includes an indication of whether the buffer period information and / or picture timing information of the output layer set appears only in one of the three or more scalable nested supplemental enhancement information messages and in another one of the three or more scalable nested supplemental enhancement information messages that immediately follows the one of the three or more scalable nested supplemental enhancement information messages.
[0248] In an embodiment, all scalable nested supplemental enhancement information messages within an access unit whose identifiers have the same value and identify an output layer set may, for example, carry the same buffer period information and / or the same picture timing information.
[0249] According to an embodiment, for a particular access unit of a plurality of access units, any picture timing supplemental enhancement information messages that apply to a set of layers and sub-layers of the output layer set may, for example, carry the same picture timing information, and / or for a particular access unit of a plurality of access units, any buffer periodicity supplemental enhancement information messages that apply to a set of layers and sub-layers of the output layer set may, for example, carry the same buffer periodicity information, and / or for a particular access unit of a plurality of access units, any decoding unit supplemental enhancement information messages that apply to a set of layers and sub-layers of the output layer set may, for example, carry the same decoding unit information.
[0250] In an embodiment, two scalable nested supplemental enhancement information messages of a particular payload type within an access unit with the same identifier value and identifying an output tier set may, for example, carry the same content.
[0251] According to an embodiment, if a device finds a buffer period scalable nested supplemental enhancement information message for an output layer set, the device may for example be configured to use the content of the buffer period scalable nested supplemental enhancement information for the output layer set without searching for further buffer period scalable nested supplemental enhancement information messages for the output layer set, and / or if a device finds a picture timing scalable nested supplemental enhancement information message for an output layer set, the device may for example be configured to use the content of the picture timing scalable nested supplemental enhancement information for the output layer set without searching for further picture timing scalable nested supplemental enhancement information messages for the output layer set.
[0252] According to an embodiment, the device may be configured to, for example, decode the sub-bitstream to decode the video.
[0253] Further, according to an embodiment, there is provided a system for encoding video into a video data stream, for decoding video. The system includes a video encoder as described above and an apparatus as described above. The video encoder may, for example, be configured to encode video into a video data stream, such that the video data stream has the video encoded therein. The apparatus may, for example, be configured to receive the video data stream as an input bitstream. Further, the apparatus may, for example, be configured to process the input bitstream to obtain a sub-bitstream. Further, the apparatus may, for example, be configured to decode the sub-bitstream to decode the video.
[0254] If scalable nested SEI messages can be used to replace the BP and PT SEI messages for all OLSs in a bitstream, the current state-of-the-art allows an encoder to distribute BPs and PTs across multiple scalable nesting SEI messages. An extractor processing such a bitstream must scan all scalable nested SEI messages in an access unit (which may be large given the amount of OLSs and repetitions) until it finds the appropriate scalable nested BP and P SEI messages.
[0255] If the encoder places the first scalably nested BP or PTSEI message into an access unit of the bitstream and, in the process of encoding the images of that access unit, finds that the BP or PTSEI message parameters need to be updated, writing another BP or PTSEI message into that access unit of the bitstream will in turn burden the extractor with the need to ensure that it uses the latest or most recent BP or PT SEI message placed in the access unit in the bitstream when abstracting that access unit.
[0256] The present invention simplifies the operation of the extractor by imposing constraints that remove the above nonsensical options.
[0257] In an embodiment, for example, it may be a requirement for bitstream compatibility that all scalable SEI messages (i.e., identifiers / indexes identifying OLSs) within an access unit with the same nesting value carry the same content (e.g., the same buffering period information and / or the same picture / timing information). This allows the BP and PT SEI messages of an OLS to be found in one scalable nesting SEI message, ensuring that an extractor obtains all the necessary information once it finds the scalable nesting SEI message in the target OLS.
[0258] In embodiments, it may be a bitstream compatibility requirement that all scalable nested SEI messages within an access unit that have the same value of an identifier, e.g., an index, identifying the OLS, carry the same content (e.g., the same buffering period information and / or the same picture / timing information). For example, for a particular access unit, any PT SEI messages (e.g., for a particular OLS) that apply to a set of layers and sublayers must have the same payload (e.g., for two PT SEI messages contained in two separate scalable nested SEI messages). The same applies to BPSEI messages and DUI SEI messages. This allows the BP and PT SEI messages of an OLS to be found in one scalable nesting SEI message, ensuring that an extractor obtains all the necessary information once it finds the scalable nesting SEI message in the target OLS.
[0259] In embodiments, it may be a bitstream compatibility requirement that all scalable nested SEI messages within an access unit having the same value of an identifier, e.g., an index, OLS identification, carry the same content (e.g., the same buffering period information and / or the same picture / timing information). For example, for a particular access unit in a bitstream that carries two scalably nested PT SEI messages that apply to the same OLS, the payloads of these two scalably nested PT SEI messages are equal, and the extractor can be sure that when it finds the scalable nested SEI message in the target OLS, it will have all the information it needs.
[0260] In an embodiment, it may be a bitstream compatibility requirement that all scalable nested SEI messages within an access unit that have the same value of an identifier, e.g., an index, identifying the OLS, carry the same content (e.g., the same buffering period information and / or the same picture / timing information). For example, the BP and PT SEI messages of an OLS may be found in one scalable nesting SEI message, and an extractor can be sure to obtain all the necessary information once it finds the scalable nesting SEI message in the target OLS.
[0261] According to an embodiment, for example, it may be a bitstream conformance requirement that BP and PT SEI messages applicable to OLS are not applicable to OLS without any other scalable nesting SEI message NAL units falling within the two scalable nesting SEI messages.
[0262] In the following, we will discuss low latency and temporal scalability of DU timing.
[0263] According to an embodiment, there is provided a video data stream having video encoded therein, the video data stream including a plurality of access units, the video data stream including a spreading factor that depends on the number of sub-bitstreams of the video data stream, or the video data stream including a clock sub-tick value that depends on the highest sub-bitstream among the sub-bitstreams of the video data stream.
[0264] According to an embodiment, each of the sub-bitstreams is output layer set dependent and / or sub-layer dependent and / or sub-picture dependent.
[0265] In an embodiment, the decoding unit removal time depends on the access unit removal time and the spreading factor.
[0266] According to an embodiment, the video data stream includes a temporal distance, the temporal distance being multiplied by a derived clock subtick value, the derived clock subtick value being derivable using the spreading factor.
[0267] In an embodiment, each of the plurality of spreading factors is assigned to a sub-layer of a plurality of sub-layers. A video data stream may, for example, include a plurality of access units.
[0268] The clock subtick value depends on the clock tick, which in turn depends on the spread.
[0269] The clock subtick value is defined as follows: ClockSubTick= = ClockTick÷( tick_divisor_minus2+2 ) * ( tick_divisor_factor_minus1[HTid]+ 1 ) ClockSubTick is the clock subtick value, Clock Tick is the clock tick, tick_divisor_minus2 is the additional tick divisor, and tick_divisor_factor_minus1[HTidU+00A0] is the timestamp of the video data stream.
[0270] Further, according to an embodiment, there is provided a video encoder for encoding video into a video data stream, such that the video data stream has video encoded therein. The video encoder generates the video data stream, such that the video data stream includes a plurality of access units. Further, the video encoder generates the video data stream, such that the video data stream includes a spreading factor that depends on the number of sub-bitstreams of the video data stream, or the video encoder generates the video data stream, such that the video data stream includes a clock sub-tick value that depends on the highest sub-bitstream among the sub-bitstreams of the video data stream.
[0271] According to an embodiment, each of the sub-bitstreams is output layer set dependent and / or sub-layer dependent and / or sub-picture dependent.
[0272] In embodiments, a video encoder may for example be configured to generate said video data stream such that a decoding unit and removal time depends on the removal time of an access unit and said spreading factor.
[0273] According to an embodiment, the video data stream may, for example, include a temporal distance, which is multiplied by a derived clock subtick value, which is derivable using the spreading factor.
[0274] In an embodiment, each of the plurality of spreading factors is assigned to a sub-layer of the plurality of sub-layers. A video encoder may be configured, for example, to generate a video data stream such that the video data stream may include, for example, the plurality of spreading factors.
[0275] The clock subtick value depends on the clock tick, which in turn depends on the spread.
[0276] The clock subtick value is defined as follows: ClockSubTick= = ClockTick÷( tick_divisor_minus2+2 ) * ( tick_divisor_factor_minus1[HTid]+ 1 ) ClockSubTick is the clock subtick value, Clock Tick is the clock tick, tick_divisor_minus2 is the additional tick divisor, and tick_divisor_factor_minus1[HTidU+00A0] is the video data stream.
[0277] Further, according to an embodiment, there is provided a video encoder for encoding video into a video data stream, such that the video data stream has video encoded therein, wherein the video encoder generates the video data stream such that the video data stream includes sub-layer specific frame rate information for the sub-layer, and / or the video encoder generates the video data stream such that the video data stream includes sub-layer specific frame display time information for the sub-layer.
[0278] Further, according to an embodiment, there is provided a video decoder for receiving a video data stream as an input bitstream, the video data stream having video encoded therein, the video data stream including a plurality of access units, the video decoder being configured to decode the video data stream to decode the video, the video data stream including a spreading factor that depends on a number of sub-bitstreams of the video data stream, the video decoder being configured to decode the video using the spreading factor, or the video data stream including a clock sub-tick value that depends on a highest sub-bitstream of the video data stream, the video decoder being configured to decode the video using the clock sub-tick value.
[0279] According to an embodiment, each of the sub-bitstreams is output layer set dependent and / or sub-layer dependent and / or sub-picture dependent.
[0280] 144. A video decoder according to claim 142 or 143, in an embodiment wherein the decoding unit removal time depends on the access unit removal time and the spreading factor.
[0281] In an embodiment, the video data stream may, for example, include a temporal distance. The video decoder may, for example, be configured to multiply the temporal distance by a derived clock sub-tick value. The video decoder may, for example, be configured to derive the derived clock sub-tick value using a spreading factor.
[0282] In an embodiment, the spreading factor is one of a plurality of spreading factors. Each of the plurality of spreading factors may, for example, be assigned to a sub-layer of a plurality of sub-layers. The video data stream may, for example, include a plurality of access units. The video decoder may, for example, be configured to decode the video using the plurality of spreading factors.
[0283] In an embodiment, the clock subtick value depends on the clock ticks, which in turn depends on the spreading factor.
[0284] In an embodiment, the clock subtick value is defined as follows: ClockSubTick= = ClockTick÷( tick_divisor_minus2+2 ) * ( tick_divisor_factor_minus1[HTid]+ 1 ) ClockSubTick is the clock subtick value, Clock Tick is the clock tick, tick_divisor_minus2 is the additional tick divisor, and tick_divisor_factor_minus1[HTidU+00A0] is the timestamp of the video data stream.
[0285] Further, according to one embodiment, there is provided a method for receiving a video data stream as an input bitstream, the video data stream having video encoded therein. A video decoder is configured to decode the video data stream to decode the video. The video data stream includes sub-layer-specific frame rate information for a sub-layer and / or the video data stream includes sub-layer-specific frame display period information for a sub-layer. The decoder is configured to determine a spreading factor using the sub-layer-specific frame rate information for the sub-layer and / or using the sub-layer-specific frame display period information.
[0286] Further, according to an embodiment, there is provided a system for encoding video into a video data stream and for decoding video. The system includes a video encoder as described above and an apparatus as described above. The video encoder encodes video into a video data stream such that the video data stream has the video encoded therein. The video decoder receives the video data stream and decodes the video data stream into video.
[0287] Removing a temporal sub-layer modifies the removal times of the remaining access units of the lower temporal sub-layers, especially in the presence of frame reordering. However, in low-latency configurations, especially when DU timing is provided, the removal times are modified in a very structured way, i.e., the resulting decode times are modified to scale or spread over time. Figure 4 illustrates the problem when there are three DUs per access unit.
[0288] Note that the final decoding time of an AU when the entire bitstream is decoded will be different from that of only its substream, thus losing the ultra-low latency property. Typically, this is the result of different decoding capabilities, i.e., a decoder capable of decoding 60 fps can process a frame in 1 / 60th of a second, while a decoder capable of decoding 30 fps can only process a frame in 1 / 30th of a second.
[0289] Note that the removal time of a DU is indicated as a delta of the removal time of an AU. deltaTime is signaled in the Picture Timing SEI message or the Decoding Unit Information SEI message in number of clock subticks. This deltaTime indicates the removal time of the DU compared to the removal time of the AU. Instead of looping over multiple temporal sublayers to indicate the DU removal time (as deltaTimes compared to the AU removal time) from the CPB, a spreading factor can be derived. In one embodiment, the spreading factor is derived from sublayer-specific frame rate information (e.g., sublayer T0 30 fps, sublayer L1 60 fps) or sublayer-specific frame display period (FrameTimeInterval described in Section 6 below, e.g., sublayer L01 / 30 s and sublayer L11 / 60 s), for example, according to the ratio of such information of two such sublayers. Alternatively, such a spreading factor can be indicated. This embodiment adds information of the DU removal time indicating the spreading factor, and calculates the corresponding DU time as a delta to the AU removal time.
[0290] Or as shown in one of the following: [Table 3]
[0291] Or, in an alternative embodiment, do the following: [Table 4]
[0292] The derivation of the relevant variable ClockSubTick in the HRD specification is changed as follows: The variable ClockSubTick is derived as follows and is called Clock Subtick: lockSubTick=ClockTick÷( tick_divisor_minus2+2 ) * ( tick_divisor_factor_minus1[HTid]+ 1 ) (C-2)
[0293] DeltaTime then uses a larger or smaller ClockSubTick depending on the highest sub-layer present in the bitstream.
[0294] Alternatively, the syntax element tick_divisor_minus2 can be replicated for each sub-layer to indicate the correct tick divisor when the highest temporal sub-layer (HTid) is set equal to the respective sub-layer. In that case, there is no spreading factor, but multiple ClockSubTicks are signaled, each sent for a different value of the highest sub-layer present in the bitstream, so that if a sub-layer is removed from the original bitstream, the ClockSubTick used will be different.
[0295] Below we explain the relevance of OLS extraction when using common DU timing.
[0296] If the bitstream contains extractable sub-streams specific to OLS and common DU timing (non-scalable nested PT SEI messages containing common DU timing), it is preferable, in terms of bitrate overhead and processing complexity, to reuse these PT SEI messages in the form of comparisons against the above indications (use_orig_pic_timing_flag and same_pic_timing_within_ols_flag) instead of providing scalable nested PT SEI messages instead.
[0297] Figure 5 shows two layers shown on top and the removal times of AUs and DUs for both layers shown below for the complete bitstream (e.g., 0th OLS) and for a single layer extracted with layer ID L0 (e.g., 1st OLS).
[0298] In this case, the common timing needs to be scaled in the same way as above by tick_divisor_factor_minus1[] or the absolute representation of the adjusted tick divisor (see DU time spread in the figure). Thus, in one embodiment, each OLS-specific HRD parameter syntax structure in a VPS carries either the relative factor or the absolute value of the tick divisor to scale the common DU removal timing.
[0299] However, in this usage scenario, the number of DUs per AU changes due to the extraction because the AU contains both L0 and L1 images, which means that aspects described below in aspect 6 are also required to derive the correct number of DUs remaining after extraction.
[0300] The following describes the relevance of the use of common DU timing to sub-picture extraction.
[0301] If the bitstream contains extractable sub-streams specific to subpictures (subpicture boundary processing for motion compensated prediction is enabled) and common DU timing (non-scalable nested PT SEI messages containing common timing), it is equally desirable, from the standpoint of bitrate overhead and processing complexity, to reuse these SEI messages in the form of comparisons against the above indications (use_orig_pic_timing_flag and same_pic_timing_within_ols_flag) instead of providing scalable nested PT SEI messages instead.
[0302] In this case, the common timing needs to be scaled in the same way as above by tick_divisor_factor_minus1[] or the absolute representation of the adjusted tick divisor to derive the correct number of DUs remaining after extraction (see DU time spread in the figure). Therefore, in one embodiment, the sub-picture specific HRD parameters syntax structure in the VPS / SPS carries either the relative factor or the absolute value of the tick divisor to scale the common DU removal timing.
[0303] Alternatively, an SEI message (eg, a sub-picture level information SEI message) may convey scaling information or absolute values so that the HRD parameters in the extracted sub-bitstream can be properly derived.
[0304] The following describes how to derive the CPB / bitrate size of a subpicture.
[0305] According to an embodiment, a video decoder is provided for receiving a video data stream as an input bitstream, the video data stream having video encoded therein. The video decoder is configured to decode the video data stream to decode the video. To decode the video, the video decoder estimates a coded picture buffer size for a sub-picture depending on information in the video data stream indicating current coded picture buffer size information.
[0306] According to an embodiment, the video data stream may include a signaled coded picture buffer size of the reference level.
[0307] In an embodiment, if the current level is equal to a reference level, the video decoder may for example be configured to determine the coded picture buffer size using the coded picture buffer size of the reference level.
[0308] According to an embodiment, the video decoder may for example be configured to estimate the coded picture buffer size according to the syntax element cpb_size_value_minus1[i][j] of the video data stream.
[0309] In an embodiment, a video decoder may for example be configured to estimate the coded picture buffer size according to the syntax element cpb_size_scale of the video data stream.
[0310] According to an embodiment, the video decoder may be configured to estimate the coded picture buffer size of the sub-picture, e.g., the coded picture buffer size of the video coding layer, e.g., the sub-picture can further estimate the coded picture buffer size with the picture buffer size coded at the network abstraction layer.
[0311] In an embodiment, the video decoder may for example be configured to estimate a video coding layer coded picture buffer size and / or a network abstraction layer coded picture buffer size depending on the value of the reference level fragment.
[0312] According to an embodiment, a video decoder may be configured, for example, to estimate a video coding layer coded picture buffer size according to: SubPicCbpSizeVcl[s]= = Floor(( cpb_size_value_minus1[i][j]+ 1 ) *2 ( 4 + cpb_size_scale ) *RefLevelFraction[i][j]÷256) The device may for example be configured to estimate the network abstraction layer coded picture buffer size according to: SubPicCbpSizeNal[s]= = Floor(( cpb_size_value_minus1[i][j]+ 1 ) *2 ( 4 + cpb_size_scale )*RefLevelFraction[i][j]÷256), RefLevelFraction is the value of the fraction of the reference level.
[0313] In an embodiment, a video decoder may be configured, for example, to estimate a video coding layer coded picture buffer size according to: SubpicCpbSizeVcl[i][j][k]= Floor(CpbVclFactor*MaxCPB*OlsRefLevelFraction[i][j][k]÷256) SubpicCpbSizeNal[i][j][k]=Floor(CpbNalFactor*MaxCPB*OlsRefLevelFraction[i][j][k]÷256) i, j, and k are indices, and OlsRefLevelFraction[i][j][k] are real numbers.
[0314] Further, according to an embodiment, there is provided a video decoder for receiving a video data stream as an input bitstream, the video data stream having video encoded therein, the video decoder being configured to decode the video data stream to decode said video, and for decoding the video, the video decoder being configured to estimate a bitrate of a sub-picture as a function of information in the video data stream indicating bitrate information of a video sequence currently being coded.
[0315] According to an embodiment, the video data stream may, for example, include an indication indicating whether the bit rate of the sub-picture should be estimated using bit rate information of the currently coded video sequence. If the indication of the video data stream indicates that the bit rate of the sub-picture should be estimated using bit rate information of the currently coded video sequence, the video decoder estimates the bit rate using the bit rate information of the currently coded video sequence. If the indication of the video data stream indicates that the bit rate of the sub-picture should be estimated without using bit rate information of the currently coded video sequence, the video decoder estimates the bit rate using a predefined value or a worst-case value without using bit rate information of the currently coded video sequence.
[0316] In an embodiment, the bitrate information of the currently coded video sequence is the signaled bitrate of a reference level. The video data stream may, for example, include the signaled bitrate of the reference level. If the current level is equal to the reference level, the video decoder may, for example, be configured to use the signaled bitrate of the reference level to determine the bitrate of the sub-picture.
[0317] According to an embodiment, the video decoder may for example be configured to estimate the bit rate of the sub-picture depending on the syntax element bit_rate_value_minus1[i][j] of the video data stream.
[0318] In an embodiment, the video decoder may for example be configured to estimate the bit rate in response to a syntax element bit_rate_scale of the video data stream.
[0319] According to an embodiment, the video decoder may be configured to estimate the coded picture buffer size of the sub-picture, e.g., the coded picture buffer size of the video coding layer, e.g., the sub-picture can further estimate the coded picture buffer size with the picture buffer size coded at the network abstraction layer.
[0320] In an embodiment, the video decoder may for example be configured to estimate a video coding layer coded picture buffer size and / or a network abstraction layer coded picture buffer size depending on the value of the reference level fragment.
[0321] According to an embodiment, a video decoder may be configured, for example, to estimate the video coding layer bit rate for a sub-picture by: SubPicBitRateVcl[s]= = Floor(( bit_rate_value_minus1[i][j]+1 ) * 2 (6 + bit_rate_scale) *RefLevelFraction[i][j]÷256)
[0322] The device may for example be configured to estimate the network abstraction layer bit rate for a sub-picture by: SubPicBitRateNal[s]= = Floor(( bit_rate_value_minus1[i][j]+1 ) * 2 (6 + bit_rate_scale) *RefLevelFraction[i][j]÷256)
[0323] RefLevelFraction is the value of the reference level fraction.
[0324] In an embodiment, a video decoder may be configured, for example, to estimate the video coding layer bit rate for a sub-picture by: SubpicBitRateVcl[i][j][k]=Floor(CpbVclFactor*ValBR*OlsRefLevelFraction[0][j][k]÷256)
[0325] SubpicBitRateNal[i][j][k]=Floor(CpbNalFactor*ValBR*OlsRefLevelFraction[0][j][k]÷256) i, j, and k are indices, and OlsRefLevelFraction[i][j][k] are real numbers.
[0326] In an embodiment, i may, for example, indicate the index of a particular indicated reference level, j may, for example, indicate the index of a particular sub-picture of a picture of an access unit of a video data stream, and k may, for example, indicate the index of the maximum temporal sub-layer included in the video data stream and / or at which the video decoder operates.
[0327] According to an embodiment, OlsRefLevelFraction[i][j][k] may for example depend on the variable sli_non_subpic_layers_fraction[i][k], which indicates the i-th fraction of the bitstream level constraint associated with the layer of targetCvss for which sps_num_subpics_minus1 is equal to 0 when Htid is equal to k.
[0328] In an embodiment, If vps_max_layers_minus1 is equal to 0, or there are no layers in the bitstream with sps_num_subpics_minus1 equal to 0, For example, sli_non_subpic_layers_fraction[i][k] is 0,
[0329] If k is less than sli_max_sublayers_minus1 and sli_non_subpic_layers_fraction[i][k] does not exist, For example, sli_non_subpic_layers_fraction[i][k]= sli_non_subpic_layers_fraction[i][k+1].
[0330] According to an embodiment, for example, OlsRefLevelFraction[i][j][k]= =sli_non_subpic_layers_fraction[i][k]+ ( n - sli_non_subpic_layers_fraction[i][k]) ÷ n * (sli_ref_level_fraction_minus1[i][j][k]+ 1). In the formula, n represents a positive integer.
[0331] According to an embodiment, for example, n=256; or n=128; or n=512; or n=1024; or n=2048; or n=4096.
[0332] According to an embodiment, i, j, and k are defined according to sli_ref_level_fraction_minus1, and sli_ref_level_fraction_minus1[i][j][k] plus 1 specifies the ith part of the level constraint associated with sli_ref_level_idc[i][k] for subpictures whose subpicture index is equal to j for layers of targetCvss where sps_num_subpics_minus1 is greater than 0 if the Htid considered as the sublayer index is equal to k.
[0333] Further, according to an embodiment, there is provided a method for receiving a video data stream as an input bitstream, the video data stream having video encoded therein, and a video decoder for decoding the video data stream to decode the video, wherein to decode the video, the video decoder receives a coded picture buffer size of a sub-picture coded in the video data stream and uses the coded picture buffer size of the sub-picture to decode the video, and / or the video decoder receives a bitrate of the sub-picture coded in the video data stream and decodes the video using the bitrate of the sub-picture.
[0334] According to an embodiment, the video data stream may, for example, include an indication indicating whether a sub-picture coded picture buffer size is encoded in the video data stream or whether the sub-picture coded picture buffer size should be estimated. If the video data stream indication indicates that the sub-picture coded picture buffer size should be estimated, the video decoder estimates the sub-picture coded picture buffer size. If the video data stream indication indicates that the sub-picture coded picture buffer size is encoded in the video data stream, the video decoder uses the sub-picture coded picture buffer size encoded in the video data stream.
[0335] In an embodiment, the video data stream may, for example, include an indication indicating whether the sub-picture bit rate is encoded in the video data stream or whether the sub-picture bit rate should be estimated. If the video data stream indication indicates that the sub-picture bit rate should be estimated, the video decoder estimates the sub-picture bit rate. If the video data stream indication indicates that the sub-picture bit rate is encoded in the video data stream, the video decoder uses the coded picture buffer size of the sub-picture that is encoded in the video data stream.
[0336] According to an embodiment, each of the plurality of extractable sub-bitstreams is specific to an output layer set, and a subpicture is assigned to at least one extractable sub-bitstream of the plurality of extractable sub-bitstreams. If the video data stream may, for example, include common decoding unit removal timing information and a plurality of extractable sub-bitstreams, parsing each output layer set-specific hypothesis reference decoder parameter in a video parameter set or a sequence parameter set or a supplemental enhancement information message of the video data stream may, for example, include either a spreading factor or an absolute value of a tick divisor in scaling the removal timing of the common decoding unit. The video decoder may, for example, be configured to process the absolute value of the spreading factor or the tick divisor.
[0337] Further, according to an embodiment, there is provided a video data stream having video encoded therein. The video data stream may, for example, include a syntax element cpb_size_value_minus1[i][j] and a syntax element cpb_size_scale. Alternatively, the video data stream may, for example, include a syntax element bit_rate_value_minus1[i][j] and a syntax element bit_rate_scale.
[0338] Further, according to an embodiment, a video data stream having video encoded therein is provided, the video data stream including an indication indicating whether the coded picture buffer size of the sub-picture should be estimated using current coded picture buffer size information, and / or the video data stream generating an indication indicating whether the bit rate of the sub-picture is estimated using bit rate information of the video sequence currently being coded.
[0339] According to an embodiment, the current coded picture buffer size information is a reference level signaled coded picture buffer size, and the video data stream may, for example, include the reference level signaled coded picture buffer size, and / or the bitrate information of the currently coded video sequence is a reference level signaled bitrate, and the video data stream may, for example, include the reference level signaled bitrate.
[0340] Further, according to an embodiment, a video data stream having video encoded therein is provided, wherein the video data stream includes an indication indicating whether a coded picture buffer size for a sub-picture has been encoded in the video data stream or whether the coded picture buffer size for a sub-picture should be estimated, and / or the video data stream includes an indication indicating whether a bit rate for a sub-picture has been encoded in the video data stream or whether the bit rate for a sub-picture needs to be estimated.
[0341] According to an embodiment, each of the plurality of extractable sub-bitstreams is specific to an output layer set, and a sub-picture is assigned to at least one extractable sub-bitstream of the plurality of extractable sub-bitstreams. If the video data stream may, for example, include common decoding unit removal timing information and a plurality of extractable sub-bitstreams, parsing each output layer set-specific hypothesis reference decoder parameter in a video parameter set or a sequence parameter set or a supplemental enhancement information message of the video data stream may, for example, include either a spreading factor or an absolute value of a tick divisor in scaling the removal timing of the common decoding unit.
[0342] Further, according to an embodiment, there is provided a video encoder for encoding video into a video data stream, such that the video data stream has video encoded therein. The video data stream may for example include the syntax element cpb_size_value_minus1[i][j] and the syntax element cpb_size_scale. Alternatively, the video encoder generates the video data stream such that the video data stream includes the syntax element bit_rate_value_minus1[i][j] and the syntax element bit_rate_scale.
[0343] Further, according to an embodiment, there is provided a video encoder for encoding video into a video data stream, such that the video data stream has video encoded therein, wherein the video encoder generates the video data stream such that the video data stream includes an indication indicating whether a coded picture buffer size of a sub-picture should be estimated using current coded picture buffer size information, and / or the video encoder generates the video data stream such that the video data stream includes generating an indication indicating whether a bitrate of a sub-picture is estimated using currently coded video sequence bitrate information.
[0344] According to an embodiment, the video encoder may be configured, for example, to generate a video data stream such that the current coded picture buffer size information is a reference level signaled coded picture buffer size, and the video data stream may, for example, include the reference level signaled coded picture buffer size, and / or the video encoder may be configured, for example, to generate a video data stream such that the bitrate information of a currently coded video sequence is a reference level signaled bitrate, and the video data stream may, for example, include the reference level signaled bitrate.
[0345] Further, according to an embodiment, there is provided a video encoder for encoding video into a video data stream, such that the video data stream has video encoded therein, wherein the video encoder generates the video data stream such that the video data stream includes an indication indicating whether a coded picture buffer size for a sub-picture has been encoded in the video data stream or whether the coded picture buffer size for a sub-picture should be estimated, and / or the video encoder generates the video data stream such that the video data stream includes an indication indicating whether a bit rate for a sub-picture has been encoded in the video data stream or whether the bit rate for a sub-picture needs to be estimated.
[0346] Further, according to an embodiment, there is provided a video encoder for encoding video into a video data stream, such that the video data stream has video encoded therein. The video encoder generates the video data stream such that, when the video data stream includes common decoding unit removal timing information and multiple extractable sub-bitstreams, each of the multiple sub-bitstreams is specific to an output layer set, and parsing each output layer set-specific hypothetical reference decoder parameter in a video parameter set or a sequence parameter set or a supplemental enhancement information message of the video data stream includes either a spreading factor or an absolute value of a tick divisor in scaling the removal timing of the common decoding unit.
[0347] Further, according to an embodiment, there is provided a system for encoding video into a video data stream and for decoding video. The system includes a video encoder as described above and a video decoder as described above. The video encoder may, for example, be configured to encode video into a video data stream such that the video data stream has the video encoded therein. The video decoder is configured to receive the video data stream and decode the video data stream to decode it into video. The current VVC draft specification includes an SEI message that provides additional information useful for estimating the level information of subpictures and subpicture set levels. This is achieved by signaling the contribution of a subpicture to a particular signal reference level. Based on the percentage of each subpicture's level, the variables CPB size and bitrate, and the level of the bitstream of subsequent subpictures, are estimated. The percentage of the cumulative level is also used to derive the CPB size and bitrate of a bitstream consisting of a set of subpictures and to estimate their levels. However, all of these derivations use the MaxCPBSize and / or MaxBitrate of the reference level, which is problematic because most bitstreams may not fully occupy the CPB and bitrate budget up to the MaxCPBSize and / or MaxBitrate of a particular level. Using these values to estimate the level of a subpicture bitstream or a merged bitstream of a subpicture set bitstream increases the likelihood of over-provisioning the (cumulative) CPB size and bitrate.
[0348] The VVC draft specification includes the derivation of the variables SubPicCpbSizeVcl[i][j] and SubPicCpbSizeNal[i][j], as follows: SubPicCpbSizeVcl[i][j]=Floor(CpbVclFactor*MaxCPB*RefLevelFraction[i][j]÷256) SubPicCpbSizeNal[i][j]=Floor(CpbNalFactor*MaxCPB*RefLevelFraction[i][j]÷256)
[0349] This is because the original bitstream containing all sub-pictures may already have more accurate information about the exact CPB size and bitrate of the bitstream (cpb_size_value_minus1 and cpb_size_scale), in addition to the information about the respective maximum values derived by reference from the bitstream level. Therefore, the object of the present invention is to derive much more accurate values for the respective CBP size and bitrate per sub-picture, which can be used for approximation at the sub-picture set level. In a further embodiment, the sub-picture CPB is derived as follows: SubPicCbpSizeVcl[s]= Floor(( cpb_size_value_minus1[i][j]+ 1 ) * 2 ( 4 + cpb_size_scale ) *RefLevelFraction[i][j]÷256) SubPicCbpSizeNal[s] = Floor(( cpb_size_value_minus1[i][j]+ 1 ) * 2 ( 4 + cpb_size_scale ) *RefLevelFraction[i][j]÷256)
[0350] Note that the syntax elements cpb_size_value_minus1[i][j] are sent separately for the Vcl and Nal HRD parameters and are used respectively in the above derivation, so the values of SubPicCbpSizeVcl[s] and SubPicCbpSizeVcl[s] potentially come from different values.
[0351] Because CPB size is signaled for a specific level, the above derivation can only be performed if that level is included as a reference level.
[0352] For bitrates, the respective derivations are modified from using the maximum bitrate as a reference as follows (Be [Vcl / Nal]Factor * May BE), SubPicBitRateVcl[s]= Floor( BrVclFactor * MaxBR *RefLevelFraction[i][j]÷256) SubPicBitRateNal[s]= Floor( BrNalFactor * MaxBR *RefLevelFraction[i][j]÷256) The actual bit rate of the signaled bitstream is used ((bit_rate_value_minus1[i][j]+1) * 2 (6 + bit_rate_scale ), listed below SubPicBitRateVcl[s] = Floor(( bit_rate_value_minus1[i][j]+1 ) * 2 (6 + bit_rate_scale) *RefLevelFraction[i][j]÷256) SubPicBitRateNal[s] = Floor(( bit_rate_value_minus1[i][j]+1 ) * 2 (6 + bit_rate_scale) *RefLevelFraction[i][j]÷256) Note that the same as above about CPB size applies here, and that the value of the syntax element bit_rate_value_minus1[i][j] depends on whether the Nal or Vcl HRD is considered, and therefore the values of SubBitRateVcl[s] and SubPicBitrateVcl[s] are derived to potentially different values.
[0353] In an embodiment, the variables SubpicCpbSizeVcl[i][j][k] and SubpicCpbSizeNal[i][j][k] are derived as follows: SubpicCpbSizeVcl[i][j][k]=Floor(CpbVclFactor*MaxCPB*OlsRefLevelFraction[i][j][k]÷256) SubpicCpbSizeNal[i][j][k]=Floor(CpbNalFactor*MaxCPB*OlsRefLevelFraction[i][j][k]÷256) In embodiments where MaxCPB is derived from sli_ref_level_idc[i][k], the variables SubpicCpbSizeVcl[i][j][k] and SubpicCpbSizeNal[i][j][k] are derived as follows: SubpicBitRateVcl[i][j][k]=Floor(CpbVclFactor*ValBR*OlsRefLevelFraction[0][j][k]÷256) SubpicBitRateNal[i][j][k]=Floor(CpbNalFactor*ValBR*OlsRefLevelFraction[0][j][k]÷256)
[0354] For example, the variables OlsRefLevelFraction[i][j][k] are numbers, eg, real numbers.
[0355] For example, the variable OlsRefLevelFraction[i][j][k] may be derived according to, for example, sli_non_subpic_layers_fraction[i][k]+ ( n - sli_non_subpic_layers_fraction[i][k]) ÷ n * (sli_ref_level_fraction_minus1[i][j][k]+ 1) In the formula, n represents a positive integer, for example, n=256; or, for example, n=128; or, for example, n=512; or, for example, n=1024; or, for example, n=2048; or, for example, n=4096.
[0356] So, for example, OlsRefLevelFraction[i][j][k]= =sli_non_subpic_layers_fraction[i][k]+ ( 256 - sli_non_subpic_layers_fraction[i][k]) ÷ 256 * (sli_ref_level_fraction_minus1[i][j][k]+ 1).
[0357] sli_non_subpic_layers_fraction[i][k] may indicate the i-th portion of the bitstream-level constraints associated with layers of targetCvss for which sps_num_subpics_minus1 is equal to 0, e.g., when Htid is equal to k. If vps_max_layers_minus1 is equal to 0, or if no layer in the bitstream has sps_num_subpics_minus1 equal to 0, sli_non_subpic_layers_fraction[i][k] is equal to 0. When k is less than sli_max_sublayers_minus1 and sli_non_subpic_layers_fraction[i][k] is not present, it is inferred to be equal to sli_non_subpic_layers_fraction[i][k+1], and sli_ref_level_fraction_minus1[i][j][k] plus 1 specifies the ith part of the level restriction, pertaining to sli_ref_level_idc[i][k], for subpictures, if Htid is equal to k, the subpicture index of the layer in targetCvss where sps_num_subpics_minus1 is greater than 0 is equal to j. When k is less than sli_max_sublayers_minus1 and sli_ref_level_fraction_minus1[i][j][k] is not present, it is inferred to be equal to sli_ref_level_fraction_minus1[i][j][k+1].
[0358] Alternatively, in another embodiment, the CPB size and / or bit rate for each sub-picture can be directly signaled instead of being derived, or in addition, there can be a gating flag that indicates whether the value can be derived or whether the value is explicitly signaled.
[0359] The following describes DU timing signaling in Picture Timing SEI.
[0360] According to an embodiment, a video data stream having video encoded therein is provided, the video data stream including a plurality of access units, and further including delta time information for each of two or more decoding units of the access units, wherein a decoding unit removal time for each decoding unit of the two or more decoding units of the access units depends on the access unit removal time of the access unit and depends on the delta time information for the decoding unit.
[0361] According to an embodiment, the video data stream may for example comprise picture timing supplemental extension information, which may for example comprise delta time information of two or more decoding units of said access unit.
[0362] In an embodiment, the delta time information indicates a difference in removal time between two of the two or more decoding units of the access unit.
[0363] According to an embodiment, the last decoding unit of the two or more decoding units of said access unit has a removal time equal to the removal time of said access unit.
[0364] In an embodiment, the access unit may, for example, include three or more decoding units, and the difference in removal time is equal for each pair of two consecutive decoding units of the three or more decoding units of the access unit.
[0365] In an embodiment, the picture timing supplemental enhancement information is applied to a sub-bitstream derived from the video data stream, and the number of decoding units remains constant.
[0366] In an embodiment, the frame time interval is signaled in the parameter set of the video data stream, the HRD parameter of the sequence parameter set.
[0367] In an embodiment, the frame time interval can be derived as the difference between the removal times of two consecutive access units at the highest temporal level.
[0368] In an embodiment, each decoding unit of the two or more decoding units of the access unit may include, for example, one video coding layer network abstraction layer unit.
[0369] In an embodiment, the picture timing supplemental enhancement information is applied to sub-bitstreams derived from the video data stream, where there are different numbers of decoding units.
[0370] In an embodiment, the frame time interval can be derived as follows: Multiply (elemental_duration_in_tc_minus1[maxTiD]+1) by ClockTicks.
[0371] In an embodiment, the video data stream may for example include an indication of whether there is a variation for the video data stream in the number of decoding units.
[0372] In an embodiment, the video data stream does not include an indication of whether the number of decoding units is variable for the video data stream.
[0373] In an embodiment, the number of decoding units in an access unit depends on the frame time interval and the common delay increment.
[0374] In an embodiment, the video data stream includes decoding unit information supplemental enhancement information messages for two or more decoding units of the access unit, where the decoding unit information supplemental enhancement information messages for the decoding units may include, for example, delta time information for the decoding units.
[0375] In an embodiment, a video data stream may, for example, include a minimum picture duration flag in a video parameter set or a sequence parameter set of said video data stream, said minimum picture duration flag indicating whether frame time interval information is present in the absence of a constant frame rate.
[0376] Further, according to an embodiment, there is provided a video encoder for encoding video into a video data stream, such that the video data stream has video encoded therein. The video encoder generates the video data stream such that the video data stream includes a plurality of access units. Furthermore, the video encoder generates the video data stream such that the video data stream includes delta time information for each of two or more decoding units of an access unit of the plurality of access units, and a decoding unit removal time of each decoding unit of the two or more decoding units of the access unit depends on the access unit removal time of the access unit and depends on the delta time information of the decoding unit.
[0377] According to an embodiment, a video encoder may be configured to generate a video data stream such that the video data stream may include, for example, picture timing supplemental extension information. The video encoder may be configured to generate a video data stream such that the picture timing supplemental extension information may include, for example, delta time information of two or more decoding units of said access unit.
[0378] In an embodiment, the video encoder may for example be configured to generate the video data stream such that the delta time information indicates a difference in removal time between two of the two or more decoding units of said access unit.
[0379] According to an embodiment, the last decoding unit of the two or more decoding units of said access unit has a removal time equal to the removal time of said access unit.
[0380] In an embodiment, the video encoder may be configured to generate the video data stream such that the access unit may include, for example, three or more decoding units, and the removal time difference is equal for each pair of two consecutive decoding units of the three or more decoding units of the access unit.
[0381] In an embodiment, the picture timing supplemental enhancement information is applied to a sub-bitstream derived from the video data stream, and the number of decoding units remains constant.
[0382] In an embodiment, a video encoder may for example be configured to generate a video data stream such that the frame time interval is signaled in a parameter set of the video data stream, an HRD parameter of a sequence parameter set.
[0383] According to an embodiment, a video encoder may for example be configured to generate a video data stream such that the frame time interval is derivable as the difference between the removal times of two consecutive access units at the highest temporal level.
[0384] In an embodiment, the video encoder may, for example, be configured to generate a video data stream such that each decoding unit of two or more decoding units of the access unit may, for example, include one video coding layer network abstraction layer unit.
[0385] According to an embodiment, a video encoder may be configured, for example, to generate a video data stream such that picture timing supplemental extension information is applied to sub-bitstreams derived from the video data stream and there are different numbers of decoding units.
[0386] In an embodiment, the frame time interval can be derived as follows: Multiply (elemental_duration_in_tc_minus1[maxTiD]+1) by ClockTicks.
[0387] In an embodiment, the video encoder may, for example, be configured to generate the video data stream, such that the video data stream may, for example, include an indication of whether the video data stream is variable in number of decoding units.
[0388] In an embodiment, the video encoder may for example be configured to generate the video data stream such that the video data stream does not include an indication of whether the video data stream is variable in number of decoding units for the video data stream.
[0389] According to an embodiment, the video encoder may for example be configured to generate the video data stream such that the number of decoding units in an access unit depends on the frame time interval and the common delay increment.
[0390] In an embodiment, a video encoder may, for example, be configured to generate a video data stream such that the video data stream may, for example, include decoding unit information supplemental enhancement information messages of two or more decoding units of the access unit, and the video encoder may, for example, be configured to generate a video data stream such that the decoding unit information supplemental enhancement information messages of the decoding units may, for example, include the delta time information of the decoding units.
[0391] In an embodiment, a video data stream may, for example, include a minimum picture duration flag in a video parameter set or a sequence parameter set of said video data stream, said minimum picture duration flag indicating whether frame time interval information is present in the absence of a constant frame rate.
[0392] Further, according to one embodiment, there is provided a method for receiving a video data stream as an input bitstream, the video data stream having video encoded therein. The video data stream includes a plurality of access units. A video decoder is configured to decode the video data stream to decode the video. Furthermore, the video data stream includes delta time information for each of two or more decoding units of the access units, wherein a decoding unit removal time of each decoding unit of the two or more decoding units of the access units depends on the access unit removal time of the access unit and depends on the delta time information of the decoding unit, and the video decoder decodes the video data stream using the delta time information for each of the two or more decoding units of the access units.
[0393] According to an embodiment, the video data stream may for example comprise picture timing supplemental extension information, which may for example comprise delta time information of two or more decoding units of said access unit.
[0394] In an embodiment, the delta time information indicates a difference in removal time between two of two or more decoding units of the access unit, and the video decoder may be configured, for example, to decode the video data stream using the difference in removal time between the two decoding units.
[0395] According to an embodiment, the last decoding unit of the two or more decoding units of the access unit has a removal time equal to the removal time of the access unit, and a video decoder may be configured, for example, to decode the video data stream using the removal time of the access unit.
[0396] In an embodiment, the access unit may, for example, include three or more decoding units, and the difference in removal time is equal for each pair of two consecutive decoding units of the three or more decoding units of the access unit.
[0397] In an embodiment, the picture timing supplemental enhancement information is applied to a sub-bitstream derived from the video data stream, and the number of decoding units remains constant.
[0398] In an embodiment, the frame time interval is signaled in a parameter set of a video data stream, an HRD parameter of a sequence parameter set, and a video decoder may for example be configured to decode said video data stream using said frame time interval.
[0399] According to an embodiment, the frame time interval can be derived as the difference between the removal times of two consecutive access units at the highest temporal level, and a video decoder can be configured, for example, to decode the video data stream using the frame time interval.
[0400] In an embodiment, each decoding unit of the two or more decoding units of the access unit may include, for example, one video coding layer network abstraction layer unit.
[0401] In an embodiment, the picture timing supplemental enhancement information is applied to sub-bitstreams derived from the video data stream, where there are different numbers of decoding units.
[0402] In an embodiment, the video decoder may be configured to derive the frame time interval, for example, according to the following: (elemental_duration_in_tc_minus1[maxTiD]+1) multiplied by ClockTicks.
[0403] According to an embodiment, the video data stream may, for example, include an indication indicating whether the number of decoding units is variable for the video data stream, and the video decoder may, for example, be configured to decode the video data stream by processing the indication.
[0404] In an embodiment, the video data stream does not include an indication of whether the number of decoding units is variable for the video data stream.
[0405] According to an embodiment, the number of decoding units in an access unit depends on the frame time interval and the common delay increment.
[0406] In an embodiment, the video data stream may include decoding unit information supplemental enhancement information messages for two or more decoding units of the access unit, where the decoding unit information supplemental enhancement information messages for the decoding units may include, for example, delta time information for the decoding units, and a video decoder may, for example, be configured to decode the video data stream using the delta time information for the decoding units.
[0407] In an embodiment, a video data stream may, for example, include a minimum picture duration flag in a video parameter set or a sequence parameter set of said video data stream, said minimum picture duration flag indicating whether frame time interval information is present in the absence of a constant frame rate.
[0408] Further, according to an embodiment, there is provided a system for encoding video into a video data stream and for decoding video. The system includes a video encoder as described above and a video decoder as described above. The video encoder may, for example, be configured to encode the video into the video data stream such that the video data stream has the video encoded therein. The video decoder may, for example, be configured to receive the video data stream and decode the video data stream into video.
[0409] As mentioned above, DU timing is given as a delta of AU timing. More specifically, the removal time of a DU is indicated by giving a delta time in either the PictureTimingSEI message or the Decoding Unit Information SEI message, which is related to the removal time of the AU that contains the particular DU.
[0410] If the information is included in the picture timing SEI message, the information signaled is the removal time difference between two DUs. The removal time of the last DU of an AU is the same as that of the AU removal time, and any other DU is signaled as the removal time difference to the next DU. There are two options to signal it (highlighted in different colors): [Table 5] The first way to signal it is when the DUs all have the same removal time difference in common. The second case is when the removal time differences between DUs within an AU are not the same.
[0411] The current syntax of the Picture Timing SEI message when DU timing is present prevents the application of aspect 2 described in this invention because the number of DUs may change when sub-bitstreams are extracted from the bitstream.
[0412] In principle, if the common removal time difference between CUs is the same, it may be possible to make the picture timing SEI message still apply even if the number of DUs is changed. In one embodiment, the PT SEI message has a mode, where the PT SEI message applies to sub-bitstreams with different numbers of DUs. The number of DUs is derived from other syntax elements. The PT SEI message is modified as follows: [Table 6]
[0413] If du_not_constraint_flag is equal to 0, the value of du_common_cpb_removal_delay_flag is inferred to be 1. The value of num_decoding_units_minus1 is inferred to be equal to FrameTimeInterval divided by (du_common_cpb_removal_delay_increment_minus1 + 1) * ClockSubTicks, minus 1.
[0414] The FrameTimeInterval can be signaled in the parameter set, the HRD parameter of the SPS, or derived as the difference between the deletion times of two consecutive access units at the highest temporal level.
[0415] In this case, there is an additional constraint that each DU contains one VCL NAL unit.
[0416] To explicitly signal the FrameTimeInterval, you can do the following: [Table 7]
[0417] min_pic_duration_within_cvs_present_flag is an SPS or VPS flag that indicates the presence of a FrameTimeInterval when there is no constant frame rate, allowing the FrameTimeInterval to be indicated both when the frame rate is constant and when the frame rate is not constant.
[0418] In the second case, du_not_constraint_flag can be set to 1 only if the frame rate is constant. In that case, the value of FrameTimeInterval is derived as (elemental_duration_in_tc_minus1 [maxTiD] + 1) multiplied by ClockTicks.
[0419] Alternatively, the additional signaling flag can be omitted by merging the respective indications with a common DU timing mode signaling as follows: [Table 8]
[0420] Note that the above derivation of the number of DUs in an AU is based on the FrameTimeInterval and a common delay increment (seconds given in number of clock-sub-ticks). Note also that in aspect 4, the clock-sub-ticks vary depending on the temporal sub-layer present in the bitstream. The number of DUs is derived using the clock-sub-ticks of the highest temporal layer, or the spreading factor described in aspect 4 is also taken into account in the FrameTimeInterval, resulting in the same result.
[0421] While some aspects are described in terms of apparatus, it will be apparent that these aspects also represent a description of a corresponding method, where a block or device corresponds to a method step or feature of a method step. Similarly, aspects described in terms of method steps also represent a description of a corresponding block or item or feature of a corresponding apparatus. Some or all method steps may be performed by (or using) a hardware apparatus, such as, for example, a microprocessor, a programmable computer, or an electronic circuit. In some embodiments, one or more of the most important method steps may be performed by such an apparatus.
[0422] Depending on certain implementation requirements, embodiments of the present invention may be implemented in hardware or software, or at least partially in hardware or at least partially in software. This implementation may be performed using a digital storage medium, such as a floppy disk, DVD, Blu-ray, CD, ROM, PROM, EPROM, EEPROM, or FLASH memory, having electronically readable control signals that may be stored thereon and that cooperate (or may be able to cooperate) with a programmable computer system to perform the respective methods. Thus, the digital storage medium may be computer-readable.
[0423] Some embodiments according to the invention include a data carrier having electronically readable control signals that can cooperate with a programmable computer system to perform one of the methods described herein.
[0424] Generally, embodiments of the present invention may be implemented as a computer program product having program code operable to perform one of the methods when the computer program product runs on a computer, which may for example be stored on a machine-readable carrier.
[0425] Other embodiments comprise the computer program for performing one of the methods described herein, stored on a machine readable carrier.
[0426] In other words, an embodiment of the inventive methods is, therefore, a computer program having a program code for performing one of the methods described herein, when the computer program runs on a computer.
[0427] A further embodiment of the inventive method is therefore a data carrier (or digital storage medium, or computer-readable medium) comprising, recorded on it, the computer program for performing one of the methods described herein. The data carrier, digital storage medium, or recorded medium is typically tangible and / or non-transitory.
[0428] A further embodiment of the inventive method is, therefore, a data stream or a sequence of signals representing the computer program for performing one of the methods described herein, The data stream or the sequence of signals can be adapted to be transferred via a data communication connection, for example via the Internet.
[0429] A further embodiment comprises a processing means, such as for example a computer, or a programmable logic device, configured to or adapted to perform one of the methods described herein.
[0430] A further embodiment comprises a computer having installed thereon the computer program for performing one of the methods described herein.
[0431] Further embodiments according to the invention include an apparatus or system configured to transfer (e.g., electronically or optically) a computer program for performing one of the methods described herein to a receiver. The receiver may be, for example, a computer, a mobile device, a memory device, etc. The apparatus or system may, for example, comprise a file server for transferring the computer program to the receiver.
[0432] In some embodiments, a programmable logic device (e.g., a field programmable gate array) may be used to perform some or all of the functions of the methods described herein. In some embodiments, a field programmable gate array may cooperate with a microprocessor to perform one of the methods described herein. In general, the methods are preferably performed by any hardware apparatus.
[0433] The devices described herein may be implemented using a hardware device, or using a computer, or using a combination of a hardware device and a computer.
[0434] The methods described herein may be performed using a hardware apparatus, or using a computer, or using a combination of a hardware apparatus and a computer.
[0435] The above-described embodiments are merely illustrative of the principles of the present invention. It is understood that modifications and variations of the arrangements and details described herein will be apparent to those skilled in the art. It is therefore the intention to be limited only by the scope of the appended claims and not by the specific details presented by way of description and illustration of the embodiments herein.
[0436] References [1] ISO / IEC, ITU-T. High efficiency video coding. ITU-T Recommendation H.265 | ISO / IEC 23008 10 (HEVC), edition 1, 2013;edition 2, 2014.
Claims
1. 1. A method for decoding a picture from a bitstream, comprising: decoding, from the bitstream, information indicating a bitrate of the bitstream, the information including a first syntax element, a second syntax element, and a reference level fraction of the bitstream; deriving a bit rate of a sub-picture of the picture based on the information, the deriving of the bit rate of the sub-picture comprising: The bit rate of the bit stream is (first syntax element + 1) * 2 (6+第2の構文要素) and determining multiplying the bitrate of the bitstream by a reference level fraction; A method comprising:
2. the first syntax element indicates a bit rate value of the picture; the second syntax element indicates a bit rate scale of the picture; The method of claim 1.
3. The method includes deriving a video coding layer (VCL) bitrate for the sub-picture; deriving a Network Abstraction Layer (NAL) bitrate for the sub-picture; The method of claim 1 , comprising:
4. the value indicated by the first syntax element depends on the NAL of the sub-picture or the VCL of the sub-picture; The method of claim 3.
5. the reference level fraction is the fraction that the subpicture contributes to a given signaled reference level, and is signaled in a supplemental enhancement information (SEI) message in the bitstream. The method of claim 1.
6. deriving the bit rate of the subpicture includes deriving Floor(bit rate * reference level fraction / 256); The method of claim 1.
7. decoding additional information indicating a current coded picture buffer size from the bitstream, the additional information including a third syntax element and a fourth syntax element; deriving the current coded picture buffer size for the sub-picture based on the third syntax element and the fourth syntax element; The method of claim 1 further comprising:
8. 1. An electronic device for decoding pictures from a bitstream, comprising: decoding, from the bitstream, information indicating a bitrate of the bitstream, the information including a first syntax element, a second syntax element, and a reference level fraction of the bitstream; deriving a bit rate for a sub-picture based on the information; deriving the bit rate for the sub-picture; The bit rate of the bit stream is (first syntax element + 1) * 2 (6+第2の構文要素) It was decided that multiplying the bit rate of the bitstream by the reference level fraction; 11. An electronic device comprising: at least one processor configured to:
9. the first syntax element indicates a bit rate value of the picture; the second syntax element indicates a bit rate scale of the picture; 9. The electronic device of claim 8.
10. Deriving a video coding layer (VCL) bit rate for the subpicture; deriving a Network Abstraction Layer (NAL) bitrate for the sub-picture; 9. The electronic device of claim 8, further comprising the at least one processor configured to:
11. the value indicated by the first syntax element depends on the NAL of the sub-picture or the VCL of the sub-picture; The electronic device of claim 10.
12. the reference level fraction is the fraction that the subpicture contributes to a given signaled reference level, and is signaled in a supplemental enhancement information (SEI) message in the bitstream.
9. The electronic device of claim 8.
13. deriving Floor(bitrate*reference level fraction / 256) to derive the bitrate of the subpicture; 9. The electronic device of claim 8, comprising the at least one processor configured to:
14. decoding additional information from the bitstream indicating a current coded picture buffer size, the additional information including a third syntax element and a fourth syntax element; deriving the current coded picture buffer size for the sub-picture based on the third syntax element and the fourth syntax element; 9. The electronic device of claim 8, further comprising the at least one processor configured to:
15. 1. A non-transitory computer-readable medium comprising instructions that, when executed by at least one processor, cause an electronic device to operate to decode pictures from a bitstream as follows: The electronic device is decoding, from the bitstream, information indicating a bitrate of the bitstream, the information including a first syntax element, a second syntax element, and a reference level fraction of the bitstream; Derive a bit rate for a sub-picture based on the information, the instructions configuring at least one processor to derive the bit rate for the sub-picture as follows: The bit rate of the bit stream is (first syntax element + 1) * 2 (6+第2の構文要素) It was decided that multiplying the bit rate of the bitstream by the reference level fraction; a non-transitory computer-readable medium configured to:
16. the first syntax element indicates a bit rate value of the picture; the second syntax element indicates a bit rate scale of the picture; 16. The non-transitory computer-readable medium of claim 15.
17. Deriving a video coding layer (VCL) bit rate for the subpicture; deriving a Network Abstraction Layer (NAL) bitrate for the sub-picture; 20. The non-transitory computer-readable medium of claim 15, further comprising instructions for causing the at least one processor to operate as follows:
18. the value indicated by the first syntax element depends on the NAL of the sub-picture or the VCL of the sub-picture; 20. The non-transitory computer-readable medium of claim 17.
19. the reference level fraction is the fraction that the subpicture contributes to a given signaled reference level, and is signaled in a supplemental enhancement information (SEI) message in the bitstream.
16. The non-transitory computer-readable medium of claim 15.
20. deriving Floor(bitrate*reference level fraction / 256) to derive the bitrate of the subpicture; 20. The non-transitory computer-readable medium of claim 15, further comprising instructions for causing the at least one processor to operate as follows:
21. decoding additional information from the bitstream indicating a current coded picture buffer size, the additional information including a third syntax element and a fourth syntax element; deriving the current coded picture buffer size for the sub-picture based on the third syntax element and the fourth syntax element; 20. The non-transitory computer-readable medium of claim 15, further comprising instructions for causing the at least one processor to operate as follows:
22. 1. A method for encoding at least one picture into a bitstream, comprising: The bit rate of the bit stream is (first syntax element + 1) * 2 (6+第2の構文要素) and determining encoding the first syntax element and the second syntax element; deriving, for a sub-picture of said at least one picture, a fraction that said sub-picture contributes to a given signaled reference level; encoding the fraction of the signaled reference level that indicates a bit rate for the subpicture; A method comprising:
23. 1. An electronic device for encoding at least one picture into a bitstream, comprising: The bit rate of the bit stream is (first syntax element + 1) * 2 (6+第2の構文要素) It was decided that encoding the first syntax element and the second syntax element; deriving, for a sub-picture of said at least one picture, a fraction that said sub-picture contributes to a given signaled reference level; encoding the fraction of the signaled reference level that indicates a bit rate for the subpicture; 11. An electronic device comprising: at least one processor configured to:
24. 1. A non-transitory computer-readable medium for encoding pictures from a bitstream, comprising instructions that, when executed by at least one processor, cause an electronic device to operate as follows: The electronic device is The bit rate of the bit stream is (first syntax element + 1) * 2 (6+第2の構文要素) It was decided that encoding the first syntax element and the second syntax element; deriving, for a sub-picture of said at least one picture, a fraction that said sub-picture contributes to a given signaled reference level; encoding the fraction of the signaled reference level that indicates a bit rate for the subpicture; a non-transitory computer-readable medium configured to:
Citation Information
Patent Citations
Electronic device for signaling a sub-picture buffer parameter
WO2014050059A1