Video encoder, video decoder, encoding and decoding methods and video data streams for implementing advanced video encoding concepts
By introducing indication and buffer delay offset information into the video data stream, the video encoding and decoding process is optimized, solving the problem of insufficient parallel processing capability in the HEVC standard and improving encoding efficiency and decoding accuracy.
Patent Information
- Application Number
- CN202511329572.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2020-05-22
- Filing Date
- 2021-05-21
- Publication Date
- 2025-12-09
AI Technical Summary
Existing video coding technologies are insufficient in terms of parallel processing capabilities and video quality/rate adaptation, especially in the HEVC standard, which struggles to efficiently support parallel processing of video encoders and decoders and image segmentation.
The video encoding and decoding process is optimized by introducing an indication in the video data stream that determines whether to output images that depend on random access images, as well as encoding image buffer delay offset information and non-scalable nested image timing supplementation enhancement information.
It improves the parallel processing capability of video encoding and decoding, enhances video quality/rate adaptation, and improves encoding efficiency and decoding accuracy.
Smart Images

Figure CN121099073A_ABST
Abstract
Description
Division
[0001] This application is a divisional application of the application patent application with the application date of 21 May 2021, the application number 202180036764.4, and the invention name “Video encoder, video decoder, encoding and decoding method, and video data stream for implementing advanced video coding concepts”. TECHNICAL FIELD
[0002] The present invention relates to video encoding and video decoding, and in particular to a video encoder, a video decoder, an encoding and decoding method, and a video data stream for implementing advanced video coding concepts. BACKGROUND
[0003] H.265 / HEVC (HEVC = High Efficiency Video Coding) is a video codec that already provides tools for boosting or even enabling parallel processing at the encoder and / or decoder. For example, HEVC supports a subdivision of a picture into a set of tiles that are coded independently from each other. Another concept supported by HEVC relates to WPP, according to which CTU rows or CTU lines of a picture can be processed in parallel from left to right, e.g. in stripes, under the premise that some minimum CTU offset is respected when processing consecutive CTU rows. However, it would be advantageous to have a video codec that even more efficiently supports the parallel processing capabilities of a video encoder and / or a video decoder.
[0004] In the following, an introduction to VCL partitioning (VCL = Video Coding Layer) according to the prior art is described.
[0005] Generally, in video coding, the coding process of picture samples requires a small partitioning, in which samples are divided into rectangular regions for joint processing, e.g. prediction or transform coding. Therefore, pictures are partitioned into blocks of a certain size, which is constant during the encoding of a video sequence. In the H.264 / AVC (AVC = Advanced Video Coding) standard, a fixed size block of 16x16 samples, so-called macroblock, is used.
[0006] In the prior art HEVC standard (cf. [1]), there are coding tree blocks (CTB) or coding tree units (CTU) of a maximum size of 64x64 samples. In further descriptions of HEVC, the more common term CTU is used for such blocks.
[0007] CTUs are processed in a raster scan order (starting with the CTU in the top left, processing the CTUs in a picture row by row until the CTU in the bottom right).
[0008] Encoded CTU data is organized into a container called a slice. Initially, in previous video coding standards, a slice referred to a segment containing one or more consecutive CTUs of a picture. A slice was used for the segmentation of the encoded data. From another perspective, a complete picture can also be defined as one large segment, and therefore, historically, the term slice still applies. In addition to the encoded picture samples, a slice also includes additional information related to the encoding process of the slice itself, which is placed into a so-called slice header.
[0009] According to the state of the art, VCL (Video Coding Layer) also comprises techniques for segmentation and spatial partitioning. Such partitioning can be applied to video coding for various reasons, including among others parallelization of processing load balancing, CTU size matching in network transmission, error mitigation, etc.
[0010] Other examples are related to RoI (RoI = Region of Interest) coding, where for example in the middle of a picture there is a region that an observer can select, for example by a zoom operation (decoding of the RoI only) or gradual decoder refresh (GDR), where intra data (typically put into one frame of a video sequence) is distributed in time over several consecutive frames, for example as a column of intra blocks, which slide over the picture plane and locally reset the temporal prediction chain in the same way as an intra picture resets the temporal prediction chain for the whole picture plane. For the latter, there are two regions in each picture, one that is recently reset, and one that can be affected by errors and error propagation.
[0011] Reference Picture Resampling (RPR) is a technique for video coding that adapts the quality / rate of a video not only by using coarser quantization parameters but also by adapting the resolution of possibly each transmitted picture. Thus, the reference for inter prediction can have a different size than the picture that is currently being predicted for encoding. Basically, RPR requires a resampling process in the prediction loop, for example to define up- and down-sampling filters.
[0012] Depending on the style, RPR can lead to a change in the encoded picture size at any picture or be limited to occur only at certain specific pictures, for example only at specific locations that are delimited by for example segment boundary adaptive HTTP streaming. SUMMARY
[0013] It is an object of the present invention to provide an improved concept for video coding and video decoding.
[0014] It is an object of the present invention to provide an improved concept for video coding and video decoding.
[0015] According to a first aspect of the application, a device for receiving an input video data stream is provided. The input video data stream has a video encoded therein. The device is configured to generate an output video data stream from the input video data stream. Furthermore, the device is to determine whether pictures of the video that precede a dependent random access picture should be output.
[0016] Furthermore, a video data stream is provided. The video data stream has a video encoded therein. The video data stream comprises an indication whether pictures of the video that precede a dependent random access picture should be output.
[0017] Furthermore, a video encoder is provided. The video encoder is configured to encode a video into a video data stream. The video encoder is configured to generate the video data stream such that the video data stream comprises an indication whether pictures of the video that precede a dependent random access picture should be output.
[0018] Furthermore, a video decoder for receiving a video data stream having a video stored therein is provided. The video decoder is configured to decode the video from the video data stream. The video decoder is configured to decode the video depending on an indication whether pictures of the video that precede a dependent random access picture should be output.
[0019] Furthermore, a method for receiving an input video data stream is provided. The input video data stream has a video encoded therein. The method comprises generating an output video data stream from the input video data stream. Furthermore, the method comprises determining whether pictures of the video that precede a dependent random access picture should be output.
[0020] Furthermore, a method of encoding a video into a video data stream is provided. The method comprises generating the video data stream such that the video data stream comprises an indication whether pictures of the video that precede a dependent random access picture should be output.
[0021] Furthermore, a method for receiving a video data stream having a video stored therein is provided. The method comprises decoding the video from the video data stream. The video is decoded depending on an indication whether pictures of the video that precede a dependent random access picture should be output.
[0022] Furthermore, a computer program for implementing one of the above methods when the computer program is executed by a computer or signal processor is provided.
[0023] According to a second aspect of the application, an apparatus for receiving one or more input video data streams is provided. Each of the one or more input video data streams has an input video encoded therein. The apparatus is configured to generate an output video data stream from the one or more input video data streams, the output video data stream encoding an output video, wherein the apparatus is configured to generate the output video data stream such that the output video is the input video encoded in one of the one or more input video data streams or such that the output video depends on the input video of at least one of the one or more input video data streams. Further, the apparatus is configured to determine an access unit removal time in a current picture of a plurality of pictures of the output video from a coded picture buffer. The apparatus is configured to determine whether the access unit removal time of the current picture from the coded picture buffer is determined using coded picture buffer delay offset information.
[0024] Further, a video data stream is provided. The video data stream has a video encoded therein. The video data stream comprises coded picture buffer delay offset information.
[0025] Further, a video decoder for receiving a video data stream having a video stored therein is provided. The video decoder is configured to decode the video from the video data stream. Further, the video decoder is configured to decode the video depending on an access unit removal time of a current picture of a plurality of pictures of the video from a coded picture buffer. The video decoder is configured to decode the video depending on an indication indicating whether the access unit removal time of the current picture from the coded picture buffer is determined using coded picture buffer delay offset information.
[0026] Further, a method for receiving one or more input video data streams is provided. Each of the one or more input video data streams has an input video encoded therein. The method comprises generating an output video data stream from the one or more input video data streams, the output video data stream encoding an output video, wherein generating the output video data stream is performed such that the output video is the input video encoded in one of the one or more input video data streams or such that the output video depends on the input video of at least one of the one or more input video data streams. Further, the method comprises determining an access unit removal time in a current picture of a plurality of pictures of the output video from a coded picture buffer. Further, the method comprises determining whether the access unit removal time of the current picture from the coded picture buffer is determined using coded picture buffer delay offset information.
[0027] Furthermore, a method for encoding a video into a video data stream according to an embodiment is provided. The method comprises generating the video data stream such that the video data stream comprises coded picture buffer delay offset information.
[0028] Furthermore, a method for receiving a video data stream having a video stored is provided. The method comprises decoding the video from the video data stream. Decoding the video depends on an access unit removal time of a current picture from a plurality of pictures of the video from a coded picture buffer. Furthermore, decoding the video is performed depending on an indication indicating whether to use coded picture buffer delay offset information for determining the access unit removal time of the current picture from the coded picture buffer.
[0029] Furthermore, a computer program for implementing one of the above-mentioned methods when the computer program is executed by a computer or signal processor is provided.
[0030] According to a third aspect of the present application, a video data stream is provided. The video data stream has a video encoded therein. Furthermore, the video data stream comprises an initial coded picture buffer removal delay. Furthermore, the video data stream comprises an initial coded picture buffer removal offset. Furthermore, the video data stream comprises information indicating whether a sum of the initial coded picture buffer removal delay and the initial coded picture buffer removal offset is defined to be constant over two or more buffering periods.
[0031] Furthermore, a video encoder is provided. The video encoder is configured to encode a video into a video data stream. Furthermore, the video encoder is configured to generate the video data stream such that the video data stream comprises an initial coded picture buffer removal delay. Furthermore, the video encoder is configured to generate the video data stream such that the video data stream comprises an initial coded picture buffer removal offset. Furthermore, the video encoder (100) is configured to generate the video data stream such that the video data stream comprises information indicating whether a sum of the initial coded picture buffer removal delay and the initial coded picture buffer removal offset is defined to be constant over two or more buffering periods.
[0032] Further, an apparatus for receiving two input video data streams is provided, the two input video data streams being a first input video data stream and a second input video data stream. An input video is encoded in each of the two input video data streams. The apparatus is configured to generate an output video data stream from the two input video data streams, the output video data stream encoding an output video, wherein the apparatus is configured to generate the output video data stream by concatenating the first input video data stream and the second input video data stream. Further, the apparatus is configured to generate the output video data stream such that the output video data stream comprises an initial coded picture buffer removal delay. Further, the apparatus is configured to generate the output video data stream such that the output video data stream comprises an initial coded picture buffer removal offset. Further, the apparatus is configured to generate the output video data stream such that the output video data stream comprises information indicating whether a sum of the initial coded picture buffer removal delay and the initial coded picture buffer removal offset is defined to be constant over two or more buffering periods.
[0033] Further, a video decoder for receiving a video data stream having a video stored therein is provided. The video decoder is configured to decode the video from the video data stream. Further, the video data stream comprises an initial coded picture buffer removal delay. Further, the video data stream comprises an initial coded picture buffer removal offset. Further, the video data stream comprises information indicating whether a sum of the initial coded picture buffer removal delay and the initial coded picture buffer removal offset is defined to be constant over two or more buffering periods. Further, the video decoder is configured to decode the video depending on the information indicating whether the sum of the initial coded picture buffer removal delay and the initial coded picture buffer removal offset is defined to be constant over two or more buffering periods.
[0034] Further, a method of encoding a video into a video data stream is provided. The method comprises generating the video data stream such that the video data stream comprises an initial coded picture buffer removal delay. The method comprises generating the video data stream such that the video data stream comprises an initial coded picture buffer removal offset. Further, the method comprises generating the video data stream such that the video data stream comprises information indicating whether a sum of the initial coded picture buffer removal delay and the initial coded picture buffer removal offset is defined to be constant over two or more buffering periods.
[0035] Further, a method for receiving two input video data streams is provided, the two input video data streams being a first input video data stream and a second input video data stream. An input video is encoded in each of the two input video data streams. The method comprises generating an output video data stream from the two input video data streams, the output video data stream encoding an output video, wherein the apparatus is configured to generate the output video data stream by concatenating the first input video data stream and the second input video data stream. Further, the method comprises generating the output video data stream such that the output video data stream comprises an initial coded picture buffer removal delay. Further, the method comprises generating the output video data stream such that the output video data stream comprises an initial coded picture buffer removal offset. Further, the method comprises generating the output video data stream such that the output video data stream comprises information indicating whether a sum of the initial coded picture buffer removal delay and the initial coded picture buffer removal offset is defined to be constant over two or more buffering periods.
[0036] Further, a method for receiving a video data stream storing a video is provided. The method comprises decoding the video from the video data stream. The video data stream comprises an initial coded picture buffer removal delay. Further, the video data stream comprises an initial coded picture buffer removal offset. Further, the video data stream comprises information indicating whether a sum of the initial coded picture buffer removal delay and the initial coded picture buffer removal offset is defined to be constant over two or more buffering periods. The method comprises decoding the video depending on the information indicating whether a sum of the initial coded picture buffer removal delay and the initial coded picture buffer removal offset is defined to be constant over two or more buffering periods.
[0037] Further, a computer program is provided for implementing one of the above methods when the computer program is executed by a computer or signal processor.
[0038] According to a fourth aspect of the present application, a video data stream is provided. The video data stream has encoded video. Furthermore, the video data stream comprises an indication indicating whether a non-scalable nested picture timing supplemental enhancement information message of a network abstraction layer unit of one access unit of a plurality of access units of one coded video sequence of one or more coded video sequences of the video data stream is defined to apply to all output layer sets of the plurality of output layer sets of the access unit. If the indication has a first value, the non-scalable nested picture timing supplemental enhancement information message of the network abstraction layer unit of the access unit is defined to apply to all output layer sets of the plurality of output layer sets of the access unit. If the indication has a value different from the first value, the indication does not define whether the non-scalable nested picture timing supplemental enhancement information message of the network abstraction layer unit of the access unit applies to all output layer sets of the plurality of output layer sets of the access unit.
[0039] Furthermore, a video encoder is provided. The video encoder is configured to encode video into a video data stream. Furthermore, the video encoder is configured to generate the video data stream such that the video data stream comprises an indication indicating whether a non-scalable nested picture timing supplemental enhancement information message of a network abstraction layer unit of one access unit of a plurality of access units of one coded video sequence of one or more coded video sequences of the video data stream is defined to apply to all output layer sets of the plurality of output layer sets of the access unit. If the indication has a first value, the non-scalable nested picture timing supplemental enhancement information message of the network abstraction layer unit of the access unit is defined to apply to all output layer sets of the plurality of output layer sets of the access unit. If the indication has a value different from the first value, the indication does not define whether the non-scalable nested picture timing supplemental enhancement information message of the network abstraction layer unit of the access unit applies to all output layer sets of the plurality of output layer sets of the access unit.
[0040] Furthermore, an apparatus for receiving an input video data stream is provided. The input video data stream has a video encoded therein. The apparatus is configured to generate a processed video data stream from the input video data stream. Furthermore, the apparatus is configured to generate the processed video data stream such that the processed video data stream comprises an indication indicating whether a non-scalable nested picture timing supplemental enhancement information message of a network abstraction layer unit of one access unit of one coded video sequence of one or more coded video sequences of the processed video data stream is defined to be applied to all output layer sets of a plurality of output layer sets of the access unit. If the indication has a first value, the non-scalable nested picture timing supplemental enhancement information message of the network abstraction layer unit of the access unit is defined to be applied to all output layer sets of the plurality of output layer sets of the access unit. If the indication has a value different from the first value, the indication does not define whether the non-scalable nested picture timing supplemental enhancement information message of the network abstraction layer unit of the access unit is applicable to all output layer sets of the plurality of output layer sets of the access unit.
[0041] Furthermore, a video decoder for receiving a video data stream having a video stored therein is provided. The video decoder is configured to decode the video from the video data stream. The video data stream comprises an indication indicating whether a non-scalable nested picture timing supplemental enhancement information message of a network abstraction layer unit of one access unit of one coded video sequence of one or more coded video sequences of the video data stream is defined to be applied to all output layer sets of a plurality of output layer sets of the access unit. If the indication has a first value, the non-scalable nested picture timing supplemental enhancement information message of the network abstraction layer unit of the access unit is defined to be applied to all output layer sets of the plurality of output layer sets of the access unit. If the indication has a value different from the first value, the indication does not define whether the non-scalable nested picture timing supplemental enhancement information message of the network abstraction layer unit of the access unit is applicable to all output layer sets of the plurality of output layer sets of the access unit.
[0042] Furthermore, a method of encoding a video into a video data stream is provided. The method comprises generating the video data stream such that the video data stream comprises an indication indicating whether a non-scalable nested picture timing supplemental enhancement information message of a network abstraction layer unit of one access unit of one coded video sequence of one or more coded video sequences of the video data stream is defined to apply to all output layer sets of a plurality of output layer sets of the access unit. If the indication has a first value, the non-scalable nested picture timing supplemental enhancement information message of the network abstraction layer unit of the access unit is defined to apply to all output layer sets of the plurality of output layer sets of the access unit. If the indication has a value different from the first value, the indication does not define whether the non-scalable nested picture timing supplemental enhancement information message of the network abstraction layer unit of the access unit applies to all output layer sets of the plurality of output layer sets of the access unit.
[0043] Furthermore, a method for receiving an input video data stream is provided. The input video data stream has a video encoded therein. The method comprises generating a processed video data stream from the input video data stream. Furthermore, the method comprises generating the processed video data stream such that the processed video data stream comprises an indication indicating whether a non-scalable nested picture timing supplemental enhancement information message of a network abstraction layer unit of one access unit of one coded video sequence of one or more coded video sequences of the processed video data stream is defined to apply to all output layer sets of a plurality of output layer sets of the access unit. If the indication has a first value, the non-scalable nested picture timing supplemental enhancement information message of the network abstraction layer unit of the access unit is defined to apply to all output layer sets of the plurality of output layer sets of the access unit. If the indication has a value different from the first value, the indication does not define whether the non-scalable nested picture timing supplemental enhancement information message of the network abstraction layer unit of the access unit applies to all output layer sets of the plurality of output layer sets of the access unit.
[0044] Furthermore, a method for receiving a video data stream having a video stored is provided. The method comprises decoding the video from the video data stream. The video data stream comprises an indication, the indication indicating whether a non-scalable nested picture timing supplemental enhancement information message of a network abstraction layer unit of one access unit of one coded video sequence of one or more coded video sequences of the video data stream is defined to apply to all output layer sets of a plurality of output layer sets of the access unit. If the indication has a first value, the non-scalable nested picture timing supplemental enhancement information message of the network abstraction layer unit of the access unit is defined to apply to all output layer sets of the plurality of output layer sets of the access unit. If the indication has a value different from the first value, the indication does not define whether the non-scalable nested picture timing supplemental enhancement information message of the network abstraction layer unit of the access unit applies to all output layer sets of the plurality of output layer sets of the access unit.
[0045] Furthermore, a computer program for implementing one of the above methods when the computer program is executed by a computer or signal processor is provided.
[0046] According to a fifth aspect of the present application, a video data stream is provided. The video data stream has a video encoded therein. Furthermore, the video data stream comprises one or more scalable-nested supplemental enhancement information messages. The one or more scalable-nested supplemental enhancement information messages comprise a plurality of syntax elements. Each syntax element of one or more syntax elements of the plurality of syntax elements is defined to have a same size in each scalable-nested supplemental enhancement information message of scalable-nested supplemental enhancement information messages of the video data stream or a portion of the video data stream.
[0047] Furthermore, a video encoder is provided. The video encoder is configured to encode a video into a video data stream. Furthermore, the video encoder is configured to generate the video data stream such that the video data stream comprises one or more scalable-nested supplemental enhancement information messages. Furthermore, the video encoder is configured to generate the video data stream such that the one or more scalable-nested supplemental enhancement information messages comprise a plurality of syntax elements. Furthermore, the video encoder is configured to generate the video data stream such that each syntax element of one or more syntax elements of the plurality of syntax elements is defined to have a same size in each scalable-nested supplemental enhancement information message of scalable-nested supplemental enhancement information messages of the video data stream or a portion of the video data stream.
[0048] Furthermore, an apparatus for receiving an input video data stream is provided. The input video data stream has video encoded therein. The apparatus is configured to generate an output video data stream from the input video data stream. The video data stream comprises one or more scalable-nested supplemental enhancement information messages. The one or more scalable-nested supplemental enhancement information messages comprise a plurality of syntax elements. Each syntax element of one or more syntax elements of the plurality of syntax elements is defined to have a same size in each scalable-nested supplemental enhancement information message of the video data stream or a portion of the video data stream. The apparatus is configured to process the one or more scalable-nested supplemental enhancement information messages.
[0049] Furthermore, a video decoder for receiving a video data stream having video stored therein is provided. The video decoder is configured to decode the video from the video data stream. The video data stream comprises one or more scalable-nested supplemental enhancement information messages. The one or more scalable-nested supplemental enhancement information messages comprise a plurality of syntax elements. Each syntax element of one or more syntax elements of the plurality of syntax elements is defined to have a same size in each scalable-nested supplemental enhancement information message of the video data stream or a portion of the video data stream. The video decoder is configured to decode the video depending on the one or more syntax elements of the plurality of syntax elements.
[0050] Furthermore, a method of encoding video into a video data stream is provided. The method comprises generating the video data stream such that the video data stream comprises one or more scalable-nested supplemental enhancement information messages. Furthermore, the method comprises generating the video data stream such that the one or more scalable-nested supplemental enhancement information messages comprise a plurality of syntax elements. Furthermore, the method comprises generating the video data stream such that each syntax element of one or more syntax elements of the plurality of syntax elements is defined to have a same size in each scalable-nested supplemental enhancement information message of the video data stream or a portion of the video data stream.
[0051] Furthermore, a method for receiving an input video data stream is provided. The input video data stream has video encoded therein. The method comprises generating an output video data stream from the input video data stream. The video data stream comprises one or more scalable-nested supplemental enhancement information messages. The one or more scalable-nested supplemental enhancement information messages comprise a plurality of syntax elements. Each syntax element of one or more syntax elements of the plurality of syntax elements is defined to have a same size in each scalable-nested supplemental enhancement information message of the video data stream or a portion of the video data stream. The method comprises processing the one or more scalable-nested supplemental enhancement information messages.
[0052] Furthermore, a method for receiving a video data stream having a video stored therein is provided. The method comprises decoding the video from the video data stream. The video data stream comprises one or more scalable-nested supplemental enhancement information messages. The one or more scalable-nested supplemental enhancement information messages comprise a plurality of syntax elements. Each syntax element of one or more syntax elements of the plurality of syntax elements is defined to have a same size in each scalable-nested supplemental enhancement information message of the video data stream or a portion of the video data stream. The video is decoded depending on the one or more syntax elements of the plurality of syntax elements.
[0053] Furthermore, a computer program for implementing one of the above methods when the computer program is executed by a computer or signal processor is provided.
[0054] Preferred embodiments are provided in dependent claims. BRIEF DESCRIPTION OF DRAWINGS
[0055] In the following, embodiments of the present application will be described in detail with reference to the attached drawings, in which:
[0056] Figure 1 A video encoder for encoding a video into a video data stream according to an embodiment is shown.
[0057] Figure 2 An apparatus for receiving an input video data stream according to an embodiment is shown.
[0058] Figure 3 A video decoder for receiving a video data stream having a video stored therein according to an embodiment is shown.
[0059] Figure 4 A raw bitstream (top) and a bitstream after discarding pictures (bottom) according to an embodiment are shown. Figure 4 Figure 4
[0060] Figure 5 A concatenation of two bitstreams after pictures have been discarded from one of the two bitstreams according to an embodiment is shown.
[0061] Figure 6 A concatenation of two bitstreams according to another embodiment is shown.
[0062] Figure 7 Two sets of HRD SEI (scalable-nested SEI and non-scalable-nested SEI) in a two-layer bitstream according to an embodiment are shown.
[0063] Figure 8 A video encoder is shown.
[0064] Figure 9 A video decoder is shown.
[0065] Figure 10 The relationship between, on the one hand, a reconstructed signal (e.g. a reconstructed picture) and, on the other hand, a combination of a prediction residual signal and a prediction signal signaled in a data stream is shown. DETAILED DESCRIPTION
[0066] The following description of the drawings starts with a presentation of a description of an encoder and a decoder for block-based prediction coding of video pictures in order to form an example of an encoding framework into which embodiments of the present application can be built in. The respective encoder and decoder are described with respect to Figures 8 to 10 are described. In the following, the description of embodiments of the inventive concept is presented together with a description of how this concept can be built into Figure 8 and Figure 9 encoders and decoders, although the embodiments described using Figures 1 to 3 and in the following can also be used to form encoders and decoders operating under an encoding framework not according to Figure 8 and Figure 9 encoders and decoders.
[0067] Figure 8 A video encoder is shown, an apparatus for predictively encoding a picture 12 into a data stream 14 using, exemplarily, transform-based residual coding. The apparatus or encoder is indicated using reference sign 10. Figure 9 A corresponding video decoder 20 is shown, e.g. an apparatus 20 configured to predictively decode a picture 12' from the data stream 14 also using transform-based residual decoding, wherein the prime is used to indicate that the picture 12' reconstructed by the decoder 20 deviates from the picture 12 originally encoded by the apparatus 10 in terms of encoding losses introduced by quantization of the prediction residual signal. Figure 8 and Figure 9 uses transform-based prediction residual coding, although embodiments of the present application are not limited to such prediction residual coding. As will be outlined in the following, this is also true for other details described with respect to Figure 8 and Figure 9 .
[0068] The encoder 10 is configured to spatial-to-spectral transform the prediction residual signal and to encode the thus obtained prediction residual signal into the data stream 14. Likewise, the decoder 20 is configured to decode the prediction residual signal from the data stream 14 and to perform a spectral-to-spatial transformation on the thus obtained prediction residual signal.
[0069] Internally, the encoder 10 can comprise a prediction residual signal former 22 which generates a prediction residual 24 to measure a deviation of a prediction signal 26 from an original signal, e.g. from the picture 12. The prediction residual signal former 22 can for example be a subtractor which subtracts the prediction signal from the original signal, e.g. from the picture 12. The encoder 10 then further comprises a transformer 28 which spatially to spectrally transforms the prediction residual signal 24 to obtain a spectral domain prediction residual signal 24', which is then quantized by a quantizer 32, which is also comprised in the encoder 10. The thus quantized prediction residual signal 24" is encoded into the bitstream 14. To this end, the encoder 10 can optionally comprise an entropy encoder 34 which entropy encodes the prediction residual signal 24" which is transformed and quantized into the data stream 14. The prediction signal 26 is generated by a prediction stage 36 of the encoder 10 on the basis of the prediction residual signal 24" which is encoded into the data stream 14 and decodable from the data stream 14. To this end, as Figure 8 indicated, the prediction stage 36 can internally comprise a dequantizer 38 which dequantizes the prediction residual signal 24" in order to obtain a spectral domain prediction residual signal 24"' corresponding to the signal 24' except for a quantization loss, followed by an inverse transformer 40 which inverse transforms, e.g. spectrally to spatially, the latter prediction residual signal 24"' in order to obtain a prediction residual signal 24"" corresponding to the original prediction residual signal 24 except for a quantization loss. A combiner 42 of the prediction stage 36 then recombines the prediction signal 26 and the prediction residual signal 24"" by addition, e.g., to obtain a reconstructed signal 46, e.g. a reconstruction of the original signal 12. The reconstructed signal 46 can correspond to the signal 12'. A prediction module 44 of the prediction stage 36 then generates the prediction signal 26 on the basis of the signal 46 by using, e.g., spatial prediction, e.g. intra-picture prediction, and / or temporal prediction, e.g. inter-picture prediction.
[0070] Likewise, as Figure 9 indicated, the decoder 20 can internally comprise components corresponding to and interconnected in a manner corresponding to the prediction stage 36. In particular, an entropy decoder 50 of the decoder 20 can entropy decode the quantized spectral domain prediction residual signal 24" from the data stream, followed by a dequantizer 52, an inverse transformer 54, a combiner 56 and a prediction module 58 interconnected and cooperating in a manner described above with respect to the prediction stage 36 to recover a reconstructed signal on the basis of the prediction residual signal 24"', as Figure 9 indicated, the output of the combiner 56 yields the reconstructed signal, i.e. the picture 12'.
[0071] Although not specifically described above, it is clear that the encoder 10 can set coding parameters including, for example, prediction modes, motion parameters, etc., in accordance with some optimization scheme, e.g., in a way that optimizes some rate and distortion related criterion, e.g., the coding cost. For example, the encoder 10 and the decoder 20 and the respective corresponding modules 44, 58 can support different prediction modes such as intra coding modes and inter coding modes. The granularity at which the encoder and the decoder switch between these prediction mode types can correspond to a respective subdivision of the pictures 12 and 12' into coding segments or coding blocks. For example, in units of these coding segments, the pictures can be subdivided into blocks that are intra coded and blocks that are inter coded. As outlined in more detail below, intra coded blocks are predicted based on a spatial, already coded / decoded neighborhood of the respective block. Several intra coding modes can exist and be selected for respective intra coded segments including directional or angular intra coding modes according to which the respective segment is filled by extrapolating sample values of the neighborhood into the respective intra coded segment along a specific direction that is specific to the respective directional intra coding mode. For example, the intra coding modes can also include one or more other modes such as a DC coding mode according to which the prediction for the respective intra coded block assigns a DC value to all samples within the respective intra coded segment and / or a planar intra coding mode according to which the prediction for the respective block is approximated or determined as a spatial distribution of sample values described by a two-dimensional linear function over the sample positions of the respective intra coded block, wherein the driving tilt and offset of the plane are defined by the two-dimensional linear function based on neighboring samples. In contrast, inter coded blocks can be predicted, e.g., in time. For inter coded blocks, a motion vector can be signaled within the data stream that indicates a spatial displacement of a portion of a previously coded picture of the video to which the picture 12 belongs, from which portion of the previously coded / decoded picture the picture is sampled to obtain a prediction signal for the respective inter coded block. This means that, in addition to the residual signal coding included in the data stream 14, e.g., the entropy encoded transform coefficient levels representing the quantized spectral domain prediction residual signal 24'', the data stream 14 can have coded therein coding mode parameters for assigning coding modes to the various blocks, prediction parameters for some of the blocks, e.g., motion parameters for inter coded segments, and optionally other parameters, e.g., parameters for controlling and signaling the subdivision of the pictures 12 and 12', respectively, into segments. The decoder 20 uses these parameters to subdivide the pictures in the same way, to assign the same prediction modes to the segments, and to perform the same predictions to produce the same prediction signals.
[0072] Figure 10 The combination between the reconstructed signal, e.g., the reconstructed picture 12', of one aspect and the prediction residual signal 24'''' and the prediction signal 26 signaled in the data stream of the other aspect is shown. As already described above, this combination can be an addition. The prediction signal 26 is in the data stream 14, e.g., in the form of a prediction residual signal 24''', which is added to the quantized transform coefficients 24' of the prediction residual signal 24 to obtain the prediction signal 26.Figure 10 The picture area is shown to be subdivided into intra coded blocks, which are illustratively indicated using hatching, and inter coded blocks, which are illustratively indicated without hatching. The subdivision can be an arbitrary subdivision, e.g. a regular subdivision of the picture area into square blocks or rows and columns of non-square blocks, or a recursive multi-tree subdivision of the picture 12 from a tree root block into a plurality of leaf blocks of variable size, e.g. a quad-tree subdivision, etc., wherein, Figure 10 A hybrid thereof is shown, wherein the picture area is first subdivided into rows and columns of tree root blocks, and then further subdivided into one or more leaf blocks according to a recursive multi-tree subdivision.
[0073] Likewise, the data stream 14 can have intra coding modes encoded therein for the intra coded blocks 80, which assign one of a number of supported intra coding modes to a respective intra coded block 80. For the inter coded blocks 82, the data stream 14 can have one or more motion parameters encoded therein. Generally, the inter coded blocks 82 are not limited to being temporally coded. Alternatively, the inter coded blocks 82 can be any blocks predicted from a previously coded portion other than the current picture 12 itself, e.g. a previously coded picture of a video to which the picture 12 belongs, or a picture of another view or lower level in case the encoder and decoder are a scalable encoder and decoder, respectively.
[0074] Figure 10 The prediction residual signals 24'''' in the data stream 14 are also shown to subdivide the picture area into blocks 84. These blocks can be referred to as transform blocks in order to distinguish them from the coding blocks 80 and 82. In fact, Figure 10 It is shown that the encoder 10 and the decoder 20 can use two different subdivisions, which subdivide the picture 12 and the picture 12', respectively, into blocks, namely one into coding blocks 80 and 82, and the other into transform blocks 84. The two subdivisions can be identical, e.g. each coding block 80 and 82 can at the same time form a transform block 84, but Figure 10 It is shown that the case is possible that the subdivision into transform blocks 84 forms an extension of the subdivision into coding blocks 80 and 82, such that any boundary between two of the blocks 80 and 82 covers a boundary between two of the blocks 84, or in other words, each block 80, 82 coincides with one of the transform blocks 84, or with a cluster of transform blocks 84. However, it is also possible to determine or select the subdivisions independently from each other, such that the transform blocks 84 can alternatively span across block boundaries between the blocks 80, 82. Statements analogous to the ones made with respect to the subdivision into blocks 80, 82 are thus true with respect to the subdivision into transform blocks 84, e.g. the blocks 84 can be the result of a regular subdivision of the picture area into blocks (arranged or not arranged into rows and columns), a recursive multi-tree subdivision of the picture area, or a combination or any other type of blockation. Incidentally, it is noted that the blocks 80, 82 and 84 are not limited to being square, rectangular or any other shape.
[0075] Figure 10 It is also shown that the combination of prediction signal 26 and prediction residual signal 24'''' directly generates the reconstructed signal 12'. However, it should be noted that, according to an alternative embodiment, more than one prediction signal 26 can be combined with prediction residual signal 24'''' to generate image 12'.
[0076] exist Figure 10 In this context, transform block 84 should have the following meaning. Transformer 28 and inverse transformer 54 perform their transforms on a unit basis using these transform blocks 84. For example, many codecs use some kind of DST or DCT for all transform blocks 84. Some codecs allow skipping transforms, such that for some transform blocks 84, the prediction residual signal is directly encoded in the spatial domain. However, according to the embodiments described below, encoder 10 and decoder 20 are configured in a manner that they support several transforms. For example, the transforms supported by encoder 10 and decoder 20 may include:
[0077] oDCT-II (or DCT-III), where DCT stands for Discrete Cosine Transform
[0078] oDST-IV, where DST represents Discrete Sine Transform
[0079] oDCT-IV
[0080] oDST-VII
[0081] oIdentity Transformation (IT)
[0082] Of course, while transformer 28 will support all forward transform versions of these transforms, decoder 20 or inverse transformer 54 will support their corresponding backward or inverse versions:
[0083] o Inverse DCT-II (or anti-DCT-III)
[0084] oInverse DST-IV
[0085] oReverse DCT-IV
[0086] oInverse DST-VII
[0087] oIdentity Transformation (IT)
[0088] The following description provides further details about which transforms the encoder 10 and decoder 20 can support. In any case, it should be noted that the set of supported transforms may include only one transform, such as a spectrum-to-space or space-to-spectrum transform.
[0089] As already outlined above, Figures 8 to 10Examples have been presented where the inventive concepts described further below can be implemented to form specific examples of encoders and decoders according to this application. In this regard, Figure 8 and Figure 9 The encoder and decoder can represent possible implementations of the encoder and decoder described below, respectively. However, Figure 8 and Figure 9 This is merely an example. However, the encoder according to embodiments of this application can perform block-based encoding of image 12 using the more detailed overview below, and differs from... Figure 8 The encoder, for example, differs in that it is not a video encoder but a still image encoder, does not support inter-frame prediction, or is subdivided into 80 blocks to match... Figure 10 The methods illustrated in the examples are different from those described above. Similarly, the decoder according to embodiments of this application can perform block-based decoding of image 12' from data stream 14 using the coding concepts further outlined below, but can be performed in a manner different from, for example... Figure 9 The decoder 20 differs from the others in that it is not a video decoder but a still image decoder, and does not support intra-frame prediction, or anything related to... Figure 10 The description may differ in how image 12' is subdivided into blocks, and / or the decoder may derive the prediction residuals from data stream 14 in the spatial domain rather than in the transform domain.
[0090] Figure 1 A video encoder 100 for encoding video into a video data stream according to an embodiment is shown. The video encoder 100 is configured to generate a video data stream such that the video data stream includes an indication of whether video should be output before relying on random access images.
[0091] Figure 2 An apparatus 200 for receiving an input video data stream according to an embodiment is shown. The input video data stream contains encoded video. The apparatus 200 is configured to generate an output video data stream from the input video data stream.
[0092] Figure 3 A video decoder 300 according to an embodiment is shown for receiving a video data stream in which video is stored. The video decoder 300 is configured to decode video from the video data stream. The video decoder 300 is configured to decode the video based on an indication of whether video should be output before a picture that depends on random access.
[0093] Furthermore, a system according to an embodiment is provided. The system includes... Figure 2 The device and Figure 3 The video decoder. Figure 3 The video decoder 300 is configured to receive Figure 2The output video data stream of the device. Figure 3 The video decoder 300 is configured to... Figure 2 The device 200 decodes video in the output video data stream.
[0094] In an embodiment, the system may include, for example, further Figure 1 100 video encoders. Figure 2 The device 200 can, for example, be configured to... Figure 1 The video encoder 100 receives video data streams as input video data streams.
[0095] The (optional) intermediate device 210 of apparatus 200 may be configured, for example, to receive a video data stream from video encoder 100 as an input video data stream and to generate an output video data stream from the input video data stream. For example, the intermediate device may be configured to modify the (header / metadata) information of the input video data stream, and / or may be configured to remove images from the input video data stream, and / or may be configured to mix / concatenate the input video data stream with an additional second bitstream having a second video encoded therein.
[0096] (Optional) The video decoder 221 can be configured, for example, to decode video from the output video data stream.
[0097] (Optional) Assume that the reference decoder 222 can be configured, for example, to determine the timing information of the video based on the output video data stream, or can be configured, for example, to determine the buffer information of the buffer to store the video or a portion of the video.
[0098] The system includes Figure 1 Video encoder 101 and Figure 2 The video decoder 151.
[0099] Video encoder 101 is configured to generate an encoded video signal. Video decoder 151 is configured to decode the encoded video signal to reconstruct images of the video.
[0100] The first aspect of the invention will now be described in detail below.
[0101] According to a first aspect of the invention, an apparatus 200 is provided for receiving an input video data stream. The input video data stream contains encoded video. The apparatus 200 is configured to generate an output video data stream from the input video data stream. Furthermore, the apparatus 200 is used to determine whether the video should be output before a random access image.
[0102] According to an embodiment, the device 200 may be configured, for example, to determine a first variable (e.g., NoOutputBeforeDrapFlag) that indicates whether the video should be output before the images that rely on random access.
[0103] In an embodiment, the apparatus 200 may be configured, for example, to generate an output video data stream such that the output video data stream may include, for example, an indication of whether the video should be output before the images that depend on random access.
[0104] According to an embodiment, the apparatus 200 may be configured, for example, to generate an output video data stream such that the output video data stream may include, for example, supplementary enhancement information, including, for example, an indication of whether the video should be output before the images that rely on random access.
[0105] In an embodiment, the image preceding the dependent random access image in the video can be an independent random access image. The apparatus 200 can, for example, be configured to generate an output video data stream such that the output video data stream may include, for example, a flag (e.g., ph_pic_output_flag) with a predefined value (e.g., 0) in the image header of the independent random access image, such that the predefined value (e.g., ph_pic_output_flag) of the flag (e.g., ph_pic_output_flag) can, for example, indicate that the independent random access image is directly preceding the dependent random access image in the video data stream and should not be output.
[0106] According to an embodiment, the flag may be, for example, a first flag, wherein the device 200 may be configured to generate an output video data stream such that the output video data stream may include, for example, another flag in the picture parameter set of the video data stream, wherein the other flag may indicate, for example, whether the first flag (e.g., ph_pic_output_flag) exists in the picture header of an independently randomly accessed picture.
[0107] In an embodiment, the device 200 may be configured, for example, to generate an output video data stream, such that the output video data stream may include, for example, the following flags as indicators of whether the video should be output before the random access image: a supplemental enhancement information flag in the supplemental enhancement information of the output video data stream, or a picture parameter set flag in the picture parameter set of the output video data stream, or a sequence parameter set flag in the sequence parameter set of the output video data stream, or an external device flag, wherein the value of the external device flag may be set, for example, by an external unit located outside the device 200.
[0108] According to an embodiment, the device 200 may be configured, for example, to determine the value of a second variable (e.g., PictureOutputFlag) of the video before the random access image based on a first variable (e.g., NoOutputBeforeDrapFlag), wherein the second variable (e.g., PictureOutputFlag) may, for example, indicate whether the image should be output for the image, and wherein the device 200 may, for example, be configured to output or not output the image based on the second variable (e.g., PictureOutputFlag).
[0109] In this embodiment, the images in the video preceding the random access images can be independent random access images. The first variable (e.g., NoOutputBeforeDrapFlag) could, for example, indicate that independent random access images should not be output.
[0110] According to an embodiment, the images in the video preceding the random access images can be independent random access images. The device 200 can, for example, be configured to set a first variable (e.g., NoOutputBeforeDrapFlag) such that the first variable (e.g., NoOutputBeforeDrapFlag) can, for example, indicate that independent random access images should be output.
[0111] In an embodiment, the device 200 may be configured, for example, to signal to the video decoder 300 whether the video should be output before the images that rely on random access.
[0112] In addition, a video data stream is provided. This video data stream contains encoded video. The video data stream includes an indication of whether the video should be output before the images obtained through random access.
[0113] According to an embodiment, the video data stream includes supplemental enhancement information, which includes, for example, an indication of whether the video should be output before relying on randomly accessed images.
[0114] In one embodiment, the image preceding the dependent random access image in the video can be an independent random access image. The video data stream may, for example, include a flag (e.g., ph_pic_output_flag) with a predefined value (e.g., 0) in the image header of the independent random access image, such that the predefined value (e.g., ph_pic_output_flag) of the flag (e.g., ph_pic_output_flag) indicates, for example, that the independent random access image occurs directly before the dependent random access image within the video data stream and should not be output.
[0115] According to an embodiment, the flag may be, for example, a first flag, wherein the video data stream may include, for example, another flag in the picture parameter set of the video data stream, wherein the other flag may, for example, indicate whether the first flag (e.g., ph_pic_output_flag) exists in the picture header of an independently randomly accessed picture.
[0116] In an embodiment, the video data stream may include, for example, the following flags as indicators of whether the video should be output before the random access image: a supplemental enhancement information flag in the supplemental enhancement information of the output video data stream, or a picture parameter set flag in the picture parameter set of the output video data stream, or a sequence parameter set flag in the sequence parameter set of the output video data stream, or an external device flag, wherein the value of the external device flag may be set by an external unit located outside the device 200, for example.
[0117] In addition, a video encoder 100 is provided. This video encoder 100 can, for example, be configured to encode video into a video data stream. Furthermore, the video encoder 100 can, for example, be configured to generate a video data stream such that the video data stream includes an indication of whether video should be output before relying on random access images.
[0118] According to an embodiment, the video encoder 100 may be configured, for example, to generate a video data stream such that the video data stream may include, for example, supplemental enhancement information, including, for example, an indication of whether the video should be output before the random access images.
[0119] In an embodiment, the image preceding the dependent random access image in the video can be an independent random access image. The video encoder 100 can, for example, be configured to generate a video data stream such that the video data stream may include, for example, a flag (e.g., ph_pic_output_flag) with a predefined value (e.g., 0) in the image header of the independent random access image, such that the predefined value (e.g., ph_pic_output_flag) of the flag (e.g., ph_pic_output_flag) can, for example, indicate that the independent random access image is directly preceding the dependent random access image within the video data stream, and that the independent random access image should not be output.
[0120] According to an embodiment, the flag may be, for example, a first flag, wherein the video encoder 100 may be configured to generate a video data stream such that the video data stream may include, for example, another flag in the picture parameter set of the video data stream, wherein the other flag may indicate, for example, whether the first flag (e.g., ph_pic_output_flag) exists in the picture header of an independently randomly accessed picture.
[0121] In an embodiment, the video encoder 100 may be configured, for example, to generate a video data stream such that the video data stream may include, for example, the following flags as indicators of whether the video should be output before the random access image: a supplemental enhancement information flag in the supplemental enhancement information of the output video data stream, or a picture parameter set flag in the picture parameter set of the output video data stream, or a sequence parameter set flag in the sequence parameter set of the output video data stream, or an external device flag, wherein the value of the external device flag may be set, for example, by an external unit located outside the device 200.
[0122] In addition, a video decoder 300 is provided for receiving a video data stream in which video is stored. The video decoder 300 is configured to decode video from the video data stream. The video decoder 300 is configured to decode the video based on an indication of whether video should be output before a picture that depends on random access.
[0123] According to an embodiment, the video decoder 300 may be configured, for example, to decode the video based on a first variable (e.g., NoOutputBeforeDrapFlag), which indicates whether the video should output images prior to those relying on random access.
[0124] In one embodiment, the video data stream may include, for example, an indication of whether the video should be output before the random access images. The video decoder 300 may be configured, for example, to decode the video based on the indication within the video data stream.
[0125] According to an embodiment, the video data stream includes supplemental enhancement information, which includes, for example, an indication of whether the video should be output before the random access images. The video decoder 300 may be configured, for example, to decode the video based on the supplemental enhancement information.
[0126] In one embodiment, the image preceding the dependent random access image in the video can be an independent random access image. The video data stream may, for example, include a flag (e.g., ph_pic_output_flag) with a predefined value (e.g., 0) in the header of the independent random access image, such that the predefined value (e.g., ph_pic_output_flag) of the flag (e.g., ph_pic_output_flag) indicates, for example, that the independent random access image precedes the dependent random access image directly within the video data stream and should not be output. The video decoder 300 may, for example, be configured to decode the video based on this flag.
[0127] According to an embodiment, the flag may be, for example, a first flag, wherein the video data stream may include, for example, another flag in the picture parameter set of the video data stream, wherein the other flag may, for example, indicate whether the first flag (e.g., ph_pic_output_flag) exists in the picture header of an independently accessed random picture. The video decoder 300 may, for example, be configured to decode the video based on the other flag.
[0128] In an embodiment, the video data stream may include, for example, the following flags as indicators of whether the video should be output before the random access image: a supplemental enhancement information flag within the supplemental enhancement information of the output video data stream, or a picture parameter set flag within the picture parameter set of the output video data stream, or a sequence parameter set flag within the sequence parameter set of the output video data stream, or an external device flag, wherein the value of the external device flag may be set, for example, by an external unit located outside of device 200. The video decoder 300 may be configured, for example, to decode the video according to the indications within the video data stream.
[0129] According to an embodiment, the video decoder 300 may be configured, for example, to reconstruct the video from the video data stream. The video decoder 300 may be configured, for example, to output or not output images of the video preceding the randomly accessed images, based on a first variable (e.g., NoOutputBeforeDrapFlag).
[0130] In an embodiment, the video decoder 300 may be configured, for example, to determine the value of a second variable (e.g., PictureOutputFlag) of the video before the random access image based on a first variable (e.g., NoOutputBeforeDrapFlag), wherein the second variable (e.g., PictureOutputFlag) may, for example, indicate whether the image should be output for the image, and wherein the device 200 may, for example, be configured to output or not output the image based on the second variable (e.g., PictureOutputFlag).
[0131] According to an embodiment, the images in the video preceding the dependent random access images can be independent random access images. The video decoder 300 can, for example, be configured to decode the video based on a first variable (e.g., NoOutputBeforeDrapFlag) that may indicate that independent random access images should not be output.
[0132] In an embodiment, the images in the video preceding the dependent random access images can be independent random access images. The video decoder 300 can, for example, be configured to decode the video based on a first variable (e.g., NoOutputBeforeDrapFlag) that may indicate that an independent random access image should be output.
[0133] In addition, a system is provided. This system includes the device 200 as described above and the video decoder 300 as described above. The video decoder 300 is configured to receive the output video data stream of the device 200. Furthermore, the video decoder 300 is configured to decode video from the output video data stream of the device 200.
[0134] According to an embodiment, the system may also include, for example, a video encoder 100. The device 200 may be configured, for example, to receive a video data stream from the video encoder 100 as an input video data stream.
[0135] Specifically, the first aspect of the invention relates to starting CVS at DRAP and omitting IDR output in decoding and conformance testing.
[0136] When a bitstream includes pictures labeled DRAP (i.e., in the bitstream, only the previous IRAP is used as a reference for the DRAP and from there is used), these DRAP pictures can be used for random access functions with low rate overhead. However, when using a target DRAP to randomly access the stream, it is not expected that any initial pictures preceding the target DRAP (i.e., the associated IRAP of the target DRAP) will be displayed at the decoder output because the time distance between these pictures will cause video playback instability / jitter when played back at the rate of the original video until the video plays back smoothly from the target DRAP.
[0137] Therefore, it is desirable to omit the output of images preceding the DRAP image. This aspect of the invention proposes means for correspondingly controlling the decoder.
[0138] In one embodiment, an external means for setting the PicOutputFlag variable of the IRAP image is available, as follows:
[0139] – If an external means not specified in the specification is available to set the image's variable NoOutputBeforeDrapFlag to a value, then the image's NoOutputBeforeDrapFlag is set to the value provided by the external means.
[0140] [...]
[0141] – The variable PictureOutputFlag for the current image is exported as follows:
[0142] – If sps_video_parameter_set_id is greater than 0 and the current layer is not an output layer (i.e., for any value of i in the range 0 to NumOutputLayersInOls[TargetOlsIdx]-1 (inclusive), nuh_layer_id is not equal to OutputLayerIdInOls[TargetOlsIdx][i]), or one of the following conditions is true, then PictureOutputFlag is set to 0:
[0143] – The current image is a RASL image, and the NoOutputBeforeRecoveryFlag of the associated IRAP image is equal to 1.
[0144] – The current image is either a GDR image with NoOutputBeforeRecoveryFlag equal to 1, or a recovered image of a GDR image with NoOutputBeforeRecoveryFlag equal to 1.
[0145] – The current image is an IRAP image with NoOutputBeforeDrapFlag equal to 1.
[0146] Otherwise, PictureOutputFlag is set to equal ph_pic_output_flag.
[0147] In another embodiment, NoOutputBeforeDrapFlag is set externally only for the first IRAP image in CVS; otherwise, it is set to 0.
[0148] – If an external means not specified in the specification is available to set the NoOutputBeforeDrapFlag variable of an image to a value, then the NoOutputBeforeDrapFlag of the first image in CVS is set to the same value provided by the external means. Otherwise, NoOutputBeforeDrapFlag is set to 0.
[0149] For cases where images are removed between IRAP and DRAP images, the aforementioned flag NoOutputBeforeDrapFlag can also be associated with the use of alternative HRD timings transmitted in the bitstream, such as the UseAltCpbParamsFlag in the VVC specification.
[0150] In an alternative embodiment, a constraint is that the output flag ph_pic_output_flag in the image header should have a value of 0 for the IRAP image that is directly before the DRAP image and has no non-DRAP images in between. In this case, whenever the extractor or player uses DRAP for random access, i.e., when it removes the intermediate image between the IRAP and DRAP from the bitstream, it is also necessary to verify or adjust that the corresponding output flag is set to 0 and the IRAP output is omitted.
[0151] To simplify this operation, the raw bitstream needs to be prepared accordingly. More specifically, `pps_output_flag_present_flag` (which determines the presence of the flag `ph_pic_output_flag` in the image header) should be equal to 1, allowing the image header to be easily modified without needing to change the parameter set. That is:
[0152] The requirement for bitstream consistency is that if a PPS is referenced by an image within a CVSS AU that has an associated DRAP AU, then the value of pps_output_flag_present_flag should be equal to 1.
[0153] In addition to the options listed above, in another embodiment, the parameter set PPS or SPS indicates whether the first AU in the bitstream, i.e., the CRA or IDR that constitutes the start of CLVS, should be output after decoding. Therefore, system integration is simpler because, for example, when parsing files in ISOBMFF format, only the parameter set needs to be adjusted, rather than changing relatively low-level syntax (e.g., PHs).
[0154] The example is shown below:
[0155]
[0156] A value of 1 for `sps_pic_in_cvss_au_no_output_flag` indicates that images referencing SPS in the CVSS AU should not be output. A value of 0 for `sps_pic_in_cvss_au_no_output_flag` indicates that images referencing SPS in the CVSS AU may or may not be output.
[0157] The requirement for bitstream consistency is that the value of sps_pic_in_cvss_au_no_output_flag should be the same for any SPS referenced by any output layer in OLS.
[0158] In 8.1.2
[0159] – The variable PictureOutputFlag for the current image is exported as follows:
[0160] – If sps_video_parameter_set_id is greater than 0 and the current layer is not an output layer (i.e., for any value of i in the range 0 to NumOutputLayersInOls[TargetOlsIdx]-1 (inclusive), nuh_layer_id is not equal to OutputLayerIdInOls[TargetOlsIdx][i]), or one of the following conditions is true, then PictureOutputFlag is set to 0:
[0161] – The current image is a RASL image, and the NoOutputBeforeRecoveryFlag of the associated IRAP image is equal to 1.
[0162] – The current image is either a GDR image with NoOutputBeforeRecoveryFlag equal to 1, or a recovered image of a GDR image with NoOutputBeforeRecoveryFlag equal to 1.
[0163] Otherwise, if the current AU is a CVSS AU and sps_pic_in_cvss_au_no_output_flag is equal to 1, then PictureOutputFlag is set to 0.
[0164] Otherwise, PictureOutputFlag is set to equal ph_pic_output_flag.
[0165] Note that in the implementation, the decoder can output images that do not belong to the output layer. For example, when there is only one output layer and the image of the output layer is unavailable in the AU, such as due to loss or layer downswitching, the decoder can set PictureOutputFlag to 1 for the image with the highest nuh_layer_id value and ph_pic_output_flag equal to 1 among all images available to the decoder in the AU, and set PictureOutputFlag to 0 for all other images available to the decoder in the AU.
[0166] In another embodiment, for example, the requirement can be defined as follows:
[0167] The requirement for bitstream consistency is that if the image belongs to an IRAP AU and the IRAP AU is directly before the DRAP AU, then the value of ph_pic_output_flag should be equal to 0.
[0168] The second aspect of the invention will now be described in detail below.
[0169] According to a second aspect of the invention, an apparatus is provided for receiving one or more input video data streams. Each of the one or more input video data streams encodes an input video. The apparatus 200 is configured to generate an output video data stream from the one or more input video data streams, the output video data streams encoding the output video, wherein the apparatus is configured to: generate the output video data stream such that the output video is an input video encoded within one of the one or more input video data streams, or such that the output video depends on the input video of at least one of the one or more input video data streams. Furthermore, the apparatus 200 is configured to determine the removal time of the current image from an access unit of an encoded image buffer among a plurality of images of the output video. The apparatus 200 is configured to: determine whether to use encoded image buffer delay offset information to determine the removal time of the current image from an access unit of the encoded image buffer.
[0170] According to an embodiment, the apparatus 200 may be configured, for example, to discard a group of one or more images of the input video of a first video data stream in one or more input video data streams to generate an output video data stream. The apparatus 200 may also be configured, for example, to determine, based on encoded image buffer delay offset information, the removal time of at least one image from the access unit of the encoded image buffer of the output video.
[0171] In an embodiment, the first video received by the device 200 may be, for example, a preprocessed video obtained from an original video, from which one or more images have been discarded to generate the preprocessed video. The device 200 may, for example, be configured to determine, based on encoded image buffer delay offset information, the removal time of at least one image from the access unit of the encoded image buffer in the output video.
[0172] According to an embodiment, the buffer delay offset information depends on the number of images that have been discarded in the input video.
[0173] In an embodiment, one or more input video data streams are two or more input video data streams. The apparatus 200 may, for example, be configured to concatenate the processed video with the input video of a second video data stream from the two or more input video data streams to obtain an output video, and the apparatus 200 may, for example, be configured to encode the output video into an output video data stream.
[0174] According to an embodiment, the device 200 may be configured, for example, to determine whether to use encoded image buffer delay offset information to determine the access unit removal time of the current image based on the position of the current image within the output video. Alternatively, the device 200 may be configured, for example, to determine whether to set the encoded image buffer delay offset value of the encoded image buffer delay offset information to 0 to determine the access unit removal time of the current image based on the position of the current image within the output video.
[0175] In an embodiment, the device 200 may be configured, for example, to determine whether to use encoded image buffer delay offset information to determine the access unit removal time of the current image based on the position of a previously non-discardable image preceding the current image in the output video.
[0176] According to an embodiment, the device 200 may be configured, for example, to determine whether to use encoded image buffer delay offset information to determine the access unit removal time of the current image based on whether a previous non-discardable image preceding the current image in the output video is, for example, the first image in a previous buffering cycle.
[0177] In an embodiment, the device 200 may be configured, for example, to determine whether to use encoded image buffer delay offset information to determine the access unit removal time of the current image, which is the first image of the input video of the second video data stream, based on a concatenation flag.
[0178] According to an embodiment, the device 200 may be configured, for example, to determine the access unit removal time of the current image based on the removal time of the previous image.
[0179] In an embodiment, the device 200 may be configured, for example, to determine the access unit removal time of the current image based on the initial encoded image buffer removal delay information.
[0180] According to an embodiment, the device 200 may be configured, for example, to update the initial encoded image buffer removal delay information according to a clock tick, so as to obtain temporary encoded image buffer removal delay information to determine the access unit removal time of the current image.
[0181] According to an embodiment, if the concatenation flag is set to a first value, the device 200 is configured to use encoded image buffer delay offset information to determine one or more removal times. If the concatenation flag is set to a second value different from the first value, the device 200 is configured not to use encoded image buffer delay offset information to determine one or more removal times.
[0182] In an embodiment, the device 200 may be configured, for example, to signal the video decoder 300 whether to use the encoded image buffer delay offset information to determine the time when the current image is removed from the access unit of the encoded image buffer.
[0183] According to an embodiment, the current image may be located, for example, at the splicing point of the output video, where two input videos have already been spliced together.
[0184] In addition, a video data stream is provided. This video data stream contains encoded video. The video data stream includes encoded image buffer delay offset information.
[0185] According to an embodiment, the video data stream may include, for example, a cascading flag.
[0186] In one embodiment, the video data stream may include, for example, initial encoded image buffer removal delay information.
[0187] According to an embodiment, for example, when it is known that some images (e.g., RASL images) have been discarded, if the concatenation flag is set to a first value (e.g., 0), the concatenation flag indicates that encoded image buffer delay offset information needs to be used to determine the removal time of one or more (images or access units). If the concatenation flag is set to a second value different from the first value (e.g., 1), the concatenation flag indicates that the indicated offset is not used to determine the removal time of one or more (images or access units), for example, regardless of offset signaling and, for example, regardless of whether the RASL images have been discarded. If the images have not been discarded, then, for example, the offset is not used.
[0188] In addition, a video encoder 100 is provided. The video encoder 100 is configured to encode video into a video data stream. The video encoder 100 is configured to generate a video data stream such that the video data stream includes encoded image buffer delay offset information.
[0189] According to an embodiment, the video encoder 100 may be configured, for example, to generate a video data stream such that the video data stream may include, for example, a cascading flag.
[0190] In an embodiment, the video encoder 100 may be configured, for example, to generate a video data stream such that the video data stream may include, for example, encoded image buffer delay offset information.
[0191] In addition, a video decoder 300 is provided for receiving a video data stream in which video is stored. The video decoder 300 is configured to decode video from the video data stream. Furthermore, the video decoder 300 is configured to decode the video based on the removal time of the current image from the access unit of the encoded image buffer among a plurality of images of the video. The video decoder 300 is configured to decode the video based on an indication indicating whether encoded image buffer delay offset information is used to determine the removal time of the current image from the access unit of the encoded image buffer.
[0192] According to an embodiment, the time at which at least one of the multiple images in the video is removed from the access unit of the encoded image buffer depends on the encoded image buffer delay offset information.
[0193] In this embodiment, the video decoder 300 is configured to decode the video based on the following: determining whether to use encoded image buffer delay offset information to determine the access unit removal time of the current image based on the position of the current image within the video.
[0194] According to an embodiment, the video decoder 300 may be configured, for example, to decode the video based on whether the encoded image buffer delay offset value of the encoded image buffer delay offset information can be set to 0, for example.
[0195] In an embodiment, the video decoder 300 may be configured, for example, to determine whether to use encoded image buffer delay offset information to determine the access unit removal time of the current image based on the position of a previous non-discardable image in the video preceding the current image.
[0196] According to an embodiment, the video decoder 300 may be configured, for example, to determine whether to use encoded image buffer delay offset information to determine the access unit removal time of the current image based on whether a previous non-discardable image preceding the current image in the video is, for example, the first image in a previous buffering cycle.
[0197] In an embodiment, the video decoder 300 may be configured, for example, to determine whether to use encoded image buffer delay offset information to determine the access unit removal time of the current image, which is the first image of the input video of the second video data stream, based on a concatenation flag.
[0198] According to an embodiment, the video decoder 300 may be configured, for example, to determine the access unit removal time of the current image based on the removal time of the previous image.
[0199] In an embodiment, the video decoder 300 may be configured, for example, to determine the access unit removal time of the current image based on the initial encoded image buffer removal delay information.
[0200] According to an embodiment, the video decoder 300 may be configured, for example, to update the initial encoded image buffer removal delay information according to a clock tick, in order to obtain temporary encoded image buffer removal delay information to determine the access unit removal time of the current image.
[0201] According to an embodiment, if the concatenation flag is set to a first value, the video decoder 300 is configured to use encoded image buffer delay offset information to determine one or more removal times. If the concatenation flag is set to a second value different from the first value, the video decoder 300 is configured not to use encoded image buffer delay offset information to determine one or more removal times.
[0202] In addition, a system is provided. This system includes the device 200 as described above and the video decoder 300 as described above. The video decoder 300 is configured to receive the output video data stream of the device 200. Furthermore, the video decoder 300 is configured to decode video from the output video data stream of the device 200.
[0203] According to an embodiment, the system may also include, for example, a video encoder 100. The device 200 may be configured, for example, to receive a video data stream from the video encoder 100 as an input video data stream.
[0204] Specifically, the second aspect of the invention relates to the possibility that prevNonDiscardable may already include the alternative offset (CpbDelayOffset) when the alternative is not at the start of BP, so for AU with concatenation_flag==1, CpbDelayOffset should be temporarily set to zero.
[0205] When two bitstreams are concatenated, the removal time of the derived AU from the CPB is different from that for non-concatenated bitstreams. At the concatenation point, the buffer periodic SEI message (BP SEI message; SEI = Supplementary Enhancement Information) includes a concatenationFlag equal to 1. The decoder then needs to check two values and take the larger one:
[0206] ● The removal time of previously non-discardable pictures (prevNonDiscardablePic) plus the increment of signaling notification in the BP SEI message (auCpbRemovalDelayDeltaMinus1+1), or
[0207] ●Add InitCpbRemovalDelay to the previous image removal time.
[0208] However, when the previous image with the BP SEI message is an AU for which the removal time has been derived using an alternative timing information (i.e., the second timing information used when the RASL image or the image up to DRAP has been discarded), the offset (CpbDelayOffset) is used to calculate each removal time, which is calculated as the increment of the previous image with the buffer period, i.e., AuNominalRemovalTime[firstPicInPrevBuffPeriod] plus AuCpbRemovalDelayVal – CpbDelayOffset, as follows. Figure 4 As shown.
[0209] Figure 4 The original bitstream is shown. Figure 4 (top of the image) and the bitstream after discarding the image ( Figure 4 (Bottom of the table): After discarding AUs (lines 1, 2, and 3 in the original bitstream), the offset is incorporated into the calculation of the removal delay.
[0210] The offset is added because the removal time is calculated using the increment of the removal time of the image (called firstPicInPrevBuffPeriod), after which some AUs have been discarded, so CpbDelayOffset is needed to account for (compensate) the discarded AUs.
[0211] Figure 5 This shows the image from the original first bitstream (in Figure 5 In the middle, located to the left of the middle) after discarding (at different positions) two bit streams (the first bit stream (in the middle left) Figure 5 In the middle, located in the middle left) and the second bitstream (in Figure 5 (The splicing of the middle right)
[0212] The example of using the previous image removal time as the anchor point instead of the previous non-discardable image is similar and does not require the "-3" correction factor (CpbDelayOffset).
[0213] However, in such Figure 5In the stitching scenario shown, note that this is not necessarily the case where two AUs are derived using the removal time (firstPicInPrevBuffPeriod) associated with the BP SEI message. As discussed, for the stitching scenario, the increment is added to prevNonDiscardablePic or simply the previous picture. This means that when prevNonDiscardablePic is not firstPicInPrevBuffPeriod, CpbDelayOffset cannot be used to derive the removal time of the current AU from the CPB, because the removal time of prevNonDiscardablePic already accounts for AU dropping, and no AU is dropped between prevNonDiscardablePic and the AU for which the removal time is calculated. Now, assuming the removal time of the previous picture is used instead, this will achieve equidistant removal times (when using prevNonDiscardablePic instead) for the current AU (i.e., the stitching point with the new BP SEI message) having InitialCpbRemovalDelay (which forces the removal time of the current AU to be after its expected removal time). In this case, the removal time of the current AU cannot be less than the time calculated by adding InitCpbRemovalDelay to the removal time of the previous image, because this could lead to buffer underloading (the AU is not in the buffer before it needs to be removed). Therefore, as part of this invention, CpbDelayOffset is not used for calculation or is considered equal to 0 in this case.
[0214] The embodiments described in this article rely on checking the use of CpbDelayOffset to calculate the AU removal time when RASL AUs or AUs between IRAP and DRAP AUs are dropped from the bitstream. One of the checks used to determine whether CpbDelayOffset is not used or is considered equal to 0 is:
[0215] ●prevNonDiscardablePic is not firstPicInPrevBuffPeriod
[0216] ●The removal time of the previous image plus InitCpbRemovalDelay is used to calculate the removal time of the current AU.
[0217] The implementation in the specification can be as follows:
[0218] – When AU n is the first AU of the BP that is not initialized with HRD, the following applies:
[0219] The nominal removal time of AU n from CPB is specified as follows:
[0220]
[0221] alternative sites, in Figure 6 In another embodiment shown, the CpbDelayOffset used to calculate the AU removal time when the RASL AU is dropped from the bitstream or when an AU between IRAP and DRAP AU is dropped depends on different checks, including checking the concatenationFlag.
[0222] In this scenario, when concatenationFlag is set to 1, the increment in the bitstream needs to match the correct value, as if CpbDelayOffset were taken into account (when comparing). Figure 5 and Figure 6 (This is quite obvious, because for this graph, CpbDelayOffset is not applied or is treated as 0.)
[0223] The implementation in the specification can be as follows:
[0224] – When AU n is the first AU of the BP that is not initialized with HRD, the following applies:
[0225] The nominal removal time of AU n from CPB is specified as follows:
[0226]
[0227] The third aspect of the invention will now be described in detail below.
[0228] According to a third aspect of the invention, a video data stream is provided. Video is encoded in the video data stream. Furthermore, the video data stream includes an initial encoded image buffer removal delay. Furthermore, the video data stream includes an initial encoded image buffer removal offset. Furthermore, the video data stream includes information indicating whether the sum of the initial encoded image buffer removal delay and the initial encoded image buffer removal offset is defined as being constant across two or more buffering cycles.
[0229] According to an embodiment, the initial encoded image buffer removal delay may, for example, indicate the time required for the first access unit to elapse before sending the first access unit of the image that initializes the video data stream of the video decoder 300 to the video decoder 300.
[0230] In an embodiment, the video data stream may include, for example, a single indication that indicates whether the sum of the initial encoded image buffer removal delay and the initial encoded image buffer removal offset can be defined as constant across two or more buffer cycles.
[0231] According to an embodiment, the video data stream may, for example, include a concatenation flag as a single indicator, which may indicate, for example, whether the sum of the initial encoded image buffer removal delay and the initial encoded image buffer removal offset can be defined as constant across two or more buffer cycles. If the concatenation flag is equal to a first value, then the sum of the initial encoded image buffer removal delay and the initial encoded image buffer removal offset is constant across two or more buffer cycles. If the concatenation flag is different from the first value, then the concatenation flag does not define whether the sum of the initial encoded image buffer removal delay and the initial encoded image buffer removal offset is constant across two or more buffer cycles.
[0232] In an embodiment, if the single indication does not indicate that the sum of the initial encoded image buffer removal delay and the initial encoded image buffer removal offset is defined as constant across two or more buffer cycles, then the video data stream may, for example, include continuously updated information about the initial encoded image buffer removal delay and continuously updated information about the initial encoded image buffer removal offset.
[0233] According to an embodiment, if the video data stream includes information indicating that the sum of the initial encoded image buffer removal delay and the initial encoded image buffer removal offset is defined as constant across two or more buffer cycles, then the sum of the initial encoded image buffer removal delay and the initial encoded image buffer removal offset can, for example, be defined as constant starting from the current position within the video data stream.
[0234] In addition, a video encoder 100 is provided. The video encoder 100 is configured to encode video into a video data stream. Furthermore, the video encoder 100 is configured to generate a video data stream such that the video data stream includes an initial encoded image buffer removal delay. Furthermore, the video encoder 100 is configured to generate a video data stream such that the video data stream includes an initial encoded image buffer removal offset. Furthermore, the video encoder 100 is configured to generate a video data stream such that the video data stream includes information indicating whether the sum of the initial encoded image buffer removal delay and the initial encoded image buffer removal offset is defined as constant across two or more buffer cycles.
[0235] According to an embodiment, the initial encoded image buffer removal delay may, for example, indicate the time required for the first access unit to elapse before sending the first access unit of the image that initializes the video data stream of the video decoder 300 to the video decoder 300.
[0236] In an embodiment, the video encoder 100 may be configured, for example, to generate a video data stream such that the video data stream may include, for example, a single indication that the sum of the initial encoded picture buffer removal delay and the initial encoded picture buffer removal offset may be defined, for example, as constant across two or more buffer cycles.
[0237] According to an embodiment, the video encoder 100 may, for example, be configured to generate a video data stream such that the video data stream may include, for example, a concatenation flag as a single indicator, which may indicate, for example, whether the sum of the initial encoded image buffer removal delay and the initial encoded image buffer removal offset can be defined as constant across two or more buffer cycles. If the concatenation flag is equal to a first value, then the sum of the initial encoded image buffer removal delay and the initial encoded image buffer removal offset is constant across two or more buffer cycles. If the concatenation flag is different from the first value, then the concatenation flag does not define whether the sum of the initial encoded image buffer removal delay and the initial encoded image buffer removal offset is constant across two or more buffer cycles.
[0238] In an embodiment, if the single indication does not indicate that the sum of the initial encoded image buffer removal delay and the initial encoded image buffer removal offset is defined as constant across two or more buffer cycles, then the video encoder 100 may, for example, be configured to generate a video data stream such that the video data stream may, for example, include continuously updated information about the initial encoded image buffer removal delay and continuously updated information about the initial encoded image buffer removal offset.
[0239] According to an embodiment, if the video data stream includes, for example, information that can indicate that the sum of the initial encoded image buffer removal delay and the initial encoded image buffer removal offset is defined as constant across two or more buffer cycles, then the sum of the initial encoded image buffer removal delay and the initial encoded image buffer removal offset is defined as constant starting from the current position within the video data stream.
[0240] In addition, an apparatus 200 is provided for receiving two input video data streams (as a first input video data stream and a second input video data stream). Each of the two input video data streams contains an input video. The apparatus 200 is configured to generate an output video data stream from the two input video data streams, the output video data stream encoding the output video, wherein the apparatus is configured to generate the output video data stream by concatenating the first input video data stream and the second input video data stream. Furthermore, the apparatus 200 is configured to generate the output video data stream such that the output video data stream includes an initial encoded image buffer removal delay. Furthermore, the apparatus 200 is configured to generate the output video data stream such that the output video data stream includes an initial encoded image buffer removal offset. Furthermore, the apparatus 200 is configured to generate the output video data stream such that the output video data stream includes information indicating whether the sum of the initial encoded image buffer removal delay and the initial encoded image buffer removal offset is defined as constant across two or more buffer cycles.
[0241] According to an embodiment, the initial encoded image buffer removal delay may, for example, indicate the time required for the first access unit to elapse before sending the first access unit of the image that initializes the output video data stream of the video decoder 300 to the video decoder 300.
[0242] In an embodiment, the apparatus 200 may be configured, for example, to generate an output video data stream such that the output video data stream may include, for example, a single indication that the sum of the initial encoded image buffer removal delay and the initial encoded image buffer removal offset may be defined, for example, as constant across two or more buffer cycles.
[0243] According to an embodiment, apparatus 200 may, for example, be configured to generate an output video data stream such that the output video data stream may include, for example, a concatenation flag as a single indication, which may indicate, for example, whether the sum of the initial encoded image buffer removal delay and the initial encoded image buffer removal offset can be defined as constant across two or more buffer cycles. If the concatenation flag is equal to a first value, then the sum of the initial encoded image buffer removal delay and the initial encoded image buffer removal offset is constant across two or more buffer cycles. If the concatenation flag is different from the first value, then the concatenation flag does not define whether the sum of the initial encoded image buffer removal delay and the initial encoded image buffer removal offset is constant across two or more buffer cycles.
[0244] In an embodiment, if the single indication does not indicate that the sum of the initial encoded image buffer removal delay and the initial encoded image buffer removal offset is defined as constant across two or more buffer cycles, then the apparatus 200 is configured to generate an output video data stream such that the output video data stream includes continuously updated information about the initial encoded image buffer removal delay and continuously updated information about the initial encoded image buffer removal offset.
[0245] According to an embodiment, if the video data stream includes information indicating that the sum of the initial encoded image buffer removal delay and the initial encoded image buffer removal offset is defined as constant across two or more buffer cycles, then the sum of the initial encoded image buffer removal delay and the initial encoded image buffer removal offset is defined as constant starting from the current position within the video data stream.
[0246] Furthermore, a video decoder 300 is provided for receiving a video data stream in which video is stored. The video decoder 300 is configured to decode video from the video data stream. The video data stream includes an initial encoded image buffer removal delay. The video data stream also includes an initial encoded image buffer removal offset. Furthermore, the video data stream includes information indicating whether the sum of the initial encoded image buffer removal delay and the initial encoded image buffer removal offset is defined as constant across two or more buffering cycles. Furthermore, the video decoder 300 is configured to decode the video based on the information indicating whether the sum of the initial encoded image buffer removal delay and the initial encoded image buffer removal offset is defined as constant across two or more buffering cycles.
[0247] According to an embodiment, the initial encoded image buffer removal delay may, for example, indicate the time required for the first access unit to elapse before sending the first access unit of the image that initializes the output video data stream of the video decoder 300 to the video decoder 300.
[0248] In an embodiment, the video data stream may, for example, include a single indication, which may indicate whether the sum of the initial encoded image buffer removal delay and the initial encoded image buffer removal offset can, for example, be defined as constant across two or more buffer cycles. The video decoder 300 may, for example, be configured to decode the video according to this single indication.
[0249] According to an embodiment, the video data stream may, for example, include a concatenation flag as a single indicator, which may indicate, for example, whether the sum of the initial encoded image buffer removal delay and the initial encoded image buffer removal offset can be defined as constant across two or more buffering cycles. If the concatenation flag equals a first value, then the sum of the initial encoded image buffer removal delay and the initial encoded image buffer removal offset is constant across two or more buffering cycles. If the concatenation flag differs from the first value, then the concatenation flag does not define whether the sum of the initial encoded image buffer removal delay and the initial encoded image buffer removal offset is constant across two or more buffering cycles. The video decoder 300 is configured to decode the video according to the concatenation flag.
[0250] In an embodiment, if the single indication does not indicate that the sum of the initial encoded image buffer removal delay and the initial encoded image buffer removal offset is defined as constant across two or more buffer cycles, then the video data stream includes continuously updated information regarding the initial encoded image buffer removal delay and continuously updated information regarding the initial encoded image buffer removal offset. The video decoder 300 is configured to decode the video based on the continuously updated information regarding the initial encoded image buffer removal delay and the continuously updated information regarding the initial encoded image buffer removal offset.
[0251] According to an embodiment, if the video data stream includes information indicating that the sum of the initial encoded image buffer removal delay and the initial encoded image buffer removal offset is defined as constant across two or more buffer cycles, then the sum of the initial encoded image buffer removal delay and the initial encoded image buffer removal offset is defined as constant starting from the current position within the video data stream.
[0252] In addition, a system is provided. This system includes the device 200 as described above and the video decoder 300 as described above. The video decoder 300 is configured to receive the output video data stream of the device 200. Furthermore, the video decoder 300 is configured to decode video from the output video data stream of the device 200.
[0253] According to an embodiment, the system may also include, for example, a video encoder 100. The device 200 may be configured, for example, to receive a video data stream from the video encoder 100 as an input video data stream.
[0254] Specifically, the third aspect of the invention relates to splicing, initial Cpb removal delay, and initial Cpb removal offset.
[0255] Currently, the specification indicates that the sum of the initial Cpb removal delay and the initial Cpb removal offset is constant within the CVS. The same constraint is expressed for alternative scenarios. The initial Cpb removal delay indicates the time required for the first AU in the bitstream used to initialize the decoder before it can be used for decoding. The initial Cpb removal offset is a property of the bitstream, meaning that the earliest arrival time of an AU in the decoder is not necessarily equidistant from the time 0 when the first AU arrives at the decoder. It helps determine when the first bit of an AU can arrive at the decoder earliest.
[0256] The current constraint in the VVC draft specification indicates that the sum of these two values must be constant within CVS:
[0257] Throughout the entire CVS, for every pair of values of i and j, the sum of nal_initial_cpb_removal_delay[i][j] and nal_initial_cpb_removal_offset[i][j] should be constant, and the sum of nal_initial_alt_cpb_removal_delay[i][j] and nal_initial_alt_cpb_removal_offset[i][j] should also be constant.
[0258] Problems arise when editing or splicing bitstreams to form new combined bitstreams. It is also desirable to be able to indicate whether the CVS boundary across the bitstream is satisfied, as different values for this sum can lead to buffer underloading or overflow.
[0259] Therefore, in this embodiment, an indication is carried in the bitstream, and starting from a point in the bitstream (e.g., a splice point), the constraint on the constant sum of InitCpbRemovalDelay and InitCpbRemovalDelayOffset (and alternative counterparts) is reset, and the sums in the bitstream before and after that point may differ. The sum remains constant unless the indication is present in the bitstream.
[0260] For example:
[0261] When concatenationFlag equals 0, the constraint on bitstream consistency is that the sum of InitCpbRemovalDelay and InitCpbRemovalDelayOffset is constant across buffer cycles.
[0262] Otherwise, the sum of InitCpbRemovalDelay and InitCpbRemovalDelayOffset does not necessarily need to be constant across buffer cycles. Update the values of InitCpbRemovalDelay and InitCpbRemovalDelayOffset to account for arrival time.
[0263] In an embodiment, if several bit streams are concatenated, at each concatenation point, a concatenation flag can define, for example, whether the sum remains constant.
[0264] The fourth aspect of the invention is described in detail below.
[0265] According to a fourth aspect of the invention, a video data stream is provided. Video is encoded in the video data stream. Furthermore, the video data stream includes an indication (e.g., general_same_pic_timing_in_all_ols_flag) indicating whether a non-scalable nested picture timing supplementation enhancement message for a network abstraction layer unit of a plurality of access units in one or more encoded video sequences of the video data stream is defined as applicable to all output layer sets in a plurality of output layer sets of the access unit. If the indication (e.g., general_same_pic_timing_in_all_ols_flag) has a first value, then the non-scalable nested picture timing supplementation enhancement message for the network abstraction layer unit of the access unit is defined as applicable to all output layer sets in a plurality of output layer sets of the access unit. If the indication (e.g., general_same_pic_timing_in_all_ols_flag) has a value different from the first value, then the indication does not define whether the non-scalable nested picture timing supplementation enhancement message for the network abstraction layer unit of the access unit is applicable to all output layer sets in a plurality of output layer sets of the access unit.
[0266] According to an embodiment, for example, if the indication (e.g., general_same_pic_timing_in_all_ols_flag) has a first value, the network abstraction layer unit does not include any other supplementary enhancement information messages that are different from the picture timing supplementary enhancement information message.
[0267] In an embodiment, for example, if the indication (e.g., general_same_pic_timing_in_all_ols_flag) has a first value, the network abstraction layer unit does not include any other supplementary enhancement information messages.
[0268] According to an embodiment, for example, if the indication (e.g., general_same_pic_timing_in_all_ols_flag) has a first value, then for each of the multiple access units of the encoded video sequence in one or more encoded video sequences, each network abstraction layer unit includes a non-scalable nested picture timing supplemental enhancement information message, the network abstraction layer unit does not include any other supplemental enhancement information message different from the picture timing supplemental enhancement information message, or does not include any other supplemental enhancement information message.
[0269] In an embodiment, if the indication (e.g., general_same_pic_timing_in_all_ols_flag) has a first value, then for each of the plurality of access units in one or more coded video sequences of a video data stream, each network abstraction layer unit includes a non-scalable nested picture timing supplemental enhancement message, the network abstraction layer unit does not include any other supplemental enhancement message different from the picture timing supplemental enhancement message, or does not include any other supplemental enhancement message.
[0270] Additionally, a video encoder 100 may be provided, for example. The video encoder 100 is configured to encode video into a video data stream. Furthermore, the video encoder 100 is configured to generate a video data stream such that the video data stream includes an indication (e.g., `general_same_pic_timing_in_all_ols_flag`) indicating whether a non-scalable nested picture timing supplementation enhancement message of a network abstraction layer unit of a plurality of access units in one or more encoded video sequences of the video data stream is defined as applicable to all sets of output layers in a plurality of output layer sets of said access unit. If the indication (e.g., `general_same_pic_timing_in_all_ols_flag`) has a first value, then the non-scalable nested picture timing supplementation enhancement message of the network abstraction layer unit of said access unit is defined as applicable to all sets of output layers in a plurality of output layer sets of said access unit. If the indicator (e.g., general_same_pic_timing_in_all_ols_flag) has a value different from the first value, then the indicator does not define whether the non-scalable nested picture timing supplementation enhancement information message of the network abstraction layer unit of the access unit is applicable to all output layer sets in the multiple output layer sets of the access unit.
[0271] According to an embodiment, for example, if the indication (e.g., general_same_pic_timing_in_all_ols_flag) has a first value, the video encoder 100 is configured to generate a video data stream such that the network abstraction layer unit does not include any other supplementary enhancement information messages different from the picture timing supplementary enhancement information message.
[0272] In an embodiment, for example, if the indication (e.g., general_same_pic_timing_in_all_ols_flag) has a first value, the video encoder 100 is configured to generate a video data stream such that the network abstraction layer unit does not include any other supplementary enhancement information messages.
[0273] According to an embodiment, for example, if the indication (e.g., general_same_pic_timing_in_all_ols_flag) has a first value, the video encoder 100 may be configured, for example, to generate a video data stream such that each network abstraction layer unit, for each of a plurality of access units of a encoded video sequence in one or more encoded video sequences, includes a non-scalable nested picture timing supplemental enhancement information message, the network abstraction layer unit not including any other supplemental enhancement information message different from the picture timing supplemental enhancement information message, or not including any other supplemental enhancement information message.
[0274] In an embodiment, for example, if the indication (e.g., general_same_pic_timing_in_all_ols_flag) has a first value, the video encoder 100 may be configured, for example, to generate a video data stream such that each of the plurality of access units of each of one or more encoded video sequences in the video data stream includes a non-scalable nested picture timing supplemental enhancement information message in each network abstraction layer unit, the network abstraction layer unit not including any other supplemental enhancement information message different from the picture timing supplemental enhancement information message, or not including any other supplemental enhancement information message.
[0275] In addition, an apparatus 200 is provided for receiving an input video data stream. The input video data stream contains encoded video. The apparatus 200 is configured to generate a processed video data stream from the input video data stream. Furthermore, the apparatus 200 is configured to generate the processed video data stream such that the processed video data stream includes an indication (e.g., general_same_pic_timing_in_all_ols_flag) indicating whether a non-scalable nested picture timing supplementation enhancement message of a network abstraction layer unit of a plurality of access units of an access unit in one or more encoded video sequences of the processed video data stream is defined as applicable to all output layer sets in a plurality of output layer sets of the access unit. If the indication (e.g., general_same_pic_timing_in_all_ols_flag) has a first value, then the non-scalable nested picture timing supplementation enhancement message of the network abstraction layer unit of the access unit is defined as applicable to all output layer sets in a plurality of output layer sets of the access unit. If the indicator (e.g., general_same_pic_timing_in_all_ols_flag) has a value different from the first value, then the indicator does not define whether the non-scalable nested picture timing supplementation enhancement information message of the network abstraction layer unit of the access unit is applicable to all output layer sets in the multiple output layer sets of the access unit.
[0276] According to an embodiment, for example, if the indication (e.g., general_same_pic_timing_in_all_ols_flag) has a first value, the apparatus 200 is configured to generate a processed video data stream such that the network abstraction layer unit does not include any other supplementary enhancement information messages different from the picture timing supplementary enhancement information message.
[0277] In an embodiment, for example, if the indication (e.g., general_same_pic_timing_in_all_ols_flag) has a first value, the apparatus 200 is configured to generate a processed video data stream such that the network abstraction layer unit does not include any other supplementary enhancement information messages.
[0278] According to an embodiment, for example, if the indication (e.g., general_same_pic_timing_in_all_ols_flag) has a first value, the apparatus 200 may be configured, for example, to generate a processed video data stream such that each network abstraction layer unit, for each of a plurality of access units of a encoded video sequence in one or more encoded video sequences, includes a non-scalable nested picture timing supplemental enhancement information message, the network abstraction layer unit not including any other supplemental enhancement information message different from the picture timing supplemental enhancement information message, or not including any other supplemental enhancement information message.
[0279] In an embodiment, for example, if the indication (e.g., general_same_pic_timing_in_all_ols_flag) has a first value, the apparatus 200 may be configured, for example, to generate a processed video data stream such that each of the plurality of access units of each of the one or more coded video sequences processing the video data stream includes a non-scalable nested picture timing supplemental enhancement information message for each access unit, the network abstraction layer unit not including any other supplemental enhancement information message different from the picture timing supplemental enhancement information message, or not including any other supplemental enhancement information message.
[0280] Furthermore, a video decoder 300 is provided for receiving a video data stream in which video is stored. The video decoder 300 is configured to decode video from the video data stream. The video data stream includes an indication (e.g., `general_same_pic_timing_in_all_ols_flag`) indicating whether a non-scalable nested picture timing supplementation enhancement message for a network abstraction layer unit of a network abstraction layer unit of a plurality of access units in one or more encoded video sequences of the video data stream is defined as applicable to all output layer sets in a plurality of output layer sets of the access unit. If the indication (e.g., `general_same_pic_timing_in_all_ols_flag`) has a first value, then the non-scalable nested picture timing supplementation enhancement message for the network abstraction layer unit of the access unit is defined as applicable to all output layer sets in a plurality of output layer sets of the access unit. If the indication (e.g., general_same_pic_timing_in_all_ols_flag) has a value different from the first value, then the indication does not define whether the non-scalable nested picture timing supplementation enhancement information message of the network abstraction layer unit of the access unit applies to all output layer sets in the multiple output layer sets of the access unit. The video decoder 300 is configured to decode the video according to the indication.
[0281] According to an embodiment, for example, if the indication (e.g., general_same_pic_timing_in_all_ols_flag) has a first value, the network abstraction layer unit does not include any other supplementary enhancement information messages that are different from the picture timing supplementary enhancement information message.
[0282] In an embodiment, for example, if the indication (e.g., general_same_pic_timing_in_all_ols_flag) has a first value, the network abstraction layer unit does not include any other supplemental enhancement information messages. The video decoder 300 is configured to decode the video according to the indication.
[0283] According to an embodiment, for example, if the indication (e.g., general_same_pic_timing_in_all_ols_flag) has a first value, then for each of the multiple access units of the encoded video sequence in one or more encoded video sequences, each network abstraction layer unit includes a non-scalable nested picture timing supplemental enhancement information message, the network abstraction layer unit does not include any other supplemental enhancement information message different from the picture timing supplemental enhancement information message, or does not include any other supplemental enhancement information message.
[0284] In an embodiment, for example, if the indication (e.g., general_same_pic_timing_in_all_ols_flag) has a first value, then for each of the plurality of access units in one or more coded video sequences of a video data stream, each network abstraction layer unit includes a non-scalable nested picture timing supplemental enhancement information message, the network abstraction layer unit does not include any other supplemental enhancement information message different from the picture timing supplemental enhancement information message, or does not include any other supplemental enhancement information message.
[0285] In addition, a system is provided. This system includes the apparatus 200 as described above and the video decoder 300 as described above. The video decoder 300 is configured to receive the processed video data stream from the apparatus 200. Furthermore, the video decoder 300 is configured to decode video from the output video data stream of the apparatus 200.
[0286] According to an embodiment, the system may also include, for example, a video encoder 100. The device 200 may be configured, for example, to receive a video data stream from the video encoder 100 as an input video data stream.
[0287] Specifically, the fourth aspect of the invention relates to constraining the PT SEI from being paired with other HRD SEIs when general_same_pic_timing_in_all_ols_flag equals 1.
[0288] The VVC draft specification includes a flag called general_same_pic_timing_in_all_ols_flag in the general HRD parameter structure, with the following semantics:
[0289] `general_same_pic_timing_in_all_ols_flag` equal to 1 indicates that non-scalable nested PT SEI messages in each AU apply to AUs of any OLS in the bitstream and that scalable nested PT SEI messages do not exist. `general_same_pic_timing_in_all_ols_flag` equal to 0 indicates that non-scalable nested PT SEI messages in each AU may or may not apply to AUs of any OLS in the bitstream and that scalable nested PT SEI messages may exist.
[0290] Typically, when extracting an OLS sub-bitstream from the original bitstream (including OLS data plus non-OLS data), the corresponding HRD-related timing / buffer information of the target OLS, encapsulated in the form of buffer period, picture timing, and decoding unit information SEI messages (so-called scalable nested SEI messages), is decapsulated. This decapsulated SEI message is then used to replace the non-scalable nested HRD SEI information in the original bitstream. However, in many scenarios, the content of some messages (e.g., picture timing SEI messages) can remain unchanged when the layer is discarded, i.e., from one OLS to its subset. Therefore, `general_same_pic_timing_in_all_ols_flag` provides a shortcut that allows only the BP and DUI SEI messages to be replaced, but the PTSEI in the original bitstream can remain valid; that is, when `general_same_pic_timing_in_all_ols_flag` equals 1, the PT SEI is not removed during extraction. Therefore, it is not necessary to encapsulate the replacement PT SEI message in the scalable nested SEI message carrying the replacement BP and DUI SEI messages, and no bitrate overhead is introduced for this information.
[0291] However, in the prior art, PT SEI messages are allowed to be carried together with other HRD SEI messages within a single SEINAL unit (NAL unit = Network Abstraction Layer Unit), meaning that BP, PT, and SEI messages can all be encapsulated within the same prefixed SEI NAL unit. Therefore, the extractor must perform a more in-depth examination of this SEI NAL unit to understand the included messages, and when only one of the included messages (PT) is retained during the extraction process, the explicit SEINAL unit actually needs to be rewritten (i.e., non-PT SEI messages are removed). To avoid this cumbersome low-level processing and to allow the extractor to operate entirely on the non-parametric set portion of the bitstream at the NAL unit level, a bitstream constraint prohibits this bitstream construction from being part of this invention. In one embodiment, this constraint is stated as follows:
[0292] `general_same_pic_timing_in_all_ols_flag` equal to 1 indicates that non-scalable nested PT SEI messages in each AU apply to AUs of any OLS in the bitstream and that no scalable nested PT SEI messages exist. `general_same_pic_timing_in_all_ols_flag` equal to 0 indicates that non-scalable nested PT SEI messages in each AU may or may not apply to AUs of any OLS in the bitstream and that scalable nested PT SEI messages may exist. When `general_same_pic_timing_in_all_ols_flag` equals 1, the bitstream consistency constraint is: all general SEI messages in the bitstream that contain SEI messages with `payload_type` equal to 1 (picture timing) should not contain SEI messages with `payload_type` not equal to 1.
[0293] The fifth aspect of the invention is described in detail below.
[0294] According to a fifth aspect of the invention, a video data stream is provided. Video is encoded in the video data stream. Furthermore, the video data stream includes one or more scalable nested supplemental enhancement information messages. The one or more scalable nested supplemental enhancement information messages include a plurality of syntax elements. Each of the one or more syntax elements is defined to have the same size in each scalable nested supplemental enhancement information message in or within the video data stream.
[0295] According to an embodiment, the video data stream may, for example, include one or more non-scalable nested supplemental enhancement messages. The one or more scalable nested supplemental enhancement messages and the one or more non-scalable nested supplemental enhancement messages include a plurality of syntax elements. Each of the one or more syntax elements is defined to have the same size in each scalable nested supplemental enhancement message in the video data stream or a portion thereof, and in each non-scalable nested supplemental enhancement message in the video data stream or a portion thereof.
[0296] In an embodiment, the video data stream may, for example, include a plurality of access units, wherein each of the plurality of access units may, for example, be assigned to one of a plurality of images in the video. This portion of the video data stream may, for example, be an access unit among the plurality of access units in the video data stream. Each of one or more syntax elements among the plurality of syntax elements may, for example, be defined as having the same size in each of the scalable nested supplementary enhancement information messages in the scalable nested supplementary enhancement information messages of the access unit.
[0297] According to an embodiment, the video data stream may, for example, include one or more non-scalable nested supplemental enhancement information messages. The one or more scalable nested supplemental enhancement information messages and the one or more non-scalable nested supplemental enhancement information messages include multiple syntax elements. Each syntax element of the multiple syntax elements may, for example, be defined as having the same size in each scalable nested supplemental enhancement information message of the access unit and in each non-scalable nested supplemental enhancement information message of the access unit.
[0298] In an embodiment, this portion of the video data stream may, for example, be an encoded video sequence of the video data stream. Each of one or more of the plurality of syntax elements may, for example, be defined as having the same size in each scalable nested supplementary enhancement information message in the scalable nested supplementary enhancement information message of the encoded video sequence.
[0299] According to an embodiment, the video data stream may, for example, include one or more non-scalable nested supplemental enhancement messages. The one or more scalable nested supplemental enhancement messages and the one or more non-scalable nested supplemental enhancement messages include a plurality of syntax elements. Each of the one or more syntax elements may, for example, be defined as having the same size in each scalable nested supplemental enhancement message of the encoded video sequence and in each non-scalable nested supplemental enhancement message of the encoded video sequence.
[0300] In an embodiment, each of one or more of the plurality of syntax elements may, for example, be defined as having the same size in each of the scalable nested supplemental enhancement messages in the video data stream.
[0301] According to an embodiment, each of the plurality of syntax elements may be defined, for example, as having the same size in each scalable nested supplementary enhancement message in the scalable nested supplementary enhancement message of the video data stream and in each non-scalable nested supplementary enhancement message of the non-scalable nested supplementary enhancement message of the video data stream.
[0302] In an embodiment, the video data stream or a portion thereof may include, for example, at least one buffered periodic supplemental enhancement message, wherein the buffered periodic supplemental enhancement message defines the size of each of one or more syntax elements among a plurality of syntax elements.
[0303] According to an embodiment, the buffer cycle supplementation enhancement information message includes at least one of the following for defining the size of each of one or more syntax elements among a plurality of syntax elements:
[0304] bp_cpb_initial_removal_delay_length_minus1 element,
[0305] The element bp_cpb_removal_delay_length_minus1
[0306] The element bp_dpb_output_delay_length_minus1
[0307] bp_du_cpb_removal_delay_increment_length_minus1 element,
[0308] The element is bp_dpb_output_delay_du_length_minus1.
[0309] In an embodiment, for each of the multiple access units of a video data stream that includes a scalable nested buffer periodic supplementation enhancement information message, the access unit may also include, for example, a non-scalable nested buffer periodic supplementation enhancement information message.
[0310] According to an embodiment, for each of the multiple single-level access units of a video data stream that includes a scalable nested buffer periodic supplementation enhancement information message, the single-level access unit may also include, for example, a non-scalable nested buffer periodic supplementation enhancement information message.
[0311] In addition, a video encoder 100 is provided. The video encoder 100 is configured to encode video into a video data stream. Furthermore, the video encoder 100 is configured to generate a video data stream such that the video data stream includes one or more scalable nested supplementary enhancement information messages. Furthermore, the video encoder 100 is configured to generate a video data stream such that one or more scalable nested supplementary enhancement information messages include multiple syntax elements. Furthermore, the video encoder 100 is configured to generate a video data stream such that each syntax element of one or more of the multiple syntax elements is defined to have the same size in each scalable nested supplementary enhancement information message in the video data stream or a portion thereof.
[0312] According to an embodiment, the video encoder 100 may be configured, for example, to generate a video data stream such that the video data stream may include, for example, one or more non-scalable nested supplemental enhancement messages. The video encoder 100 may be configured, for example, to generate a video data stream such that one or more scalable nested supplemental enhancement messages and one or more non-scalable nested supplemental enhancement messages include a plurality of syntax elements. The video encoder 100 may be configured, for example, to generate a video data stream such that each syntax element of one or more of the plurality of syntax elements may be defined, for example, to have the same size in each scalable nested supplemental enhancement message in the video data stream or a portion thereof, and in each non-scalable nested supplemental enhancement message in the video data stream or a portion thereof.
[0313] In an embodiment, the video encoder 100 may, for example, be configured to generate a video data stream such that the video data stream may include, for example, a plurality of access units, wherein each of the plurality of access units may, for example, be assigned to one of a plurality of images of the video. This portion of the video data stream may, for example, be an access unit among the plurality of access units of the video data stream. The video encoder 100 may, for example, be configured to generate a video data stream such that each of one or more syntax elements among a plurality of syntax elements may, for example, be defined as having the same size in each of the scalable nested supplementary enhancement information messages in the scalable nested supplementary enhancement information messages of the access unit.
[0314] According to an embodiment, the video encoder 100 may be configured, for example, to generate a video data stream such that the video data stream may include, for example, one or more non-scalable nested supplemental enhancement information messages. The video encoder 100 may be configured, for example, to generate a video data stream such that one or more scalable nested supplemental enhancement information messages and one or more non-scalable nested supplemental enhancement information messages include multiple syntax elements. The video encoder 100 may be configured, for example, to generate a video data stream such that each syntax element of one or more of the multiple syntax elements may be defined, for example, to have the same size in each scalable nested supplemental enhancement information message in the scalable nested supplemental enhancement information message of the access unit and in each non-scalable nested supplemental enhancement information message of the non-scalable nested supplemental enhancement information message of the access unit.
[0315] In an embodiment, this portion of the video data stream may, for example, be an encoded video sequence of the video data stream. The video encoder 100 may, for example, be configured to generate a video data stream such that each of one or more syntax elements among a plurality of syntax elements may, for example, be defined as having the same size in each scalable nested supplemental enhancement information message within the scalable nested supplemental enhancement information message of the encoded video sequence.
[0316] According to an embodiment, the video encoder 100 may be configured, for example, to generate a video data stream such that the video data stream may include, for example, one or more non-scalable nested supplemental enhancement messages. The video encoder 100 may be configured, for example, to generate a video data stream such that one or more scalable nested supplemental enhancement messages and one or more non-scalable nested supplemental enhancement messages include multiple syntax elements. The video encoder 100 may be configured, for example, to generate a video data stream such that each syntax element of one or more of the multiple syntax elements may be defined, for example, to have the same size in each scalable nested supplemental enhancement message of the encoded video sequence and in each non-scalable nested supplemental enhancement message of the encoded video sequence.
[0317] In an embodiment, the video encoder 100 may be configured, for example, to generate a video data stream such that each of one or more syntax elements among a plurality of syntax elements may be defined, for example, to have the same size in each scalable nested supplemental enhancement information message in a scalable nested supplemental enhancement information message of the video data stream.
[0318] According to an embodiment, the video encoder 100 may be configured, for example, to generate a video data stream such that each of one or more syntax elements among a plurality of syntax elements may be defined, for example, to have the same size in each scalable nested supplementary enhancement message in the scalable nested supplementary enhancement message of the video data stream and in each non-scalable nested supplementary enhancement message of the non-scalable nested supplementary enhancement message of the video data stream.
[0319] In an embodiment, the video encoder 100 may be configured, for example, to generate a video data stream such that the video data stream or a portion thereof may include, for example, at least one buffered periodic supplemental enhancement message, wherein the buffered periodic supplemental enhancement message defines the size of each of one or more of a plurality of syntax elements.
[0320] According to an embodiment, the video encoder 100 may, for example, be configured to generate a video data stream such that the buffer periodicity enhancement information message includes at least one of the following for defining the size of each of one or more syntax elements among a plurality of syntax elements:
[0321] bp_cpb_initial_removal_delay_length_minus1 element,
[0322] The element bp_cpb_removal_delay_length_minus1
[0323] The element bp_dpb_output_delay_length_minus1
[0324] bp_du_cpb_removal_delay_increment_length_minus1 element,
[0325] The element is bp_dpb_output_delay_du_length_minus1.
[0326] In an embodiment, the video encoder 100 may be configured, for example, to generate a video data stream such that each of a plurality of access units of the video data stream includes a scalable nested buffer periodic supplementation enhancement information message, the access unit further including, for example, a non-scalable nested buffer periodic supplementation enhancement information message.
[0327] According to an embodiment, the video encoder 100 may be configured, for example, to generate a video data stream such that each of the plurality of single-level access units of the video data stream includes a scalable nested buffer periodic supplementation enhancement information message, the single-level access unit further including, for example, a non-scalable nested buffer periodic supplementation enhancement information message.
[0328] In addition, an apparatus 200 is provided for receiving an input video data stream. The input video data stream contains encoded video. The apparatus 200 is configured to generate an output video data stream from the input video data stream. The video data stream includes one or more scalable nested supplemental enhancement information messages. Each of the one or more scalable nested supplemental enhancement information messages includes a plurality of syntax elements. Each syntax element of the plurality of syntax elements is defined to have the same size in each scalable nested supplemental enhancement information message within the video data stream or a portion of the video data stream. The apparatus 200 is configured to process the one or more scalable nested supplemental enhancement information messages.
[0329] According to an embodiment, the video data stream may, for example, include one or more non-scalable nested supplemental enhancement messages. The one or more scalable nested supplemental enhancement messages and the one or more non-scalable nested supplemental enhancement messages include a plurality of syntax elements. Each of the one or more syntax elements is defined to have the same size in each scalable nested supplemental enhancement message in the video data stream or that portion of the video data stream, and in each non-scalable nested supplemental enhancement message in the video data stream or that portion of the video data stream. The apparatus 200 is configured to process one or more scalable nested supplemental enhancement messages and one or more non-scalable nested supplemental enhancement messages.
[0330] In an embodiment, the video data stream may, for example, include a plurality of access units, wherein each of the plurality of access units may, for example, be assigned to one of a plurality of images in the video. This portion of the video data stream may, for example, be an access unit among the plurality of access units in the video data stream. Each of one or more syntax elements among the plurality of syntax elements may, for example, be defined as having the same size in each of the scalable nested supplementary enhancement information messages in the scalable nested supplementary enhancement information messages of the access unit.
[0331] According to an embodiment, the video data stream may, for example, include one or more non-scalable nested supplemental enhancement messages. The one or more scalable nested supplemental enhancement messages and the one or more non-scalable nested supplemental enhancement messages include a plurality of syntax elements. Each of the one or more syntax elements may, for example, be defined as having the same size in each scalable nested supplemental enhancement message of the access unit and in each non-scalable nested supplemental enhancement message of the access unit. The apparatus 200 may, for example, be configured to process one or more scalable nested supplemental enhancement messages and one or more non-scalable nested supplemental enhancement messages.
[0332] In an embodiment, this portion of the video data stream may, for example, be an encoded video sequence of the video data stream. Each of one or more of the plurality of syntax elements may, for example, be defined as having the same size in each scalable nested supplementary enhancement information message in the scalable nested supplementary enhancement information message of the encoded video sequence.
[0333] According to an embodiment, the video data stream may, for example, include one or more non-scalable nested supplemental enhancement messages. The one or more scalable nested supplemental enhancement messages and the one or more non-scalable nested supplemental enhancement messages include a plurality of syntax elements. Each of the one or more syntax elements may, for example, be defined as having the same size in each scalable nested supplemental enhancement message of the encoded video sequence and in each non-scalable nested supplemental enhancement message of the encoded video sequence. The apparatus 200 may, for example, be configured to process one or more scalable nested supplemental enhancement messages and one or more non-scalable nested supplemental enhancement messages.
[0334] In an embodiment, each of one or more of the plurality of syntax elements may, for example, be defined as having the same size in each of the scalable nested supplemental enhancement messages in the video data stream.
[0335] According to an embodiment, each of the plurality of syntax elements may be defined, for example, as having the same size in each scalable nested supplementary enhancement message in the scalable nested supplementary enhancement message of the video data stream and in each non-scalable nested supplementary enhancement message of the non-scalable nested supplementary enhancement message of the video data stream. The apparatus 200 may be configured, for example, to process one or more scalable nested supplementary enhancement messages and one or more non-scalable nested supplementary enhancement messages.
[0336] In an embodiment, the video data stream or a portion thereof may, for example, include at least one buffered periodicity enhancement message, wherein the buffered periodicity enhancement message defines the size of one or more syntax elements among a plurality of syntax elements. The apparatus 200 may, for example, be configured to process at least one buffered periodicity enhancement message.
[0337] According to an embodiment, the buffer cycle supplementation enhancement information message includes at least one of the following for defining the size of one or more syntax elements among a plurality of syntax elements:
[0338] bp_cpb_initial_removal_delay_length_minus1 element,
[0339] The element bp_cpb_removal_delay_length_minus1
[0340] The element bp_dpb_output_delay_length_minus1
[0341] bp_du_cpb_removal_delay_increment_length_minus1 element,
[0342] The element is bp_dpb_output_delay_du_length_minus1.
[0343] In an embodiment, for each of the multiple access units of the video data stream that includes a scalable nested buffer periodic supplementation enhancement information message, the access unit may also include, for example, a non-scalable nested buffer periodic supplementation enhancement information message. The apparatus 200 may, for example, be configured to process both the scalable nested supplementation enhancement information message and the non-scalable nested supplementation enhancement information message.
[0344] According to an embodiment, for each of a plurality of single-level access units of a video data stream that includes a scalable nested buffer periodic supplementation enhancement information message, the single-level access unit may also include, for example, a non-scalable nested buffer periodic supplementation enhancement information message. The apparatus 200 may, for example, be configured to process both the scalable nested supplementation enhancement information message and the non-scalable nested supplementation enhancement information message.
[0345] In addition, a video decoder 300 is provided for receiving a video data stream in which video is stored. The video decoder 300 is configured to decode video from the video data stream. The video data stream includes one or more scalable nested supplemental enhancement information messages. Each of the one or more scalable nested supplemental enhancement information messages includes multiple syntax elements. Each syntax element of the one or more syntax elements is defined to have the same size in each scalable nested supplemental enhancement information message in the video data stream or a part of the video data stream. The video decoder 300 is configured to decode the video based on one or more of the multiple syntax elements.
[0346] According to an embodiment, the video data stream may, for example, include one or more non-scalable nested supplemental enhancement messages. The one or more scalable nested supplemental enhancement messages and the one or more non-scalable nested supplemental enhancement messages include a plurality of syntax elements. Each of the one or more syntax elements may, for example, be defined as having the same size in each scalable nested supplemental enhancement message of the video data stream or that portion of the video data stream, and in each non-scalable nested supplemental enhancement message of the video data stream or that portion of the video data stream.
[0347] In an embodiment, the video data stream may, for example, include a plurality of access units, wherein each of the plurality of access units may, for example, be assigned to one of a plurality of images in the video. This portion of the video data stream may, for example, be an access unit among the plurality of access units in the video data stream. Each of one or more syntax elements among the plurality of syntax elements may, for example, be defined as having the same size in each of the scalable nested supplementary enhancement information messages in the scalable nested supplementary enhancement information messages of the access unit.
[0348] According to an embodiment, the video data stream may, for example, include one or more non-scalable nested supplemental enhancement information messages. The one or more scalable nested supplemental enhancement information messages and the one or more non-scalable nested supplemental enhancement information messages include multiple syntax elements. Each syntax element of the multiple syntax elements may, for example, be defined as having the same size in each scalable nested supplemental enhancement information message of the access unit and in each non-scalable nested supplemental enhancement information message of the access unit.
[0349] In an embodiment, this portion of the video data stream may, for example, be an encoded video sequence of the video data stream. Each of one or more of the plurality of syntax elements may, for example, be defined as having the same size in each scalable nested supplementary enhancement information message in the scalable nested supplementary enhancement information message of the encoded video sequence.
[0350] According to an embodiment, the video data stream may, for example, include one or more non-scalable nested supplemental enhancement messages. The one or more scalable nested supplemental enhancement messages and the one or more non-scalable nested supplemental enhancement messages include a plurality of syntax elements. Each of the one or more syntax elements may, for example, be defined as having the same size in each scalable nested supplemental enhancement message of the encoded video sequence and in each non-scalable nested supplemental enhancement message of the encoded video sequence.
[0351] In an embodiment, each of one or more of the plurality of syntax elements may, for example, be defined as having the same size in each of the scalable nested supplemental enhancement messages in the video data stream.
[0352] According to an embodiment, each of the plurality of syntax elements may be defined, for example, as having the same size in each scalable nested supplementary enhancement message in the scalable nested supplementary enhancement message of the video data stream and in each non-scalable nested supplementary enhancement message of the non-scalable nested supplementary enhancement message of the video data stream.
[0353] In an embodiment, the video data stream or the portion thereof may include, for example, at least one buffered periodic supplementation enhancement message, wherein the buffered periodic supplementation enhancement message defines the size of each of one or more syntax elements among a plurality of syntax elements.
[0354] According to an embodiment, the buffer cycle supplementation enhancement information message includes at least one of the following for defining the size of each of one or more syntax elements among a plurality of syntax elements:
[0355] bp_cpb_initial_removal_delay_length_minus1 element,
[0356] The element bp_cpb_removal_delay_length_minus1
[0357] The element bp_dpb_output_delay_length_minus1
[0358] bp_du_cpb_removal_delay_increment_length_minus1 element,
[0359] The element is bp_dpb_output_delay_du_length_minus1.
[0360] In an embodiment, for each of the multiple access units of a video data stream that includes a scalable nested buffer periodic supplementation enhancement information message, the access unit may also include, for example, a non-scalable nested buffer periodic supplementation enhancement information message.
[0361] According to an embodiment, for each of the multiple single-level access units of a video data stream that includes a scalable nested buffer periodic supplementation enhancement information message, the single-level access unit may also include, for example, a non-scalable nested buffer periodic supplementation enhancement information message.
[0362] In addition, a system is provided. This system includes the device 200 as described above and the video decoder 300 as described above. The video decoder 300 is configured to receive the output video data stream of the device 200. Furthermore, the video decoder 300 is configured to decode the video from the output video data stream of the device 200.
[0363] According to an embodiment, the system may also include, for example, a video encoder 100. The device 200 may be configured, for example, to receive a video data stream from the video encoder 100 as an input video data stream.
[0364] Specifically, the fifth aspect of the invention relates to constraining all BP SEI messages in a bitstream to indicate certain variable-length encoded syntax elements of the same length and not being scalable nested in the absence of non-scalable nested variants in the same AU.
[0365] Buffered Periodic SEI messages, Picture Timing SEI messages, and Decoding Unit Information SEI messages provide precise timing information for NAL units within the bitstream to control their transition through the decoder's buffer during conformance testing. Some syntax elements in PT and DUI SEI messages are encoded with variable lengths, and the length of these syntax elements is transmitted in the BP SEI message. This parsing dependency is a design trade-off. The benefit of saving the transmission of these length syntax elements at each PT or DUI SEI message is achieved to offset the cost of not being able to parse PT and DUI SEI messages without first parsing the associated BP SEI message. Since BP SEI messages (once per multiple frames) are transmitted much less frequently than PT messages (once per frame) or DUI SEI messages (multiple times per frame), this bit saving is achieved through a common design trade-off, similar to how the picture header structure can reduce the bit cost of the slice header when using many slices.
[0366] More specifically, the BP SEI message in the current VVC draft specification includes a syntax element that serves as the root for resolving dependencies:
[0367] ●bp_cpb_initial_removal_delay_length_minus1, which specifies the encoded length of the initial CPB removal delay for the candidate AU in the PT SEI message, and
[0368] ●bp_cpb_removal_delay_length_minus1, which specifies the encoded length of the CPB removal delay and removal delay offset of the AU in the PT SEI message, and
[0369] ●bp_dpb_output_delay_length_minus1, which specifies the encoded length of the DPB output delay of the AU in the PT SEI message, and
[0370] ●bp_du_cpb_removal_delay_increment_length_minus1, which specifies the encoded length of the individual CPB removal delay and the common CPB removal delay of the DU in the PT SEI message, and the CPB removal delay of the DU in the DUI SEI message, and
[0371] ●bp_dpb_output_delay_du_length_minus1, which specifies the encoded length of the DPB output delay of the AU in the PT SEI message and DU SEI message.
[0372] However, problems arise when the bitstream contains multiple OLSs. While the BP / PT / DUI SEI messages applicable to the OLS representing the bitstream are carried verbatim in the bitstream, tracking parsing dependencies is trivial. Other pairs of BP / PT / DUI SEI messages corresponding to the OLS representing the (sub)bitstream are carried in encapsulated form within so-called scalable nested SEI messages. Parsing dependencies still apply, and given that the number of OLSs can be very large, it is a considerable burden for the decoder or parser to track the correctly encapsulated BP SEI messages for parsing dependencies when processing encapsulated PT and DUI SEI messages. In particular, this is exacerbated by the fact that these messages can also be encapsulated within different scalable nested SEI messages.
[0373] Therefore, as part of this invention, in one embodiment, the following bitstream constraint is established: the encoded value of the corresponding syntax element describing the length must be the same in all scalable nested and non-scalable nested BP SEI messages in the AU. Thus, the decoder or parser only needs to store the corresponding length value when parsing the first non-scalable BP SEI message in the AU, and can resolve the parsing dependencies of all PT messages and DUI SEI messages within a buffer cycle starting at the corresponding AU, regardless of whether these messages are encapsulated in scalable nested SEI messages. The following is an example of the corresponding canonical text:
[0374] The requirement for bitstream consistency is that all scalable nested and non-scalable nested buffered periodic SEI messages in the AU have the same corresponding value for the syntax elements bp_cpb_initial_removal_delay_length_minus1, bp_cpb_removal_delay_length_minus1, bp_dpb_output_delay_length_minus1, bp_du_cpb_removal_delay_increment_length_minus1, and bp_dpb_output_delay_du_length_minus1.
[0375] In another embodiment, this constraint applies only to scalable nested BP SEI messages within the buffer period determined by the current non-scalable nested BP SEI message, and is described as follows:
[0376] The requirement for bitstream consistency is that all scalable nested buffered cycle SEI messages within a buffered cycle have the same corresponding value for the syntax elements bp_cpb_initial_removal_delay_length_minus1, bp_cpb_removal_delay_length_minus1, bp_dpb_output_delay_length_minus1, bp_du_cpb_removal_delay_increment_length_minus1, and bp_dpb_output_delay_du_length_minus1then, followed by non-scalable nested buffered cycle SEI messages within the buffered cycle.
[0377] Here, the bitstream's BP defines the constrained range of the scalable nested BP from one scalable nested BP to the next.
[0378] In another embodiment, this constraint applies to all AU representations of the bitstream, for example, as follows:
[0379] The requirement for bitstream consistency is that all scalable nested and non-scalable nested buffered periodic SEI messages in the bitstream have the same corresponding value for the syntax elements bp_cpb_initial_removal_delay_length_minus1, bp_cpb_removal_delay_length_minus1, bp_dpb_output_delay_length_minus1, bp_du_cpb_removal_delay_increment_length_minus1, and bp_dpb_output_delay_du_length_minus1.
[0380] In another embodiment, the constraints apply only to the AU representation in CVS, so the smart encoder can still facilitate differences in the duration of BP in the bitstream for encoding relevant delay and offset syntax elements. The canonical text will be as follows:
[0381] The requirement for bitstream consistency is that all scalable nested and non-scalable nested buffered cycle SEI messages in CVS have the same corresponding value for the syntax elements bp_cpb_initial_removal_delay_length_minus1, bp_cpb_removal_delay_length_minus1, bp_dpb_output_delay_length_minus1, bp_du_cpb_removal_delay_increment_length_minus1, and bp_dpb_output_delay_du_length_minus1.
[0382] Here, the constraint range is CVS.
[0383] More specifically, a buffer period, or BP SEI message, defines a so-called buffer period in which the timing of each picture uses the picture at the beginning of the buffer period as an anchor. For example, the start of a buffer period can help test the consistency of random access functionality in a bitstream.
[0384] Figure 7 Two sets of HRD SEIs (scalable nested SEIs and non-scalable nested SEIs) in a two-layer bitstream according to an embodiment are shown.
[0385] In such Figure 7In the multi-layered scene shown, for example, the scalable nested HRD SEI provides a different buffer cycle setting (through BP at POC 0 and POC 3) than the non-scalable nested SEI (POC 0 only) that is used when only layer L0 is extracted and played from POC 3.
[0386] However, this also increases the complexity and cost of tracking the parsing correlations between PT and individual BP messages as described above, which is undesirable. Therefore, as part of the invention, in one embodiment, scalable nested BP SEI messages are prohibited in AUs that do not have non-scalable nested BP SEI messages, as follows:
[0387] The requirement for bitstream consistency is that scalable nested BP SEI messages should not be present in AUs that do not contain non-scalable nested BP SEI messages.
[0388] Since the above use cases are limited to multi-layer bitstreams, in another embodiment, the relevant constraints are limited to single-layer bitstreams, as follows:
[0389] The requirement for bitstream consistency is that scalable nested BP SEI messages should not be present in a single-level AU that does not contain non-scalable nested BP SEI messages.
[0390] Although some aspects have been described in the context of the apparatus, it will be clear that these aspects also represent a description of the corresponding method, wherein a block or device corresponds to a method step or a feature of a method step. Similarly, aspects described in the context of method steps also represent a description of the features of the corresponding block or item or the corresponding apparatus. Some or all of the method steps may be performed by (or using) hardware devices (such as microprocessors, programmable computers, or electronic circuits). In some embodiments, one or more of the most important method steps may be performed by such an apparatus.
[0391] Depending on certain implementation requirements, embodiments of the present invention can be implemented in hardware or software, or at least partially in hardware or at least partially in software. Implementation can be performed using a digital storage medium (e.g., floppy disk, DVD, Blu-ray, CD, ROM, PROM, EPROM, EEPROM, or flash memory) on which electronically readable control signals are stored, which cooperate with (or are capable of cooperating with) a programmable computer system to perform the corresponding methods. Therefore, the digital storage medium can be computer-readable.
[0392] Some embodiments of the invention include a data carrier having electronically readable control signals, which is capable of cooperating with a programmable computer system to perform one of the methods described herein.
[0393] Typically, embodiments of the present invention can be implemented as a computer program product having program code operable to perform one of the methods when the computer program product is run on a computer. The program code may, for example, be stored on a machine-readable medium.
[0394] Other embodiments include a computer program stored on a machine-readable medium for performing one of the methods described herein.
[0395] In other words, embodiments of the method of the present invention are therefore computer programs having program code for performing one of the methods described herein when the computer program is run on a computer.
[0396] Therefore, another embodiment of the method of the present invention is a data carrier (or digital storage medium or computer-readable medium) on which a computer program is recorded, the computer program being used to perform one of the methods described herein. The data carrier, digital storage medium, or recorded medium is typically tangible and / or non-transitory.
[0397] Therefore, another embodiment of the method of the present invention represents a data stream or signal sequence of a computer program used to perform one of the methods described herein. The data stream or signal sequence may, for example, be configured to be transmitted via a data communication connection (e.g., via the Internet).
[0398] Another embodiment includes a processing means, such as a computer or a programmable logic device, which is configured or adapted to perform one of the methods described herein.
[0399] Another embodiment includes a computer having a computer program installed thereon for performing one of the methods described herein.
[0400] Another embodiment of the invention includes an apparatus or system configured to transmit a computer program to a receiver (e.g., electronically or optically) for performing one of the methods described herein. The receiver may be, for example, a computer, mobile device, storage device, etc. The apparatus or system may, for example, include a file server for transmitting the computer program to the receiver.
[0401] In some embodiments, a programmable logic device (e.g., a field-programmable gate array) may be used to perform some or all of the functions described herein. In some embodiments, the field-programmable gate array may cooperate with a microprocessor to perform one of the methods described herein. Generally, the methods are preferably performed by any hardware device.
[0402] The apparatus described herein can be implemented using hardware devices, a computer, or a combination of hardware devices and a computer.
[0403] The methods described herein can be performed using hardware devices, computers, or a combination of hardware devices and computers.
[0404] The above embodiments are merely illustrative of the principles of the present invention. It should be understood that modifications and variations of the arrangements and details described herein will be readily apparent to those skilled in the art. Therefore, the invention is intended to be limited only by the scope of the appended claims and not by the specific details given by way of the description and explanation of the embodiments herein.
[0405] References
[0406] [1] ISO / IEC, ITU-T. High efficiency video coding. ITU-T Recommendation H.265 | ISO / IEC 23008 10 (HEVC), edition 1, 2013; edition 2, 2014.
Claims
1. A video decoder for receiving a video data stream containing stored video data. in, The video decoder is configured to decode the video from the video data stream. The video data stream includes an indication indicating whether a non-scalable nested image timing supplemental enhancement information message of a network abstraction layer unit ... Wherein, if the indication has a first value, the non-scalable nested image timed supplementation enhancement information message of the network abstraction layer unit of the access unit is defined as applicable to all output layer sets in the plurality of output layer sets of the access unit. Wherein, if the indication has a value different from the first value, then the indication does not define whether the non-scalable nested image timed supplementation enhancement information message of the network abstraction layer unit of the access unit applies to all output layer sets of the multiple output layer sets of the access unit. The video decoder is configured to decode the video according to the instructions.
2. The video decoder according to claim 1, in, If the indication has the first value, then the network abstraction layer unit does not include any other supplementary enhancement information messages that are different from the image timed supplementary enhancement information message.
3. The video decoder according to claim 1 or 2, in, If the indication has the first value, then the network abstraction layer unit does not include any other supplementary enhancement information messages.
4. The video decoder according to any one of claims 1 to 3, in, If the indication has the first value, then for each of the multiple access units of the encoded video sequence in the one or more encoded video sequences, each network abstraction layer unit that includes non-scalable nested picture timing supplemental enhancement information messages does not include any other supplemental enhancement information messages different from the picture timing supplemental enhancement information messages, or does not include any other supplemental enhancement information messages.
5. The video decoder according to any one of claims 1 to 3, in, If the indication has the first value, then for each of the plurality of access units in each of the one or more coded video sequences of the video data stream, each network abstraction layer unit includes a non-scalable nested picture timing supplemental enhancement information message, the network abstraction layer unit does not include any other supplemental enhancement information message different from the picture timing supplemental enhancement information message, or does not include any other supplemental enhancement information message.
6. A video encoder, in, The video encoder is configured to encode video into a video data stream. The video encoder is configured to generate the video data stream such that the video data stream includes an indication indicating whether a non-scalable nested image timing supplementation enhancement information message of a network abstraction layer unit of a network abstraction layer unit of a plurality of access units of a coded video sequence in one or more coded video sequences of the video data stream is defined as applicable to all output layer sets in a plurality of output layer sets of the access unit. Wherein, if the indication has a first value, the non-scalable nested image timed supplementation enhancement information message of the network abstraction layer unit of the access unit is defined as applicable to all output layer sets in the plurality of output layer sets of the access unit. Wherein, if the indication has a value different from the first value, then the indication does not define whether the non-scalable nested image timed supplementation enhancement information message of the network abstraction layer unit of the access unit applies to all output layer sets of the plurality of output layer sets of the access unit.
7. The video encoder according to claim 6, in, If the indication has the first value, then the network abstraction layer unit does not include any other supplementary enhancement information messages that are different from the image timed supplementary enhancement information message.
8. The video encoder according to claim 6 or 7, in, If the indication has the first value, then the network abstraction layer unit does not include any other supplementary enhancement information messages.
9. The video encoder according to any one of claims 6 to 8, in, If the indication has the first value, the video encoder is configured to generate the video data stream such that for each of the multiple access units of the encoded video sequence in the one or more encoded video sequences, each network abstraction layer unit includes a non-scalable nested image timing supplemental enhancement information message, the network abstraction layer unit does not include any other supplemental enhancement information message different from the image timing supplemental enhancement information message, or does not include any other supplemental enhancement information message.
10. The video encoder according to any one of claims 6 to 8, in, If the indication has the first value, the video encoder is configured to generate the video data stream such that for each of the plurality of access units of each of the one or more coded video sequences in the video data stream, each network abstraction layer unit includes a non-scalable nested picture timing supplemental enhancement information message, the network abstraction layer unit does not include any other supplemental enhancement information message different from the picture timing supplemental enhancement information message, or does not include any other supplemental enhancement information message.
11. A video data stream, in, The video data stream contains encoded video. The video data stream includes an indication indicating whether a non-scalable nested image timing supplemental enhancement information message of a network abstraction layer unit ... Wherein, if the indication has a first value, the non-scalable nested image timed supplementation enhancement information message of the network abstraction layer unit of the access unit is defined as applicable to all output layer sets in the plurality of output layer sets of the access unit. Wherein, if the indication has a value different from the first value, then the indication does not define whether the non-scalable nested image timed supplementation enhancement information message of the network abstraction layer unit of the access unit applies to all output layer sets of the plurality of output layer sets of the access unit.
12. The video data stream according to claim 11, in, If the indication has the first value, then the network abstraction layer unit does not include any other supplementary enhancement information messages that are different from the image timed supplementary enhancement information message.
13. The video data stream according to claim 11 or 12, in, If the indication has the first value, then the network abstraction layer unit does not include any other supplementary enhancement information messages.
14. The video data stream according to any one of claims 11 to 13, in, If the indication has the first value, then for each of the multiple access units of the encoded video sequence in the one or more encoded video sequences, each network abstraction layer unit including non-scalable nested picture timing supplemental enhancement information messages does not include any other supplemental enhancement information messages different from the picture timing supplemental enhancement information messages, or does not include any other supplemental enhancement information messages.
15. The video data stream according to any one of claims 11 to 13, in, If the indication has the first value, then for each of the plurality of access units in each of the one or more coded video sequences of the video data stream, each network abstraction layer unit includes a non-scalable nested picture timing supplemental enhancement information message, the network abstraction layer unit does not include any other supplemental enhancement information message different from the picture timing supplemental enhancement information message, or does not include any other supplemental enhancement information message.
16. A method for receiving a video data stream containing stored video data. in, The method includes: decoding the video from the video data stream. The video data stream includes an indication indicating whether a non-scalable nested image timing supplemental enhancement information message of a network abstraction layer unit ... Wherein, if the indication has a first value, the non-scalable nested image timed supplementation enhancement information message of the network abstraction layer unit of the access unit is defined as applicable to all output layer sets in the plurality of output layer sets of the access unit. Wherein, if the indication has a value different from the first value, then the indication does not define whether the non-scalable nested image timed supplementation enhancement information message of the network abstraction layer unit of the access unit applies to all output layer sets of the multiple output layer sets of the access unit. Decoding the video is performed according to the instructions.
17. A method for encoding video into a video data stream, in, The method includes: generating the video data stream such that the video data stream includes an indication indicating whether a non-scalable nested image timing supplementation enhancement information message of a network abstraction layer unit of a plurality of access units of an access unit in one or more coded video sequences of the video data stream is defined as applicable to all output layer sets in a plurality of output layer sets of the access unit. Wherein, if the indication has a first value, the non-scalable nested image timed supplementation enhancement information message of the network abstraction layer unit of the access unit is defined as applicable to all output layer sets in the plurality of output layer sets of the access unit. Wherein, if the indication has a value different from the first value, then the indication does not define whether the non-scalable nested image timed supplementation enhancement information message of the network abstraction layer unit of the access unit applies to all output layer sets of the plurality of output layer sets of the access unit.
18. A non-transitory computer-readable medium comprising a computer program for implementing the method according to claim 16 or 17 when executed on a computer or signal processor.