Video data streams, video encoders, apparatuses and methods for hypothetical reference decoders and for output layer sets
By introducing a playback speed modification factor and seamless stitching point management, the parallel processing capability in the video encoding and decoding process is optimized, solving the technical problems of video encoding and decoding in the existing technology, realizing more efficient video data stream management and parallel processing capability, and improving the efficiency of video encoding and decoding.
Patent Information
- Application Number
- CN202511179258.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2019-12-20
- Filing Date
- 2020-12-18
- Publication Date
- 2025-11-28
AI Technical Summary
Existing video encoding and decoding technologies are insufficient in terms of parallel processing capabilities and video data stream management, especially in the H.265/HEVC standard, which struggles to effectively support parallel processing of video encoders and decoders and efficient management of data streams.
By introducing a playback speed modification factor (Nx), video encoders and decoders can generate or process video data streams based on decoding capability requirements. These video data streams include information about sampling and/or video data streams containing multiple access units, including information about bitrate limits. Furthermore, by generating or processing video data streams to include mapping information for seamless stitching points and operation points, the management of the video data streams is optimized.
It improves the parallel processing capabilities in video encoding and decoding, enables more efficient video data stream management, supports seamless splicing of video data streams and optimization of operation points, and enhances the efficiency and flexibility of video encoding and decoding.
Smart Images

Figure CN121037573A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] This application is a divisional application of application number: 202080088769.7, invention name: Video data stream, video encoder, apparatus and method for hypothetical reference decoder and output layer set. The present invention relates to video encoding and video decoding, and in particular to a video encoder, to a video decoder, to methods for encoding and decoding, and to a video data stream for implementing advanced video coding concepts. BACKGROUND
[0002] H.265 / HEVC (HEVC = High Efficiency Video Coding) is a video codec which already provides tools for improving or even enabling parallel processing at the encoder and / or decoder. For example, HEVC supports the subdivision of a picture into an array of tiles which are coded independently from each other. Another concept supported by HEVC relates to WPP according to which CTU rows or CTU columns of a picture can be processed in parallel from left to right (e.g. in form of slices) as long as a certain minimum CTU offset is respected in the processing of consecutive CTU rows (CTU = Coding Tree Unit). However, it would be advantageous to have a video codec which even more efficiently supports the parallel processing capabilities of a video encoder and / or video decoder.
[0003] Generally, in video coding, the coding process of picture samples requires smaller partitions in which samples are divided into rectangular regions for joint processing such as prediction or transform coding. Thus, pictures are partitioned into blocks of a certain size which is constant during encoding of a video sequence. In the H.264 / AVC standard, fixed size blocks of 16x16 samples, so-called macroblocks (AVC = Advanced Video Coding), are used.
[0004] In the state-of-the-art HEVC standard (see [1]), there are coding tree blocks (CTB) or coding tree units (CTU) of a maximum size of 64x64 samples. In further descriptions of HEVC, the more common terminology of CTU is used for such blocks.
[0005] CTUs are processed in a raster scan order, starting with the top-left CTU, processing the CTUs in a picture row by row down to the bottom-right CTU.
[0006] The decoded CTU data is organized into a container called a slice. Initially, in previous video coding standards, a slice referred to a segment comprising one or more consecutive CTUs of a picture. A slice was taken for the segmentation of the decoded data. From another perspective, a complete picture can also be defined as one large segment and, thus, historically, the term slice still applies. In addition to the decoded picture samples, a slice also comprises additional information related to the coding process of the slice itself, which is placed into a so-called slice header.
[0007] According to the current state of the art, the VCL (Video Coding Layer) also comprises techniques for fragmentation and spatial partitioning. For various reasons, such partitioning can be applied in video coding, among them load balancing in parallelization, CTU size matching in network transmission, error-mitigation, etc.
[0008] A bitstream as specified in a video coding standard has information associated with HRD conformance. This conformance consists of a hypothetical reference decoder (HRD) comprising a buffer model which assumes that NAL units enter a coded picture buffer (CPB) in front of the decoder and are removed therefrom at specific timing to ensure that the CPB size is not overrun or that NAL units do not arrive later than they need to be removed (buffer underrun). In addition, the model consists of a decoded picture buffer (DPB) from which decoded pictures are outputted when they are no longer needed for prediction and, in many implementations, the size of the decoded pictures is also limited. The timing information of the HRD is conveyed in the bitstream through so-called SEI messages, in particular, the buffering period (BP) SEI message which defines the timing information of the buffering period (multiple access units or AUs), the picture timing (PT) SEI message which conveys the timing information of a single associated AU, and the decoded unit information (DUI) SEI message which conveys the timing information of an associated subset of AUs, i.e. a decoding unit or DU. SUMMARY
[0009] It is an object of the present application to provide an improved concept for video encoding and video decoding.
[0010] The object of the present application is solved by the subject-matter of the independent claims.
[0011] Preferred embodiments are provided in the dependent claims.
[0012] According to embodiments, there is provided an apparatus for receiving a video data stream as an input bitstream, wherein the video data stream has a video encoded into it. For a playback speed modification factor (Nx), the apparatus is configured to determine a decoding capability requirement in dependence on a decoding capability requirement limit. The playback speed modification factor (Nx) is either a forward playback speed modification factor or a backward playback speed modification factor.
[0013] Further, according to embodiments, there is provided a video data stream having a video encoded into it. The video data stream comprises information on a sampling rate limit, and / or the video data stream comprises information on a code rate limit.
[0014] Further, according to embodiments, there is provided a video encoder for encoding a video into a video data stream, such that the video data stream has a video encoded into it. The video encoder is configured to generate the video data stream such that the video data stream comprises information on a sampling rate limit, and / or the video encoder is configured to generate the video data stream such that the video data stream comprises information on a code rate limit.
[0015] Further, according to embodiments, there is provided a method for receiving a video data stream as an input bitstream, wherein the video data stream has a video encoded into it. For a playback speed modification factor (Nx), the method comprises determining a decoding capability requirement in dependence on a decoding capability requirement limit. The playback speed modification factor (Nx) is either a forward playback speed modification factor or a backward playback speed modification factor.
[0016] Further, according to embodiments, there is provided a method for encoding a video into a video data stream, such that the video data stream has a video encoded into it. The method comprises generating the video data stream such that the video data stream comprises information on a sampling rate limit, and / or the method comprises generating the video data stream such that the video data stream comprises information on a code rate limit.
[0017] Further, there is provided a computer program for implementing one of the methods as described before when executed on a computer or signal processor.
[0018] Further, according to embodiments, there is provided a video data stream having a video encoded into it. The video data stream comprises a plurality of access units, wherein a group of access units of the video data stream comprises the plurality of access units. Further, the video data stream comprises an indication for each access unit of a subset of access units indicating that there is a seamless splice point after the access unit. The subset of access units is a proper subset of the group of access units of the video data stream.
[0019] Furthermore, according to an embodiment, a video encoder is provided for encoding video into a video data stream, such that the video data stream has video encoded therein. The video encoder is used to generate the video data stream such that the video data stream includes a plurality of access units, wherein a group of access units in the video data stream includes the plurality of access units. Additionally, the video encoder is used to generate the video data stream such that the video data stream includes an indication for each access unit of a subset of access units, indicating the existence of a seamless stitching point after the access unit. Furthermore, the video encoder is used to generate the video data stream such that the subset of access units is an appropriate subset of the group of access units in the video data stream.
[0020] Furthermore, according to an embodiment, an apparatus is provided for receiving a video data stream as an input bit stream, wherein the video data stream has video encoded therein. The apparatus is used to process the video data stream. The video data stream includes a plurality of access units, wherein a group of access units of the video data stream includes the plurality of access units. Furthermore, the video data stream includes an indication for each access unit of a subset of access units, indicating the existence of a seamless splicing point after the access unit. The subset of access units is an appropriate subset of the group of access units of the video data stream. The apparatus is used to process the indication.
[0021] Furthermore, according to embodiments, a method for encoding video into a video data stream is provided, such that the video data stream has video encoded therein. The method includes generating the video data stream such that the video data stream includes a plurality of access units, wherein a group of access units in the video data stream includes the plurality of access units. Furthermore, the method includes generating the video data stream such that the video data stream includes an indication for each access unit of a subset of access units, indicating the existence of a seamless splicing point after the access unit; additionally, the method includes generating the video data stream such that the subset of access units is an appropriate subset of the group of access units in the video data stream.
[0022] Furthermore, according to an embodiment, a method is provided for receiving a video data stream as an input bit stream, wherein the video data stream has video encoded therein. The method includes processing the video data stream, which includes a plurality of access units, wherein a group of access units in the video data stream includes the plurality of access units. Additionally, the video data stream includes an indication for each access unit of a subset of access units, indicating the existence of a seamless splicing point after the access unit. The subset of access units is an appropriate subset of the group of access units in the video data stream. The method includes processing the indication.
[0023] In addition, a computer program is provided for implementing one of the methods as described above when executed on a computer or signal processor.
[0024] Furthermore, according to an embodiment, an apparatus is provided for receiving a video data stream as an input bit stream, wherein the video data stream has video encoded therein. The apparatus is configured to process the input bit stream to obtain sub-bits based on an output layer set. If the output layer set does not include a predefined layer among a plurality of layers present in the video data stream, the apparatus is configured to remove at least one non-scalable nested supplementary enhancement information message allocated with the predefined layer. If the output layer set includes the predefined layer, the apparatus is configured not to remove any non-scalable nested supplementary enhancement information messages allocated with the predefined layer.
[0025] Furthermore, according to an embodiment, a video data stream having video encoded therein is provided. The video data stream includes a set of video parameters. The video parameter set includes multiple sets of output layers. If none of the multiple sets of output layers includes all layers present in the multiple layers of the video data stream, then the video data stream does not include any non-scalable nested supplementary enhancement information messages used for assuming a reference decoder.
[0026] Furthermore, according to an embodiment, a video encoder is provided for encoding video into a video data stream, such that the video data stream has video encoded therein. The video encoder is used to generate the video data stream such that the video data stream includes a set of video parameters. Furthermore, the video encoder is used to generate the video data stream such that the video parameter set includes a plurality of output layer sets. If none of the plurality of output layer sets includes all layers present in the plurality of layers in the video data stream, then the video encoder is used to generate the video data stream such that the video data stream does not include any non-expandable nested supplementary enhancement information messages used for assuming a reference decoder.
[0027] Furthermore, according to an embodiment, a method for receiving a video data stream as an input bitstream is provided, wherein the video data stream has video encoded therein. The method includes processing the input bitstream based on an output layer set to obtain a sub-bitstream. If the output layer set does not include a predefined layer (vps_layer_id[0]) among a plurality of layers present in the video data stream, the method includes removing at least one non-expandable nested supplemental enhancement information message allocated with the predefined layer (vps_layer_id[0]). If the output layer set includes the predefined layer (vps_layer_id[0]), the method includes not removing any non-expandable nested supplemental enhancement information messages allocated with the predefined layer (vps_layer_id[0]).
[0028] Furthermore, according to an embodiment, a method is provided for encoding video into a video data stream, such that the video data stream has video encoded therein. The method includes generating the video data stream such that the video data stream includes a set of video parameters. Furthermore, the method includes generating the video data stream such that the video parameter set includes a plurality of output layer sets. If none of the plurality of output layer sets includes all layers present in the plurality of layers in the video data stream, then the method includes generating the video data stream such that the video data stream does not include any non-expandable nested supplementary enhancement information messages used for assuming a reference decoder.
[0029] Furthermore, according to an embodiment, an apparatus is provided for receiving a video data stream as an input bit stream, wherein the video data stream has video encoded therein. The video data stream includes initial profile hierarchy information. Furthermore, the video data stream includes a first group of two or more layers. The apparatus is configured to process the input bit stream using an output layer set depending on a plurality of output layer sets to generate an output bit stream by removing at least one layer from the first group of two or more layers to obtain a second group of one or more layers, such that the output bit stream includes the second group of one or more layers but excludes the at least one layer removed from the first group. Furthermore, the apparatus is configured to remove portions of the initial profile hierarchy information that are unrelated to at least one layer in the second group of one or more layers.
[0030] Furthermore, according to an embodiment, an apparatus is provided for receiving a video data stream as an input bitstream, wherein the video data stream has video encoded therein. The video data stream includes expandable nested supplemental enhancement information, which includes decoding parameter set information for each of a plurality of output layer sets. The apparatus is configured to process the input bitstream based on the output layer sets of the plurality of output layer sets to obtain an output bitstream, such that the output bitstream includes the decoding parameter set information for the output layer sets, and such that the output bitstream does not include the decoding parameter set information for any other output layer set among the plurality of output layer sets.
[0031] Furthermore, according to an embodiment, a video data stream having video encoded therein is provided. The video data stream includes scalable nested supplemental enhancement information, which includes decoding parameter set information for each of a plurality of output layer sets.
[0032] Furthermore, according to an embodiment, a video encoder is provided for encoding video into a video data stream, such that the video data stream has video encoded therein. The video encoder is used to generate the video data stream such that the video data stream includes scalable nested supplementary enhancement information, the scalable nested supplementary enhancement information including decoding parameter set information for each of a plurality of output layer sets.
[0033] Furthermore, according to an embodiment, a video data stream having video encoded therein is provided. The video data stream includes indications of a plurality of operation points. Additionally, the video data stream includes a first mapping for each of the plurality of operation points, the first mapping assigning one or more profile levels from a plurality of profile levels to the operation point. Furthermore, the video data stream includes a plurality of second mappings, wherein each of the plurality of mappings assigns one of the plurality of operation points to a set of output levels from a plurality of output level sets.
[0034] Furthermore, according to an embodiment, a video encoder is provided for encoding video into a video data stream, such that the video data stream has video encoded therein. The video encoder is used to generate the video data stream such that the video data stream includes indications of a plurality of operation points. Furthermore, the video encoder is used to generate the video data stream such that the video data stream includes a first mapping for each of the plurality of operation points, the first mapping assigning one or more profile levels from a plurality of profile levels to the operation point. Furthermore, the video encoder is used to generate the video data stream such that the video data stream includes a plurality of second mappings, wherein each of the plurality of mappings assigns one of the plurality of operation points to an output layer set from a plurality of output layer sets.
[0035] Furthermore, according to an embodiment, an apparatus is provided for receiving a video data stream as an input bit stream, wherein the video data stream has video encoded therein. The video data stream includes indications of a plurality of operation points. Furthermore, the video data stream includes a first mapping for each of the plurality of operation points, the first mapping mapping one or more profile levels from a plurality of profile levels to the operation point. Furthermore, the video data stream includes a plurality of second mappings, wherein each of the plurality of mappings maps one of the plurality of operation points to an output level from a plurality of output level sets. The apparatus is used to process the input bit stream based on the output level set from the plurality of output level sets and based on at least one of the second mappings mapping one of the plurality of operation points to the output level set to produce an output bit stream.
[0036] Furthermore, according to an embodiment, a method is provided for receiving a video data stream as an input bitstream, wherein the video data stream has video encoded therein. The video data stream includes initial profile hierarchy information. Furthermore, the video data stream includes a first set of two or more layers. The method includes processing the input bitstream using an output layer set that depends on a plurality of output layer sets to generate an output bitstream by removing at least one layer from the first set of two or more layers to obtain a second set of one or more layers, such that the output bitstream includes the second set of one or more layers, but excludes the at least one layer removed from the first set. Furthermore, the method includes removing portions of the initial profile hierarchy information that are unrelated to at least one layer in the second set of one or more layers.
[0037] Furthermore, according to an embodiment, a method is provided for receiving a video data stream as an input bitstream, wherein the video data stream has video encoded therein. The video data stream includes expandable nested supplemental enhancement information, which includes decoding parameter set information for each of a plurality of output layer sets. The method includes processing the input bitstream based on the output layer sets of the plurality of output layer sets to obtain an output bitstream, such that the output bitstream includes the decoding parameter set information for the output layer sets, and such that the output bitstream does not include the decoding parameter set information for any other output layer set among the plurality of output layer sets.
[0038] Furthermore, according to embodiments, a method for encoding video into a video data stream is provided, such that the video data stream has video encoded therein. The method includes generating the video data stream such that the video data stream includes scalable nested supplemental enhancement information, the scalable nested supplemental enhancement information including decoding parameter set information for each of a plurality of output layer sets.
[0039] Furthermore, according to an embodiment, a method for encoding video into a video data stream is provided, such that the video data stream has video encoded therein. The method includes generating the video data stream such that the video data stream includes indications of a plurality of operation points. Furthermore, the method includes generating the video data stream such that the video data stream includes a first mapping for each of the plurality of operation points, the first mapping assigning one or more profile levels from a plurality of profile levels to the operation point. Furthermore, the method includes generating the video data stream such that the video data stream includes a plurality of second mappings, wherein each of the plurality of mappings assigns one of the plurality of operation points to an output layer set from a plurality of output layer sets.
[0040] Furthermore, according to an embodiment, a method for receiving a video data stream as an input bit stream is provided, wherein the video data stream has video encoded therein. The video data stream includes indications of a plurality of operation points. Furthermore, the video data stream includes a first mapping for each of the plurality of operation points, the first mapping mapping one or more profile levels from a plurality of profile levels to the operation point. Furthermore, the video data stream includes a plurality of second mappings, wherein each of the plurality of mappings maps one of the plurality of operation points to an output level from a plurality of output level sets. The method includes processing the input bit stream based on at least one of the second mappings that maps one of the plurality of operation points to the output level set, to produce an output bit stream.
[0041] In addition, a computer program is provided for implementing one of the methods as described above when executed on a computer or signal processor.
[0042] According to an embodiment, a video data stream having video encoded therein is provided. The video data stream includes a sequence parameter set and a decoding parameter set. The sequence parameter set includes a first indication value, which indicates a maximum number of sublayers or the maximum number of sublayers minus a constant value based on information stored in the sequence parameter set. The decoding parameter set includes a second indication value, which indicates the maximum number of sublayers or the maximum number of sublayers minus the constant value based on information stored in the decoding parameter set. The first indication value is less than or equal to the second indication value.
[0043] Furthermore, according to an embodiment, a video encoder is provided for encoding video into a video data stream, such that the video data stream has video encoded therein. The video encoder is used to generate the video data stream, such that the video data stream includes a sequence parameter set and a decoding parameter set. The video encoder is used to generate the video data stream such that the sequence parameter set includes a first indication value, the first indication value indicating a maximum number of sublayers or the maximum number of sublayers minus a constant value based on information stored in the sequence parameter set. Furthermore, the video encoder generates the video data stream such that the decoding parameter set includes a second indication value, the second indication value indicating the maximum number of sublayers or the maximum number of sublayers minus the constant value based on information stored in the decoding parameter set. Furthermore, the video encoder is used to generate the video data stream such that the first indication value is less than or equal to the second indication value.
[0044] Furthermore, according to an embodiment, a video encoder is provided for encoding video into a video data stream, such that the video data stream has video encoded therein. The video encoder is used to generate the video data stream, such that the video data stream includes a sequence parameter set and a decoding parameter set. Furthermore, the video encoder is used to generate the video data stream such that the sequence parameter set includes a first indication value, the first indication value indicating a maximum number of sublayers or the maximum number of sublayers minus a constant value based on information stored in the sequence parameter set. Furthermore, the video encoder generates the video data stream such that the decoding parameter set includes a second indication value, the second indication value indicating the maximum number of sublayers or the maximum number of sublayers minus the constant value based on information stored in the decoding parameter set. The second indication value of the decoding parameter set is an upper boundary prior to the first indication value of the sequence parameter set.
[0045] Furthermore, according to an embodiment, an apparatus is provided for receiving a video data stream as an input bit stream, wherein the video data stream has video encoded therein. The video data stream includes a sequence parameter set and a decoding parameter set. The sequence parameter set includes a first indication value, which indicates a maximum number of sub-layers or the maximum number of sub-layers minus a constant value based on information stored in the sequence parameter set. The decoding parameter set includes a second indication value, which indicates the maximum number of sub-layers or the maximum number of sub-layers minus the constant value based on information stored in the decoding parameter set. If the first indication value is greater than the second indication value, the apparatus is used to process the input bit stream to generate an output bit stream, such that the output bit stream includes the second indication value as an indication of the maximum number of sub-layers or the maximum number of sub-layers minus the constant value.
[0046] Furthermore, according to an embodiment, a method for encoding video into a video data stream is provided, such that the video data stream has video encoded therein. The method includes generating the video data stream such that the video data stream includes a sequence parameter set and a decoding parameter set. Furthermore, the method includes generating the video data stream such that the sequence parameter set includes a first indication value, the first indication value indicating a maximum number of sublayers or the maximum number of sublayers minus a constant value based on information stored in the sequence parameter set. Furthermore, the method includes generating the video data stream such that the decoding parameter set includes a second indication value, the second indication value indicating the maximum number of sublayers or the maximum number of sublayers minus the constant value based on information stored in the decoding parameter set. The method includes generating the video data stream such that the first indication value is less than or equal to the second indication value.
[0047] Furthermore, according to an embodiment, a method for encoding video into a video data stream is provided, such that the video data stream has video encoded therein. The method includes generating the video data stream such that the video data stream includes a sequence parameter set and a decoding parameter set. Furthermore, the method includes generating the video data stream such that the sequence parameter set includes a first indication value, the first indication value indicating a maximum number of sublayers or the maximum number of sublayers minus a constant value based on information stored in the sequence parameter set. Furthermore, the method includes generating the video data stream such that the decoding parameter set includes a second indication value, the second indication value indicating the maximum number of sublayers or the maximum number of sublayers minus the constant value based on information stored in the decoding parameter set. The second indication value of the decoding parameter set is an upper boundary prior to the first indication value of the sequence parameter set.
[0048] Furthermore, according to an embodiment, a method is provided for receiving a video data stream as an input bit stream, wherein the video data stream has video encoded therein. The video data stream includes a sequence parameter set and a decoding parameter set. The sequence parameter set includes a first indication value, which indicates a maximum number of sub-layers or the maximum number of sub-layers minus a constant value based on information stored in the sequence parameter set. The decoding parameter set includes a second indication value, which indicates the maximum number of sub-layers or the maximum number of sub-layers minus the constant value based on information stored in the decoding parameter set. If the first indication value is greater than the second indication value, the method includes processing the input bit stream to generate an output bit stream such that the output bit stream includes the second indication value as an indication of the maximum number of sub-layers or the maximum number of sub-layers minus the constant value.
[0049] In addition, a computer program is provided for implementing the methods described above when executed on a computer or signal processor. Attached Figure Description
[0050] Figure 1 A video encoder for encoding video into a video data stream according to an embodiment is described.
[0051] Figure 2 An apparatus for receiving an input video data stream according to an embodiment is described.
[0052] Figure 3 A video decoder for receiving a video data stream in which video is stored, according to an embodiment, is described.
[0053] Figure 4 This illustrates an example of using reference image resampling for time-dependent scalability.
[0054] Figure 5 Two splicing point options are explained, where option I (top) maintains the frame rate throughout the example, and option II (bottom) has a frame rate drop in the last group of pictures in the bitstream.
[0055] Figure 6 This describes the video encoder.
[0056] Figure 7 This describes the video decoder.
[0057] Figure 8 This illustrates the relationship between a reconstructed signal (e.g., a reconstructed image) on one hand and a combination of a predicted residual signal and a predicted signal on the other hand, such as a signal sent in the data stream. Detailed Implementation
[0058] The following description of the accompanying drawings begins with a description of the encoder and decoder of a block-based predictive codec used to decode images of a video, to form an example of a decoding framework in which embodiments of the present invention can be built. About Figures 6 to 8 The corresponding encoders and decoders are described. Thereafter, embodiments of the concepts of the invention are described, along with descriptions of how these concepts can be respectively constructed to... Figure 6 and Figure 7 The descriptions in the encoder and decoder are presented together, although using Figures 1 to 3 The embodiments described below can also be used to form configurations not based on the following. Figure 6 and Figure 7 The basic decoding framework for encoders and decoders operates on the encoder and decoder.
[0059] Figure 6A video encoder is shown, an apparatus for predictively decoding image 12 into data stream 14 using transform-based residual decoding, as exemplarily. This apparatus or encoder is indicated by reference numeral 10. Figure 7 A corresponding video decoder 20 is shown, for example, an apparatus 20 configured to also use transform-based residual decoding to predictively decode the image 12' from the data stream 14, wherein the apostrophe has been used to indicate that the image 12' reconstructed by the decoder 20 deviates from the image 12 originally encoded by the encoder 10 in terms of the decoding loss introduced by the quantization of the predictive residual signal. Figure 6 and Figure 7 Transform-based predictive residual decoding is used exemplarily, although embodiments of this application are not limited to this type of predictive residual decoding. Regarding Figure 6 and Figure 7 The same applies to other details described, as will be outlined below.
[0060] Encoder 10 is configured to perform a spatial-to-spectral transformation on the prediction residual signal and encode the resulting prediction residual signal into data stream 14. Similarly, decoder 20 is configured to decode the prediction residual signal from data stream 14 and perform a spectral-to-spatial transformation on the resulting prediction residual signal.
[0061] Internally, encoder 10 may include a prediction residual signal former 22 that generates a prediction residual 24 to measure the deviation of the prediction signal 26 from the original signal, such as the deviation from image 12. The prediction residual signal former 22 may be, for example, a subtractor that subtracts the prediction signal from the original signal (e.g., from image 12). Encoder 10 then includes a transformer 28 that performs a space-to-spectral transformation on the prediction residual signal 24 to obtain a spectral domain prediction residual signal 24', which is then quantized by a quantizer 32, also included in encoder 10. The quantized prediction residual signal 24'' is thus decoded into bitstream 14. For this purpose, encoder 10 may optionally include an entropy decoder 34 that entropy decodes the transformed and quantized prediction residual signal into data stream 14. Prediction signal 26 is generated by prediction stage 36 of encoder 10 based on the prediction residual signal 24 encoded into and decodeable from data stream 14. For this purpose, as... Figure 6As shown, prediction stage 36 may internally include a dequantizer 38 that dequantizes the prediction residual signal 24'' to obtain a spectral domain prediction residual signal 24''' corresponding to signal 24' excluding quantization loss. This is followed by an inverse transform 40 that performs an inverse transform (e.g., spectral-to-spatial transform) on the latter prediction residual signal 24''' to obtain a prediction residual signal 24'''' corresponding to the original prediction residual signal 24 excluding quantization loss. Then, combiner 42 of prediction stage 36 recombines the prediction signal 26 and the prediction residual signal 24'''', such as by addition, to obtain a reconstructed signal 46, e.g., a reconstruction of the original signal 12. The reconstructed signal 46 may correspond to signal 12'. Prediction module 44 of prediction stage 36 then generates prediction signal 26 based on signal 46 using, for example, spatial prediction (e.g., intra-picture prediction) and / or temporal prediction (e.g., inter-picture prediction).
[0062] Similarly, as Figure 7 As shown, decoder 20 can be internally composed of components corresponding to prediction stage 36 and interconnected in a manner corresponding to prediction stage 36. Specifically, the entropy decoder 50 of decoder 20 can entropy decode the quantized spectral domain prediction residual signal 24'' from the data stream, so that the dequantizer 52, inverse transformer 54, combiner 56, and prediction module 58, interconnected and cooperating in the manner described above with respect to the modules of prediction stage 36, recover the reconstructed signal based on the prediction residual signal 24'', thereby achieving... Figure 7 As shown, the output of combiner 56 results in a reconstructed signal, i.e., picture 12'.
[0063] Although not specifically described above, it is readily apparent that encoder 10 can set decoding parameters, including prediction modes, motion parameters, etc., according to some optimization schemes (such as optimizing a certain rate and distortion correlation criterion (e.g., decoding cost)). For example, encoder 10 and decoder 20, along with corresponding modules 44 and 58, can each support different prediction modes, such as intra-frame decoding mode and inter-frame decoding mode. The encoder and decoder, by the granularity of their switching between these prediction mode types, can respectively correspond to subdividing images 12 and 12' into decoding segments or decoding blocks. For example, based on these decoding segments, the image can be subdivided into blocks to be intra-frame decoded and blocks to be inter-frame decoded. As outlined in more detail below, intra-frame decoded blocks are predicted based on the spatial, already decoded / decoded neighborhood of the corresponding blocks. Several intra-decoding modes may exist, and these modes are selected for corresponding intra-decoding segments that include intra-decoding modes of direction or angle. According to these modes, the corresponding segments are filled by extrapolating neighborhood sample values along a specific direction specific to the intra-decoding mode of that direction. The intra-decoding modes may, for example, also include one or more other modes, such as a DC decoding mode (according to which the prediction of the corresponding intra-decoding block assigns DC values to all samples within the corresponding intra-decoding segment), and / or a planar intra-decoding mode (according to which the prediction of the corresponding block is approximated or determined as a spatial distribution of sample values described by a two-dimensional linear function at the sample locations of the corresponding intra-decoding block, having a driving tilt and offset of a plane defined by the two-dimensional linear function based on neighboring samples). In contrast, inter-decoding blocks can be predicted, for example, temporally. For inter-frame decoding blocks, motion vectors can be signaled within the data stream. These motion vectors indicate the spatial displacement of a portion of a previously decoded image in the video to which image 12 belongs, where the previously decoded / decoded image is sampled to obtain the prediction signal for the corresponding inter-frame decoding block. This means that, in addition to the residual signal decoding included in data stream 14 (such as the entropy decoding transform coefficient levels representing the quantized spectral domain prediction residual signal 24''), data stream 14 can also encode the following: decoding mode parameters for assigning decoding modes to the individual blocks; prediction parameters for some of the blocks (such as motion parameters for inter-frame coding segments); and optional other parameters (such as parameters for controlling and signaling the subdivision of images 12 and 12' into segments, respectively). Decoder 20 uses these parameters to subdivide the images in the same manner as the encoder, assigning the same prediction modes to the segments and performing the same predictions to result in the same prediction signal.
[0064] Figure 8This illustrates the relationship between a reconstruction signal (e.g., reconstructed image 12') and a combination of a prediction residual signal 24'''' and a prediction signal 26, as signaled in data stream 14. As already stated above, this combination can be additive. Prediction signal 26... Figure 8 The image region is described as being subdivided into intra-frame decoding blocks illustratively indicated by shading and inter-frame decoding blocks illustratively indicated by non-shading. This subdivision can be any subdivision, such as regularly subdividing the image region into rows and columns of square or non-square blocks, or subdividing the image 12 from a root block multi-tree into multiple leaf blocks of different sizes, such as quadtree subdivision, etc., wherein... Figure 8 The text illustrates their mixing, where the image region is first subdivided into rows and columns of root blocks, and then further subdivided into one or more leaf blocks based on recursive multi-branch subdivision.
[0065] Similarly, for intra-frame decoded block 80, data stream 14 may have an intra-frame decode mode decoded therein, which assigns one of several supported intra-frame decode modes to the corresponding intra-frame decoded block 80. For inter-frame decoded block 82, data stream 14 may have one or more motion parameters encoded therein. Generally speaking, inter-frame decoded block 82 is not limited to being decoded in time. Alternatively, inter-frame decoded block 82 may be any block predicted from a previously decoded portion other than the current picture 12 itself, such as a previously decoded picture of the video to which picture 12 belongs, or another view or a picture at a lower layer in the hierarchy, where the encoder and decoder are scalable encoders and decoders, respectively.
[0066] Figure 8 The predicted residual signal 24'''' is also described as subdividing the image region into blocks 84. These blocks can be called transform blocks to distinguish them from decoded blocks 80 and 82. In fact, Figure 8 This describes that encoder 10 and decoder 20 can use two different subdivisions to divide images 12 and 12' into blocks respectively: one subdividing them into decoding blocks 80 and 82 respectively, and the other subdividing them into transform blocks 84 respectively. These two subdivisions can be the same; for example, each decoding block 80 and 82 can simultaneously form transform block 84, but... Figure 8This illustrates a situation where, for example, a subdivision into transform block 84 extends the subdivision into decoding blocks 80 and 82, such that any boundary between the two blocks of 80 and 82 covers the boundary between the two blocks 84, or alternatively, each block 80 and 82 either coincides with one of the transform blocks 84 or with a cluster of transform blocks 84. However, subdivisions can also be determined or selected independently so that transform block 84 can alternatively span the block boundary between blocks 80 and 82. Regarding the subdivision into transform block 84, similar statements are correct due to those previously made regarding the subdivision into blocks 80 and 82; for example, block 84 can be the result of regularly subdividing a picture region into blocks (with or without arrangement in rows and columns), the result of a recursive multi-branch subdivision of the picture region, or a combination thereof, or any other type of blocking. Incidentally, note that blocks 80, 82, and 84 are not limited to quadratic, rectangular, or any other shape.
[0067] Figure 8 This further illustrates that the combination of prediction signal 26 and prediction residual signal 24'''' directly results in the reconstructed image 12'. However, it should be noted that, according to alternative embodiments, more than one prediction signal 26 can be combined with prediction residual signal 24'''' to result in image 12'.
[0068] exist Figure 8 In this context, transform block 84 should have the following meaning. Transformer 28 and inverse transformer 54 perform their transforms on a unit basis of these transform blocks 84. For example, many codecs use some kind of DST or DCT for all transform blocks 84. Some codecs allow skipping transforms so that the residual signal can be directly decoded in the spatial domain for some of the transform blocks 84. However, according to the embodiment described below, encoder 10 and decoder 20 are configured to support several transforms. For example, the transforms supported by encoder 10 and decoder 20 may include: o DCT-II (or DCT-III), where DCT stands for Discrete Cosine Transform o DST-IV, where DST represents Discrete Sine Transform o DCT-IV o DST-VII o Identity Transformation (IT) Naturally, while transformer 28 will support all forward transform versions of these transforms, decoder 20 or inverse transformer 54 will support their corresponding backward or inverse versions: o Inverse DCT-II (or Inverse DCT-III) o Reverse DST-IV o Reverse DCT-IV o Inverse DST-VII o Identity Transformation (IT) The following description provides further details about which transforms the encoder 10 and decoder 20 are capable of supporting. In any case, it should be noted that the set of supported transforms may include only one transform, such as a spectrum-to-space or space-to-spectrum transform.
[0069] As outlined above, Figures 6 to 8 This has been presented as an example, in which the inventive concepts further described below can be implemented to form specific examples of encoders and decoders according to this application. So far, Figure 6 and Figure 7 The encoder and decoder can represent possible implementations of the encoder and decoder described below, respectively. However, Figure 6 and Figure 7 This is merely an example. However, the encoder according to embodiments of this application can perform block-based encoding of image 12 using concepts outlined in more detail below, and in conjunction with... Figure 6 The difference between this encoder and others lies in the fact that it is a still image encoder rather than a video encoder, that it does not support inter-frame prediction, or that it differs from others. Figure 8 The subdivision of block 80 is performed in the manner illustrated in the example. Similarly, the decoder according to embodiments of this application can perform block-based decoding of image 12' from data stream 14 using the decoding concepts further outlined below, but may, for example, be similar to... Figure 7 The difference between the decoder 20 and others is that it is not a video decoder but a still image decoder, it does not support intra-frame prediction, or it differs from the one about Figure 8 The described method is used to divide the image 12' into blocks and / or, for example, to derive the prediction residuals from the data stream 14 in the spatial domain rather than in the transform domain.
[0070] Figure 1 A video encoder 100 for encoding video into a video data stream according to an embodiment is described. The video encoder 100 is configured to generate a video data stream.
[0071] Figure 2 An apparatus 200 for receiving an input video data stream according to an embodiment is described. The input video data stream has video encoded therein. The apparatus 200 is configured to generate an output video data stream from the input video data stream.
[0072] Figure 3 A video decoder 300 according to an embodiment is described for receiving a video data stream in which video is stored. The video decoder 300 is configured to decode video from the video data stream.
[0073] Furthermore, a system according to an embodiment is provided. The system includes... Figure 2 The device and Figure 3 The video decoder. Figure 3 The video decoder (300) is configured to receive Figure 2 The output video data stream of the device (200). Figure 3 The video decoder 300 is configured to... Figure 2 The device 200 decodes the output video data stream.
[0074] In an embodiment, the system may, for example, further include Figure 1 The video encoder 100. For example, Figure 2 The device 200 can be configured to... Figure 1 The video encoder 100 receives video data streams as input video data streams.
[0075] The (optional) intermediate device 210 of apparatus 200 may be configured, for example, to receive a video data stream from video encoder 100 as an input video data stream and to generate an output video data stream from the input video data stream. For example, the intermediate device may be configured to modify the (header / metadata) information of the input video data stream and / or may be configured to delete pictures from the input video data stream and / or may be configured to mix / concatenate the input video data stream with an additional second bitstream having a second video encoded therein.
[0076] (Optional) The video decoder 221 can be configured, for example, to decode video from the output video data stream.
[0077] (Optional) It is assumed that the reference decoder 222 may be configured, for example, to determine the timing information of the video based on the output video data stream, or may be configured, for example, to determine the buffer information of the buffer in which the video or a portion of the video is to be stored.
[0078] The system includes Figure 1 Video encoder 101 and Figure 2 The video decoder 151.
[0079] The video encoder 101 is configured to generate an encoded video signal. The video decoder 151 is configured to decode the encoded video signal to reconstruct images of the video.
[0080] The following text describes trick mode and fast replay.
[0081] According to an embodiment, an apparatus is provided for receiving a video data stream as an input bit stream, wherein the video data stream has video encoded therein. For a playback speed modification factor (Nx), the apparatus is used to determine a decoding capability requirement based on decoding capability requirement constraints. The playback speed modification factor (Nx) is either a forward playback speed modification factor or a backward playback speed modification factor.
[0082] In an embodiment, the decoding capability requirement may be, for example, a modified sampling rate. The decoding capability requirement limitation may be, for example, a sampling rate limitation.
[0083] According to an embodiment, the video data stream may include, for example, information about sampling rate limits.
[0084] In one embodiment, the device may be configured, for example, to determine the modified sampling rate based on the image size of the video's pictures.
[0085] According to an embodiment, the device may, for example, be configured to also depend on the frame rate and a playback speed modification factor (Nx) to determine the modified sampling rate. The video data stream may, for example, include information about the frame rate.
[0086] In this embodiment, the frame rate depends on the sublayers of one or more sublayers. The video data stream may, for example, include sublayer-specific frame rate information for each of the one or more sublayers. The apparatus may, for example, be configured to obtain the frame rate from the sublayer-specific frame rate information for one of the one or more sublayers.
[0087] According to an embodiment, the apparatus may be configured, for example, to determine the modified sampling rate based on the image size of the video's pictures, which is the highest image size among a plurality of images in the video.
[0088] In an embodiment, the video data stream may include, for example, an application factor. The apparatus may be configured, for example, to use the highest-level image size and the application factor to calculate a modified sampling rate.
[0089] According to an embodiment, the maximum image size depends on one or more sub-layers. The video data stream may, for example, include sub-layer-specific maximum image size information for each of the one or more sub-layers. The apparatus may, for example, be configured to obtain the maximum image size from the sub-layer-specific maximum image size information for one of the one or more sub-layers.
[0090] In an embodiment, the video data stream may include, for example, an application factor. The apparatus may be configured, for example, to use the video's sampling rate and use the application factor to determine a modified sampling rate.
[0091] According to an embodiment, the decoding capability requirement may be, for example, a modified bitrate. The decoding capability requirement limitation may be, for example, a bitrate limitation.
[0092] According to the embodiments, the video data stream may include, for example, information about bitrate limits.
[0093] According to an embodiment, the apparatus may be configured, for example, to determine the modified bitrate based on the initial bitrate of the video data stream and based on the playback speed modification factor (Nx).
[0094] In an embodiment, the initial bitrate depends on the sub-layer video data stream in one or more sub-layers, including sub-layer-specific initial bitrate information that can be configured, for example, for each of the one or more sub-layers. The apparatus can be configured, for example, to obtain the initial bitrate from sub-layer-specific frame rate information for one of the one or more sub-layers.
[0095] According to an embodiment, the video data stream may, for example, include an IRAP-only flag (IRAP = Intra-Frame Random Access Point), which indicates whether decoding of IRAP-only images is possible as the initial playback speed increases. Depending on the IRAP-only flag, the apparatus may, for example, be configured to read decoding capability requirements from the video data stream.
[0096] In an embodiment, the video data stream may, for example, include a reference-only image flag, which indicates whether reference-only decoding is possible as the initial playback speed increases. Depending on the reference-only image flag, the apparatus may, for example, be configured to read decoding capability requirements from the video data stream.
[0097] Furthermore, according to an embodiment, a video data stream having video encoded therein is provided. The video data stream may, for example, include information about sampling rate limits, and / or the video data stream may, for example, include information about bitrate limits.
[0098] According to an embodiment, the video data stream may include, for example, information about the frame rate.
[0099] In an embodiment, the video data stream may include, for example, sublayer-specific frame rate information for each of one or more sublayers.
[0100] According to an embodiment, the video data stream may include, for example, an application factor for determining the modified sampling rate.
[0101] In an embodiment, the video data stream may, for example, include sub-layer-specific maximum image size information for each of one or more sub-layers.
[0102] According to an embodiment, the video data stream may, for example, include sub-layer-specific initial bitrate information for each of one or more sub-layers.
[0103] In an embodiment, the video data stream may include, for example, an IRAP-only flag indicating whether IRAP-only images can be decoded as the initial playback speed increases.
[0104] According to an embodiment, the video data stream may include, for example, a reference-only flag, whereby the IRAP-only flag indicates whether reference-only decoding is possible as the initial playback speed increases.
[0105] In embodiments, depending on the IRAP-only flag, the video data stream may include, for example, level information; and / or depending on the reference image-only flag, the video data stream may include, for example, level information.
[0106] Furthermore, according to embodiments, a video encoder is provided for encoding video into a video data stream, such that the video data stream has video encoded therein. The video encoder is used to generate the video data stream such that the video data stream includes information about sampling rate limits, and / or the video encoder is used to generate the video data stream such that the video data stream includes information about bitrate limits.
[0107] In an embodiment, the video encoder may be configured, for example, to generate a video data stream such that the video data stream may include, for example, information about frame rate limits.
[0108] According to an embodiment, a video encoder may be configured, for example, to generate a video data stream such that the video data stream may include, for example, sub-layer-specific frame rate information for each of one or more sub-layers.
[0109] In an embodiment, the video encoder may be configured, for example, to generate a video data stream such that the video data stream may include, for example, an application factor for determining a modified sampling rate.
[0110] According to an embodiment, a video encoder may be configured, for example, to generate a video data stream such that the video data stream may include, for example, sub-layer-specific maximum image size information for each of one or more sub-layers.
[0111] In an embodiment, the video encoder may be configured, for example, to generate a video data stream such that the video data stream may include, for example, sub-layer-specific initial bitrate information for each of one or more sub-layers.
[0112] According to an embodiment, the video encoder may be configured, for example, to generate a video data stream such that the video data stream may include, for example, an IRAP-only flag, which indicates whether IRAP-only image decoding is possible when the initial playback speed may be increased, for example.
[0113] In an embodiment, the video encoder may be configured, for example, to generate a video data stream such that the video data stream may include, for example, a reference-only flag indicating whether reference-only decoding is possible when the initial playback speed may be increased, for example.
[0114] According to an embodiment, a video encoder may be configured, for example, to generate a video data stream such that, depending on an IRAP-only flag, the video data stream may include, for example, level information; and / or a video encoder may be configured, for example, to generate a video data stream such that, depending on a reference image-only flag, the video data stream may include, for example, level information.
[0115] In an embodiment, the device may be configured, for example, to decode an input bitstream to decode video.
[0116] Furthermore, according to an embodiment, a system for encoding video into a video data stream and for decoding video is provided. The system includes a video encoder as described above and means as described above. The video encoder is used to encode video into a video data stream such that the video data stream has the video encoded therein. The means is used to receive the video data stream as an input bit stream. Furthermore, the means is used to decode the input bit stream to decode the video.
[0117] Fast-forwarding or reverse operation (rewinding and replaying quickly) is a typical operation performed in video applications. Typically, these operations involve decoding only a subset of the bitstream and decoding it at a higher speed (e.g., frame rate) than the speed indicated by the bitstream. Furthermore, even for non-stunt mode operations, playing video at a higher speed may be of interest in certain scenarios. While there isn't a huge difference between fast-forwarding and replaying, the distinction considered in this description could be that the second objective is continuous and smooth playback (even potentially with audio playback and synchronization), while the first objective is not; for example, fast-forwarding might be a discontinuous playback with many "jumps" in the content. In any case, the result is to replay the content at Nx speed.
[0118] Clearly, replaying content at a higher speed also requires decoding the bitstream at a higher speed, which leads to higher decoding capability requirements (i.e., higher levels), as a result of the higher sampling rate and higher bit rate indicated in the HRD parameters.
[0119] When reference image resampling (RPR) is not used, i.e., all images within the CVS or bitstream have the same size, the sampling rate can be easily calculated as the image size in the sample divided by the frame rate and multiplied by the speedup factor Nx. Similarly, the bitrate will be multiplied by the speedup factor Nx.
[0120] In the first embodiment, the two values described above as derived are checked according to the level limit in VVC, and the level to which the Nx acceleration belongs is calculated as the minimum value of the derived value that is less than the level limit.
[0121] Note that the frame rate or bitrate is sublayer-specific, which means that regardless of what sublayer is being replayed (e.g., fast-forwarding is performed only by decoding temporal level 0), the corresponding value is calculated using the sublayer-specific signaled frame rate or bitrate.
[0122] However, when using Reference Image Resampling (RPR), the sampling rate cannot be easily derived; that is, only the worst-case scenario with the highest image size can be used, which will result in deriving the worst-case sampling rate. Figure 4 An example illustrating how to use RPR with temporal scalability.
[0123] Figure 4 This illustrates an example of combining reference image resampling with temporal scalability.
[0124] In this example, the lowest sub-layer image size might be 1920x1080, and the highest sub-layer image size might be 960x540. By using only the maximum image size and assuming a 60fps bitstream, approximately 124*10 pixels per second will be derived. 6 The sampling rate for each sample, and the actual value will be approximately 78*10. 6 That is, in the worst case, it will be 1.6 times the actual value.
[0125] In one embodiment, the bitstream includes an indication of a factor that needs to be applied to the sampling rate derived from the worst-case value. Using the signaled factor, the worst-case sampling rate can be derived, divided by the signaled factor, and then multiplied by a speed decoding factor to calculate the actual sampling rate by accelerating the decoding operation.
[0126] In another embodiment, each sub-layer indicates a maximum image size value, which will lead to a better approximation or even a more accurate value if the image size does not change within the sub-layer.
[0127] Furthermore, there are some cases where sublayers are not fully utilized. Sometimes, fast-forwarding is performed by simply decoding and replaying IRAP images (e.g., IDR) or images that are not non-reference images. However, it is unclear what level is required to decode such a sub-bit stream when decoding at Nx speed. In another embodiment, additional level signaling is indicated for either of the following two cases: • IRAP only: Only IRAP is decoded.
[0128] • Reference image only: Non-reference images are discarded, and the rest are decoded.
[0129] The syntax for such an embodiment can be as follows.
[0130] For profile_tier_level: Similar to irap, this only allows adding reference images to the syntax.
[0131] The bit rate and CPB size can be similarly signaled as follows: The protection of delay_for_concatenation_ensured_flag is described below.
[0132] According to an embodiment, a video data stream having video encoded therein is provided. The video data stream includes a plurality of access units, wherein a group of access units in the video data stream includes a plurality of access units. Furthermore, the video data stream includes an indication for each access unit of a subset of access units, indicating the existence of a seamless stitching point after the access unit. The subset of access units is an appropriate subset of the group of access units in the video data stream.
[0133] In an embodiment, the video data stream may, for example, include multiple temporal sublayers. The indication that a seamless stitching point exists after an access unit exists only for access units of the first set of temporal sublayers among the multiple temporal sublayers, but does not exist for access units of the second set of temporal sublayers among the multiple temporal sublayers.
[0134] According to an embodiment, the video data stream may, for example, include buffer period supplementation enhancement information. The buffer period supplementation enhancement information indicates a subset of access units among a plurality of access units for which a seamless stitching point exists after the access unit.
[0135] In an embodiment, the video data stream may include, for example, a divisor. A delay value may be assigned, for example, to each of the plurality of access units in the video data stream. An indication that a seamless stitching point exists exists only for those access units in the plurality of access units of the video data stream that have the assigned delay value, which may, for example, be equal to the remainder when divided by the divisor.
[0136] According to an embodiment, the remainder value can be, for example, 0.
[0137] In an embodiment, the video data stream may include, for example, a remainder value.
[0138] According to an embodiment, an indication of a seamless splicing point following each access unit in a subset of access units can be, for example, a flag.
[0139] In an embodiment, this flag could be, for example, delay_for_concatenation_ensured_flag.
[0140] According to an embodiment, if, for example, a constant frame rate for the output image obtained from the decoded video data stream cannot be guaranteed, the flag exhibits a first value. If, for example, a constant frame rate for the output image obtained from the decoded video data stream can be guaranteed, the flag may, for example, exhibit a second value different from the first value.
[0141] In an embodiment, if the output images obtained from the decoded video data stream have equidistant output times, the flag exhibits a second value.
[0142] Furthermore, according to embodiments, a video encoder is provided for encoding video into a video data stream, such that the video data stream has video encoded therein. The video encoder is used to generate a video data stream such that the video data stream includes a plurality of access units, wherein a group of access units in the video data stream includes a plurality of access units. Additionally, the video encoder is used to generate a video data stream such that the video data stream includes an indication for each access unit of a subset of access units, indicating the existence of a seamless stitching point after the access unit. Furthermore, the video encoder is used to generate a video data stream such that the subset of access units is an appropriate subset of the group of access units in the video data stream.
[0143] In an embodiment, the video encoder may be configured, for example, to generate a video data stream such that the video data stream may include, for example, information about a plurality of temporal sublayers. The video encoder may be configured, for example, to generate a video data stream such that an indication of a seamless stitching point after an access unit exists only for access units of a first set of temporal sublayers among the plurality of temporal sublayers; however, the video encoder may be configured, for example, to generate a video data stream such that an indication of a seamless stitching point after an access unit does not exist for access units of a second set of temporal sublayers among the plurality of temporal sublayers.
[0144] According to an embodiment, the video encoder may be configured, for example, to generate a video data stream such that the video data stream may include, for example, buffer periodicity enhancement information. The video encoder may be configured, for example, to generate a video data stream such that the buffer periodicity enhancement information indicates a subset of access units among a plurality of access units for which a seamless stitching point exists after the access unit.
[0145] In an embodiment, the video encoder may be configured, for example, to generate a video data stream such that the video data stream may include, for example, a divisor. The video encoder may be configured, for example, to generate a video data stream such that a delay value may be assigned, for example, to each of a plurality of access units in the video data stream. The video encoder may be configured, for example, to generate a video data stream such that an indication that a seamless stitching point exists exists only for those access units in the plurality of access units of the video data stream that have an assigned delay value, the delay value being, for example, equal to a remainder when divided by the divisor.
[0146] According to an embodiment, the video encoder may be configured, for example, to generate a video data stream such that the remainder value may be, for example, 0.
[0147] In an embodiment, the video encoder may be configured, for example, to generate a video data stream such that the video data stream may include, for example, a remainder value.
[0148] According to an embodiment, the video encoder may be configured, for example, to generate a video data stream such that an indication of a seamless stitching point following each access unit of a subset of access units may be, for example, a flag.
[0149] In an embodiment, the video encoder may be configured, for example, to generate a video data stream such that an indication of a seamless stitching point following each access unit of a subset of access units may be, for example, a flag.
[0150] According to an embodiment, if, for example, a constant frame rate for the output image obtained from the decoded video data stream cannot be guaranteed, the video encoder generates a video data stream such that the flag exhibits a first value. If, for example, a constant frame rate for the output image obtained from the decoded video data stream can be guaranteed, the video encoder can be configured, for example, to generate a video data stream such that the flag exhibits a second value different from the first value.
[0151] In an embodiment, if the output images obtained from the decoded video data stream have equidistant output times, the video encoder may, for example, be configured to generate a video data stream such that the flag exhibits a second value.
[0152] Furthermore, according to an embodiment, an apparatus is provided for receiving a video data stream as an input bit stream, wherein the video data stream has video encoded therein. The apparatus is used to process the video data stream. The video data stream includes a plurality of access units, wherein a group of access units of the video data stream includes a plurality of access units. Additionally, the video data stream includes an indication for each access unit of a subset of access units, indicating the existence of a seamless stitching point after the access unit. The subset of access units is an appropriate subset of the group of access units of the video data stream. The apparatus is used to process the indication.
[0153] In an embodiment, the video data stream may, for example, include multiple temporal sublayers. The indication that a seamless stitching point exists after an access unit exists only for access units of the first set of temporal sublayers among the multiple temporal sublayers, but does not exist for access units of the second set of temporal sublayers among the multiple temporal sublayers.
[0154] According to an embodiment, the video data stream may, for example, include buffer periodicity enhancement information. The buffer periodicity enhancement information indicates a subset of access units among a plurality of access units for which a seamless stitching point exists after the access unit. The apparatus may, for example, be configured to process the buffer periodicity enhancement information.
[0155] In an embodiment, the video data stream may, for example, include a divisor. The apparatus may, for example, be configured to receive the divisor. A delay value may, for example, be assigned to each of a plurality of access units in the video data stream. An indication that a seamless stitching point exists in an access unit exists only for those access units in the plurality of access units of the video data stream that have the assigned delay value, which may, for example, be equal to the remainder when divided by the divisor.
[0156] According to an embodiment, the remainder value can be, for example, 0.
[0157] In an embodiment, the video data stream may include, for example, a remainder value.
[0158] The device can be configured, for example, to receive the remainder value.
[0159] According to an embodiment, an indication of a seamless splice point following each access unit in the subset of access units can be, for example, a flag. The apparatus can be configured, for example, to process the flag.
[0160] In an embodiment, this flag could be, for example, delay_for_concatenation_ensured_flag.
[0161] According to an embodiment, if, for example, a constant frame rate for the output image obtained from the decoded video data stream cannot be guaranteed, the flag exhibits a first value. If, for example, a constant frame rate for the output image obtained from the decoded video data stream can be guaranteed, the flag exhibits a second value different from the first value.
[0162] In an embodiment, if the output images obtained from the decoded video data stream have equidistant output times, the flag exhibits a second value.
[0163] According to an embodiment, the device may be configured, for example, to decode an input bitstream to decode video.
[0164] Furthermore, according to an embodiment, a system for encoding video into a video data stream and for decoding video is provided. The system includes a video encoder as described above and an apparatus as described above. The video encoder is used to encode video into a video data stream such that the video data stream has the video encoded therein. The apparatus is used to receive the video data stream as an input bit stream. Furthermore, the apparatus is used to decode the input bit stream to decode the video.
[0165] The current VVC draft specification includes instructions to allow concatenating devices, specifically devices that concatenate portions of bitstreams A and B to find AUs in bitstream A that allow seamless concatenation of bitstream B. Seamless concatenation means that the CPB removal time distance between the last AU of bitstream A and the first AU of bitstream B is the same as the CPB removal time of two consecutive AUs of bitstream A. Conversely, non-seamless concatenation would mean that the decoding (CPB removal) of the first AU of bitstream B would be delayed because the first AU of bitstream B cannot be sent to the decoder's CPB at the correct time (i.e., the last AU of bitstream A takes longer to send than expected), preventing the AU from being decoded as quickly as expected. The Current Picture Timing SEI message in the VVC draft specification includes a flag called `delay_for_concatenation_ensured_flag`. When set to 1, this flag indicates that for a particular associated AU, splicing at that bitstream position will be seamless, provided that the following AUs (i.e., the beginning of bitstream B) have a BP SEI message with `delay_for_concatenation_ensured_flag` equal to 1, and the selected `InitCpbRemovalDelay` is less than or equal to the value of `max_initial_removal_delay_for_concatination` indicated in bitstream A. Therefore, for all these AUs in bitstream A, `delay_for_concatenation_ensured_flag` represents a variable option for seamless splicing.
[0166] The problem is that not all access units with the aforementioned characteristics will represent meaningful concatenation points, and even for meaningless concatenation points, `delay_for_concatenation_ensured_flag` will be signaled. This could lead to poorly implemented concatenation devices concatenating at such suboptimal locations. Furthermore, the indication of `delay_for_concatenation_ensured_flag` for all access units, even those for meaningless concatenation points, introduces bitrate and processing overhead on the encoder side to correctly set the flag value, even when the flag value is useless to the concatenation device. In this context, meaningless concatenation points mean that even if the frame rate is maintained at the actual concatenation point, some of these concatenation points may omit a portion of bitstream A, resulting in discontinuous playback of bitstream A before the concatenation point, for example, concatenating in the middle of a hierarchical GOP. Figure 5 An example is shown where two bitstreams (A and B) are spliced at two locations (I and II). The first option (top) results in a constant frame rate throughout the example, while the second option (bottom) causes a decrease in frame rate due to the selection of an incorrect splicing point.
[0167] Figure 5 Two splicing point options are explained, where option I (top) maintains the frame rate throughout the example, while option II (bottom) has a frame rate drop in the last GOP of bitstream A.
[0168] Therefore, the present invention reduces signaling overhead and avoids misleading the splicing device by selectively signaling delay_for_concatenation_ensured_flag for a subset of access units.
[0169] In one embodiment, delay_for_concatenation_ensured_flag is signaled only for the AU of a specific time-series inter-series sublayer, where the specific time-series inter-series sublayer is indicated in the relevant BP SEI message.
[0170] In another embodiment, delay_for_concatenation_ensured_flag is signaled only at a subset of all AUs during the buffer period, where the subset of AUs is identified by information from the associated BP SEI message.
[0171] In one embodiment, a subset of AUs is indicated by signaling the divisor and an optional remainder value. The delay_for_concatenation_ensured_flag in the PT SEI message for the buffer period is signaled only if the value of the associated AU (e.g., cpb_removal_delay or dpb_output_delay in the PT SEI message) is divided by the divisor and equals the optional remainder, or equals zero in the absence of a divisor.
[0172] In another embodiment, if the output image cannot guarantee a constant frame rate as shown, the value of delay_for_concatenation_ensured_ is set to 0. Note that the current specification currently only focuses on decoding time and ensures that the difference between the last arrival time (when the AU has fully arrived at the CPB) and the CPB removal delay (i.e., > threshold) of AUs with the flag set to 1 satisfies this condition. Therefore, equidistant decoding times for consecutive AUSs are achieved. Here, the invention will supplement that if this flag is set to 1, the output times of consecutive AUSs in bitstream A also have equidistant output times.
[0173] The following describes the HRD specific to the output layer set.
[0174] According to an embodiment, an apparatus is provided for receiving a video data stream as an input bit stream, wherein the video data stream has video encoded therein. The apparatus is configured to process the input bit stream in a manner dependent on an output layer set to obtain sub-bits. If the output layer set does not include a predefined layer (vps_layer_id[0]) among a plurality of layers present in the video data stream, the apparatus is configured to remove at least one non-expandable nested supplemental enhancement information message allocated with the predefined layer (vps_layer_id[0]). If the output layer set includes the predefined layer (vps_layer_id[0]), the apparatus is configured not to remove any non-expandable nested supplemental enhancement information messages allocated with the predefined layer (vps_layer_id[0]).
[0175] If there is no output layer set consisting of all the layers in the current decoded video sequence of the video data stream, then for example, there may be no non-expandable nested supplementary enhancement information message having at least one of the buffer period payload, the picture timing payload, and the decoding unit information payload.
[0176] According to an embodiment, if there is no output layer set consisting of all the layers in the current decoded video sequence of the video data stream, then for example, there may be no non-expandable nested supplementary enhancement information message having at least one of the buffer period payload, the picture timing payload, the decoding unit information payload, and the sub-picture level information payload.
[0177] In an embodiment, if there is no output layer set consisting of all layers in the current decoded video sequence of the video data stream, then for example, there may be no non-expandable nested supplemental enhancement information message whose payload type is equal to a first value indicating the buffer period payload, or equal to a second value indicating the picture timing payload, or equal to a third value indicating the decoding unit information payload, or equal to a fourth value indicating the sub-picture level information payload.
[0178] According to an embodiment, for a non-expandable nested supplemental enhancement information message whose payload type is equal to the first value, or equal to the second value, or equal to the third value, or equal to the fourth value, when present, the non-expandable nested supplemental enhancement information message is applied to all output layer sets consisting of all layers in the current decoded video sequence in the entire bitstream.
[0179] In an embodiment, if there is no output layer set consisting of all layers in the current decoded video sequence of the video data stream, then for example, there may be no non-expandable nested supplemental enhancement information message whose payload type is equal to 0 indicating the buffer period payload, or equal to 1 indicating the picture timing payload, or equal to 130 indicating the decoding unit payload, or equal to 203 indicating the sub-picture level information payload.
[0180] According to an embodiment, for a non-expandable nested supplemental enhancement information message with a payload type equal to 0, 1, 130, or 203, when present, the non-expandable nested supplemental enhancement information message is applied to all output layer sets consisting of all layers in the current decoded video sequence throughout the entire bitstream.
[0181] In an embodiment, if the output layer set does not include the predefined layer (vps_layer_id[0]), the device may be configured, for example, to remove all nested supplementary enhancement information messages allocated together with the predefined layer (vps_layer_id[0]).
[0182] Furthermore, according to an embodiment, a video data stream having video encoded therein is provided. The video data stream includes a set of video parameters. The video parameter set includes multiple sets of output layers. If none of the multiple sets of output layers includes all layers present in the multiple layers of the video data stream, then the video data stream does not include any non-scalable nested supplementary enhancement information messages used for assuming a reference decoder.
[0183] If there is no output layer set consisting of all the layers in the current decoded video sequence of the video data stream, then for example, there may be no non-expandable nested supplementary enhancement information message having at least one of the buffer period payload, the picture timing payload, and the decoding unit information payload.
[0184] According to an embodiment, if there is no output layer set consisting of all the layers in the current decoded video sequence of the video data stream, then for example, there may be no non-expandable nested supplementary enhancement information message having at least one of the buffer period payload, the picture timing payload, the decoding unit information payload, and the sub-picture level information payload.
[0185] In an embodiment, if there is no output layer set consisting of all layers in the current decoded video sequence of the video data stream, then for example, there may be no non-expandable nested supplemental enhancement information message whose payload type is equal to a first value indicating the buffer period payload, or equal to a second value indicating the picture timing payload, or equal to a third value indicating the decoding unit information payload, or equal to a fourth value indicating the sub-picture level information payload.
[0186] According to an embodiment, for a non-expandable nested supplemental enhancement information message whose payload type is equal to the first value, or equal to the second value, or equal to the third value, or equal to the fourth value, when present, the non-expandable nested supplemental enhancement information message is applied to all output layer sets consisting of all layers in the current decoded video sequence in the entire bitstream.
[0187] In an embodiment, if there is no output layer set consisting of all layers in the current decoded video sequence of the video data stream, then for example, there may be no non-expandable nested supplemental enhancement information message whose payload type is equal to 0 indicating the buffer period payload, or equal to 1 indicating the picture timing payload, or equal to 130 indicating the decoding unit payload, or equal to 203 indicating the sub-picture level information payload.
[0188] According to an embodiment, for a non-expandable nested supplemental enhancement information message with a payload type equal to 0, 1, 130, or 203, when present, the non-expandable nested supplemental enhancement information message is applied to all output layer sets consisting of all layers in the current decoded video sequence throughout the entire bitstream.
[0189] In an embodiment, if the video data stream includes at least one non-expandable nested supplemental enhancement information message for the hypothetical reference decoder, then at least one of the plurality of output layer sets may, for example, include all of the plurality of layers present in the video data stream.
[0190] According to an embodiment, all supplemental enhancement information messages used for the hypothetical reference decoder are scalable nested supplemental enhancement information messages, which are independent of the output layer sets of the plurality of output layer sets, which may include, for example, all layers present in the plurality of layers in the video data stream.
[0191] In an embodiment, all non-expandable nested supplemental enhancement information messages for the hypothetical reference decoder are associated with an output layer set in the plurality of output layer sets, which may include, for example, all of the plurality of layers present in the video data stream.
[0192] Furthermore, according to an embodiment, a video encoder is provided for encoding video into a video data stream, such that the video data stream has video encoded therein. The video encoder is used to generate the video data stream such that the video data stream includes a set of video parameters. Furthermore, the video encoder is used to generate the video data stream such that the video parameter set includes a plurality of output layer sets. If none of the plurality of output layer sets includes all layers present in the plurality of layers in the video data stream, then the video encoder is used to generate the video data stream such that the video data stream does not include any non-expandable nested supplementary enhancement information messages used for assuming a reference decoder.
[0193] If, for example, there is no output layer set consisting of all the layers in the plurality of layers in the current decoded video sequence of the video data stream, then for example, there is no non-expandable nested supplementary enhancement information message having at least one of the buffer period payload, the picture timing payload, and the decoding unit information payload.
[0194] According to an embodiment, if, for example, there is no output layer set consisting of all the layers in the plurality of layers in the current decoded video sequence of the video data stream, then, for example, there is no non-expandable nested supplementary enhancement information message having at least one of the buffer period payload, the picture timing payload, the decoding unit information payload, and the sub-picture level information payload.
[0195] In an embodiment, if there is no output layer set consisting of all layers in the current decoded video sequence of the video data stream, then for example, there may be no non-expandable nested supplemental enhancement information message whose payload type is equal to a first value indicating the buffer period payload, or equal to a second value indicating the picture timing payload, or equal to a third value indicating the decoding unit information payload, or equal to a fourth value indicating the sub-picture level information payload.
[0196] According to an embodiment, for a non-expandable nested supplemental enhancement information message whose payload type is equal to the first value, or equal to the second value, or equal to the third value, or equal to the fourth value, when present, the non-expandable nested supplemental enhancement information message is applied to all output layer sets consisting of all layers in the current decoded video sequence in the entire bitstream.
[0197] In an embodiment, if there is no output layer set consisting of all layers in the current decoded video sequence of the video data stream, then there is no non-expandable nested supplemental enhancement information message whose payload type is equal to 0 indicating the buffer period payload, or equal to 1 indicating the picture timing payload, or equal to 130 indicating the decoding unit payload, or equal to 203 indicating the sub-picture level information payload.
[0198] According to an embodiment, for a non-expandable nested supplemental enhancement information message with a payload type equal to 0, 1, 130, or 203, when present, the non-expandable nested supplemental enhancement information message is applied to all output layer sets consisting of all layers in the current decoded video sequence throughout the entire bitstream.
[0199] In an embodiment, if the video data stream may include, for example, at least one non-expandable nested supplemental enhancement information message for assuming a reference decoder, the video encoder may be configured, for example, to generate the video data stream such that at least one of the plurality of output layer sets includes all of the plurality of layers present in the video data stream.
[0200] According to an embodiment, the video encoder may be configured, for example, to generate the video data stream such that all supplemental enhancement information messages used for the hypothetical reference decoder are scalable nested supplemental enhancement information messages, all of which are independent of the output layer sets of the plurality of output layer sets, which may include, for example, all of the plurality of layers present in the video data stream.
[0201] In an embodiment, the video encoder may be configured, for example, to generate the video data stream such that all non-expandable nested supplemental enhancement information messages for the hypothetical reference decoder are associated with the output layer sets of the plurality of output layer sets, which may include, for example, all layers present in the plurality of layers in the video data stream.
[0202] According to an embodiment, the apparatus may be configured, for example, to decode the input bitstream to decode the video.
[0203] The current VVC draft specification provides a means to carry HRD timing information (BP SEI message, PT SEI message, DUI SEI message) in a nested form within a bitstream that is only applicable to sub-bitstreams (i.e., OLSs). Other HRD SEI messages carried directly in the bitstream (as non-expandable nested SEI messages) need to be applied to the 0th OLS, i.e., the single-layer sub-bitstream indicated in the VPS by the syntax element vps_layer_id[0]. However, as is typical of layered coding, removing this layer or any other layer from the bitstream does not require rewriting the vp, so it is not even possible for a layer with nuh_layer_id equal to vps_layer_id[0] to exist.
[0204] The problem with this requirement is that when extracting OLS other than the 0th OLS, no non-expandable nested SEI messages are allowed, even if the extracted OLS does include a layer whose nuh_layer_id is equal to vps_layer_id[0]. Such non-expandable nested messages can only be preserved in the bitstream after rewriting the VPS(vps_layer_id[0]) and the corresponding values in the NAL cell headers of each layer.
[0205] The present invention, which addresses this problem, is described below in two aspects.
[0206] 1) As the first solution, the complexity described above is not addressed, but the extraction process is adjusted such that when the extracted OLS does not include layers with nuh_layer_id equal to vps_layer_id[0], the extraction process only removes non-scalable nested SEI messages applied to layers with nuh_layer_id equal to vps_layer_id[0]. Therefore, it is not necessary to unnecessarily delete non-scalable nested SEI messages (or in other words, it is not necessary to unnecessarily copy every OLS that includes layers with nuh_layer_id equal to vps_layer_id[0]).
[0207] 2) As a second solution, when no OLS corresponding to the complete bitstream (i.e., all layers within it) is defined in the VPS, non-expandable nested SEI messages (BP, PT, DUI) are not allowed in the bitstream. Otherwise, when an OLS corresponding to all existing layers in the bitstream exists, individual HRD SEI messages are allowed in a non-expandable nested manner within the bitstream. All other HRD SEI messages (applied to other OLSs) must be within expandable nested SEI messages. Conversely, all non-expandable nested SEI messages correspond as a whole to the current bitstream.
[0208] For non-expandable nested supplemental enhancement messages, when payloadType equals 0 (indicating buffer period, BP / content), 1 (indicating picture timing, PT, content), 130 (indicating decoding unit information, DUI, content), or 203 (indicating subpicture level information, SLI, content), non-expandable nested supplemental enhancement messages are applied to all output layer sets, which, when present, consist of all layers in the current encoded video sequence throughout the bitstream.
[0209] When there is no output layer set consisting of all layers in the current encoded video sequence of the entire bitstream, there will be no non-expandable nested supplemental enhancement messages with payloadType equal to 0 (BP), 1 (PT), 130 (DUI), or 203 (SLI).
[0210] The following section describes the rewrite sublayer PTL in DPS.
[0211] According to an embodiment, an apparatus is provided for receiving a video data stream as an input bitstream, wherein the video data stream has video encoded therein. The video data stream includes initial profile hierarchy information. Furthermore, the video data stream includes a first set of two or more layers. The apparatus is configured to process the input bitstream based on an output layer set of multiple output layer sets to generate an output bitstream by removing at least one layer from the first set of two or more layers to obtain a second set of one or more layers, such that the output bitstream includes the second set of one or more layers but excludes the at least one layer removed from the first set. Furthermore, the apparatus is configured to remove portions of the initial profile hierarchy information that are unrelated to at least one layer in the second set of one or more layers.
[0212] According to an embodiment, the video data stream may, for example, include initial decoding parameter set profile hierarchy information as the initial profile hierarchy information. The apparatus may, for example, be configured to remove portions of the initial decoding parameter set profile hierarchy information that are irrelevant to at least one layer in the second group of one or more layers.
[0213] Furthermore, according to an embodiment, an apparatus is provided for receiving a video data stream as an input bit stream, wherein the video data stream has video encoded therein.
[0214] The video data stream includes expandable nested supplemental enhancement information, which includes decoding parameter set information for each of a plurality of output layer sets. The apparatus is configured to process the input bitstream based on the output layer sets of the plurality of output layer sets to obtain an output bitstream, such that the output bitstream includes the decoding parameter set information for the output layer sets, and such that the output bitstream does not include the decoding parameter set information for any other output layer set among the plurality of output layer sets.
[0215] In an embodiment, the video data stream may, for example, include a raw byte sequence payload. The raw byte sequence payload may, for example, include the decoding parameter set information for each of the plurality of output layer sets.
[0216] According to an embodiment, the video data stream may, for example, include scalable nested supplemental enhancement messages, which may include, for example, the decoding parameter set information for each of the plurality of output layer sets.
[0217] In an embodiment, the video data stream may include, for example, a nesting type flag, which may consist of one or more bits.
[0218] According to an embodiment, if the nesting type flag (nesting_type) exhibits a first value, this indicates that the Scalable Nested Supplemental Enhancement Message may, for example, include one or more sets of decoded parameters, the one or more sets of decoded parameters including the decoded parameter set information, and indicates that no other supplemental enhancement messages are nested within the Scalable Nested Supplemental Enhancement Message.
[0219] In an embodiment, if the nesting type flag (nesting_type) exhibits a second value, this indicates that the one or more supplementary enhancement messages nested within the expandable nested supplementary enhancement messages are applied to the output layer set.
[0220] According to an embodiment, if the nesting type flag (nesting_type) exhibits a third value, this indicates that the one or more supplementary enhancement messages nested within the expandable nested supplementary enhancement messages are applied to a subset of the output layer set.
[0221] In an embodiment, if the nesting type flag (nesting_type) exhibits a fourth value, this indicates that the Scalable Nest Supplemental Enhancement Message may, for example, include one or more sets of decoding parameters, which include the decoding parameter set information.
[0222] Furthermore, according to an embodiment, a video data stream having video encoded therein is provided. The video data stream may, for example, include scalable nested supplemental enhancement information, which includes decoding parameter set information for each of a plurality of output layer sets.
[0223] In an embodiment, the video data stream may, for example, include a raw byte sequence payload. The raw byte sequence payload may, for example, include the decoding parameter set information for each of the plurality of output layer sets.
[0224] According to an embodiment, the video data stream may, for example, include an expandable nested supplemental enhancement message, which may include, for example, the expandable nested supplemental enhancement information, which may include, for example, the decoding parameter set information for each of the plurality of output layer sets.
[0225] In an embodiment, the video data stream may include, for example, a nesting type flag, which may consist of one or more bits.
[0226] According to an embodiment, if the nesting type flag (nesting_type) exhibits a first value, this indicates that the Scalable Nested Supplemental Enhancement Message may, for example, include one or more sets of decoded parameters, the one or more sets of decoded parameters including the decoded parameter set information, and indicates that no other supplemental enhancement messages are nested within the Scalable Nested Supplemental Enhancement Message.
[0227] In an embodiment, if the nesting type flag (nesting_type) exhibits a second value, this indicates that the one or more supplementary enhancement messages nested within the expandable nested supplementary enhancement messages are applied to the output layer set.
[0228] According to an embodiment, if the nesting type flag (nesting_type) exhibits a third value, this indicates that one or more supplementary enhancement messages nested in the expandable nested supplementary enhancement message are applied to a subset of the output layer set, for example, where one or more sub-images are assigned to the subset.
[0229] In an embodiment, if the nesting type flag (nesting_type) exhibits a fourth value, this indicates that the Scalable Nest Supplemental Enhancement Message may, for example, include one or more sets of decoding parameters, which include the decoding parameter set information.
[0230] Furthermore, according to an embodiment, a video encoder is provided for encoding video into a video data stream, such that the video data stream has video encoded therein. The video encoder is used to generate the video data stream such that the video data stream includes scalable nested supplementary enhancement information, the scalable nested supplementary enhancement information including decoding parameter set information for each of a plurality of output layer sets.
[0231] In an embodiment, the video encoder may be configured, for example, to generate the video data stream such that the video data stream may include, for example, the original byte sequence payload. The video encoder may also be configured, for example, to generate the video data stream such that the original byte sequence payload may include, for example, the decoding parameter set information for each of the plurality of output layer sets.
[0232] According to an embodiment, a video encoder may be configured, for example, to generate the video data stream such that the video data stream may include, for example, an expandable nested supplemental enhancement message, the expandable nested supplemental enhancement message may include, for example, the expandable nested supplemental enhancement information, the expandable nested supplemental enhancement information may include, for example, the decoding parameter set information for each of the plurality of output layer sets.
[0233] In an embodiment, the video encoder may be configured, for example, to generate a video data stream such that the video data stream includes a nesting type flag, which may include, for example, one or more bits.
[0234] According to an embodiment, the video encoder may be configured, for example, to generate the video data stream such that if the nesting type flag (nesting_type) exhibits a first value, this indicates that the Scalable Nested Supplemental Enhancement Message includes one or more sets of decoding parameters, which may include, for example, decoding parameter set information, and indicates that no other supplemental enhancement messages are nested within the Scalable Nested Supplemental Enhancement Message.
[0235] In an embodiment, the video encoder may be configured, for example, to generate the video data stream such that if the nesting type flag (nesting_type) exhibits a second value, this indicates that one or more supplementary enhancement messages nested within the expandable nested supplementary enhancement messages are applied to the output layer set.
[0236] According to an embodiment, the video encoder may be configured, for example, to generate the video data stream such that if the nesting type flag (nesting_type) exhibits a third value, this indicates one or more supplementary enhancement messages nested within the expandable nested supplementary enhancement messages to a subset of the output layer set, for example, wherein one or more sub-pictures are assigned to the subset.
[0237] In an embodiment, the video encoder may, for example, generate the video data stream such that if the nesting type flag (nesting_type) exhibits a fourth value, this indicates that the Scalable Nested Supplement Enhancement Message may, for example, include one or more sets of decoding parameters, the one or more sets of decoding parameters including the decoding parameter set information.
[0238] Furthermore, according to an embodiment, a video data stream having video encoded therein is provided. The video data stream includes indications of a plurality of operation points. Additionally, the video data stream includes a first mapping for each of the plurality of operation points, the first mapping assigning one or more profile levels from a plurality of profile levels to the operation point. Furthermore, the video data stream includes a plurality of second mappings, wherein each of the plurality of mappings assigns one of the plurality of operation points to a set of output levels from a plurality of output level sets.
[0239] In an embodiment, the video data stream may, for example, include a set of decoding parameters. The set of decoding parameters may, for example, include indications of multiple operation points.
[0240] According to an embodiment, the video data stream may, for example, include a number of layer indicators indicating the number of layers in the video data stream; and / or, the video data stream may, for example, include a number of sub-layer indicators indicating the number of sub-layers in the video data stream.
[0241] In an embodiment, the number of layers in the video data stream may be, for example, constant; and / or the number of sub-layers in the video data stream may be, for example, constant.
[0242] According to an embodiment, the video data stream may, for example, include scalable nested supplemental enhancement information, which includes decoding parameter set information for each of a plurality of output layer sets.
[0243] Furthermore, according to an embodiment, a video encoder is provided for encoding video into a video data stream, such that the video data stream has video encoded therein. The video encoder is used to generate the video data stream such that the video data stream includes indications of a plurality of operation points. Furthermore, the video encoder is used to generate the video data stream such that the video data stream includes a first mapping for each of the plurality of operation points, the first mapping assigning one or more profile levels from a plurality of profile levels to the operation point. Furthermore, the video encoder is used to generate the video data stream such that the video data stream includes a plurality of second mappings, wherein each of the plurality of mappings assigns one of the plurality of operation points to an output layer set from a plurality of output layer sets.
[0244] In an embodiment, the video encoder may be configured, for example, to generate the video data stream such that the video data stream may include, for example, a set of decoding parameters. The video encoder may be configured, for example, to generate the video data stream such that the set of decoding parameters may include, for example, the indications of the plurality of operation points.
[0245] According to an embodiment, the video encoder may be configured, for example, to generate the video data stream such that the video data stream may include, for example, a number of layer indicators indicating the number of layers in the video data stream; and / or the video encoder may be configured, for example, to generate the video data stream such that the video data stream may include, for example, a number of sublayer indicators indicating the number of sublayers in the video data stream.
[0246] In an embodiment, the video encoder may be configured, for example, to generate the video data stream such that the number of layers in the video data stream may be, for example, constant; and / or the video encoder may be configured, for example, to generate the video data stream such that the number of sub-layers in the video data stream may be, for example, constant.
[0247] According to an embodiment, the video encoder may be configured, for example, to generate the video data stream such that the video data stream may include, for example, scalable nested supplemental enhancement information, which includes decoding parameter set information for each of a plurality of output layer sets.
[0248] Furthermore, according to an embodiment, an apparatus is provided for receiving a video data stream as an input bit stream, wherein the video data stream has video encoded therein. The video data stream includes indications of a plurality of operation points. Furthermore, the video data stream includes a first mapping for each of the plurality of operation points, the first mapping mapping one or more profile levels from a plurality of profile levels to the operation point. Furthermore, the video data stream includes a plurality of second mappings, wherein each of the plurality of mappings maps one of the plurality of operation points to an output level from a plurality of output level sets. The apparatus is used to process the input bit stream based on the output level set from the plurality of output level sets and based on at least one of the second mappings mapping one of the plurality of operation points to the output level set to produce an output bit stream.
[0249] In an embodiment, the video data stream may, for example, include a set of decoding parameters. The set of decoding parameters may, for example, include indications of multiple operation points. The apparatus is used to process the set of decoding parameters.
[0250] According to an embodiment, the video data stream may, for example, include a number of layer indicators indicating the number of layers in the video data stream; and / or, the video data stream may, for example, include a number of sub-layer indicators indicating the number of sub-layers in the video data stream.
[0251] In an embodiment, the number of layers in the video data stream may be, for example, constant; and / or the number of sub-layers in the video data stream may be, for example, constant.
[0252] According to an embodiment, the video data stream may, for example, include scalable nested supplemental enhancement information, which includes decoding parameter set information for each of a plurality of output layer sets. The apparatus may, for example, be configured to process the scalable nested supplemental enhancement information.
[0253] According to an embodiment, the apparatus may be configured, for example, to decode the input bitstream to decode the video.
[0254] Furthermore, according to an embodiment, a system is provided for encoding video into a video data stream and for decoding the video. The system includes a video encoder as described above and a decoding device as described above. The video encoder is used to encode the video into the video data stream such that the video data stream has the video encoded therein. The device is used to receive the video data stream as an input bit stream. Furthermore, the device is used to decode the input bit stream to decode the video.
[0255] The VVC draft specification includes a Decoding Parameter Set (DPS) that describes the constraints (PTLs) of the bitstream. It can be used for capability exchange and negotiation of SDPs in applications such as RTSP / SIP, or for stream selection in adaptive streaming scenarios. Therefore, its purpose is to indicate the maximum capability required for a location stream (a concatenation of CVSs), and thus its scope is larger than all parameter sets defined in the corresponding previous generation codec generation (HEVC and AVC) which only had the scope of coded video sequences (CVSs). Now, the VVC draft specification describes the constraints of CVSS concatenation, i.e., the bitstream. The DPS primarily carries multiple profile hierarchy information (PTLs) including all PTLs used in the bitstream; that is, all PTLs in this list need to be supported in order to successfully decode the corresponding bitstream.
[0256] When performing bitstream extraction or pruning, some layers may be removed. After such processing, the resulting bitstream may no longer carry the entire set of layers; for example, extracting 3 layers containing the version (corresponding OLS) from a 6-layer bitstream. Therefore, if the DPS cannot accurately describe the PTL of the pruned bitstream, the DPS will no longer serve its capability negotiation purpose.
[0257] Therefore, as part of this invention, three options for mitigating this problem and allowing capability negotiation using DPS are described.
[0258] The DPS rewriting process during extraction is described below.
[0259] The PTL information in the DPS is adjusted during extraction because only the PTL information corresponding to the remaining layers after extraction is retained in the DPS, while other PTL information is removed.
[0260] The following text describes the DPS nesting specific to the sub-bit stream.
[0261] An extensible nested SEI message is defined that carries the OLS-specific DPS RBSP and the extracted data. The corresponding DPS RBSP is extracted through the bit stream extraction process and written into a new DPS in the output bit stream.
[0262] In this embodiment, a new nested SEI message carrying a set of parameters is used.
[0263] In another embodiment, a single expandable nested SEI message is used to carry a parameter set or other SEI messages. In this example, syntax elements are added to the expandable nested SEI message to indicate what is being executed internally.
[0264] In the syntax above, it is assumed that when the parameter set (nesting_type == 1) is included, there are no other SEIs. In another embodiment, the parameter set and other SEIs can be included in the nested SEI message simultaneously.
[0265] In another embodiment, nesting_type can be used for several purposes. Instructions: • SEI messages nested within nested SEIs are applied to OLS • SEI messages nested within nested SEIs are applied to subsets of the OLS, such as sub-images within the OLS. • A parameter set exists within a nested SEI. The following section introduces the action points in DPS.
[0266] DPS rewriting during the extraction process and sub-bitstream-specific DPS nesting may work in certain situations (e.g., constrained bitstreams). For example, DPS rewriting during the extraction process may only work if all VPSs of the bitstream are known beforehand and the number of layers / sublayers does not change from CVS to CVS. A similar situation occurs in the case of sub-bitstream-specific DPS nesting, because in principle, OLS (e.g., their number and ID) may change from CVS to CVS.
[0267] In another embodiment, operation points are defined. These may be defined in the DPS, providing a list for each PTL. In addition to the OLS index, operation point indices are added as input to the decoding process to describe and select appropriate operation points. The associations between the operation points and the OLS are then signaled in the VPS.
[0268] The syntax elements of DPS will change as follows by removing the number of sub-layers, because this may change from CVS to CVS and include operation points.
[0269] In another embodiment, num_sub_layer and the number of layers remain constant in the bitstream. This indicator is added to the bitstream and indicates the mapping between the operation point and the sub-layer or multiple layers.
[0270] In another embodiment, the extensible nested SEI message defined in 4.2 is applied to the operation point and also provides a mapping to the OLS in the bit stream.
[0271] The MaxSubLayer in the Param set is described below.
[0272] According to an embodiment, a video data stream having video encoded therein is provided. The video data stream includes a sequence parameter set and a decoding parameter set. The sequence parameter set includes a first indication value, which indicates a maximum number of sublayers or the maximum number of sublayers minus a constant value based on information stored in the sequence parameter set. The decoding parameter set includes a second indication value, which indicates the maximum number of sublayers or the maximum number of sublayers minus the constant value based on information stored in the decoding parameter set. The first indication value is less than or equal to the second indication value.
[0273] In an embodiment, the video data stream may, for example, include a sequence parameter set and a decoding parameter set. The sequence parameter set may, for example, include a first indication value that indicates a maximum number of sublayers or the maximum number of sublayers minus a constant value, based on information stored in the sequence parameter set. The decoding parameter set may, for example, include a second indication value that indicates the maximum number of sublayers or the maximum number of sublayers minus the constant value, based on information stored in the decoding parameter set. The second indication value of the decoding parameter set is an upper boundary prior to the first indication value of the sequence parameter set.
[0274] According to an embodiment, the constant value may be, for example, 1.
[0275] Furthermore, according to an embodiment, a video encoder is provided for encoding video into a video data stream, such that the video data stream has video encoded therein. The video encoder is used to generate the video data stream, such that the video data stream includes a sequence parameter set and a decoding parameter set. The video encoder is used to generate the video data stream such that the sequence parameter set includes a first indication value, the first indication value indicating a maximum number of sublayers or the maximum number of sublayers minus a constant value based on information stored in the sequence parameter set. Furthermore, the video encoder generates the video data stream such that the decoding parameter set includes a second indication value, the second indication value indicating the maximum number of sublayers or the maximum number of sublayers minus the constant value based on information stored in the decoding parameter set. Furthermore, the video encoder is used to generate the video data stream such that the first indication value is less than or equal to the second indication value.
[0276] Furthermore, according to an embodiment, a video encoder is provided for encoding video into a video data stream, such that the video data stream has video encoded therein. The video encoder is used to generate the video data stream, such that the video data stream includes a sequence parameter set and a decoding parameter set. Furthermore, the video encoder is used to generate the video data stream such that the sequence parameter set includes a first indication value, the first indication value indicating a maximum number of sublayers or the maximum number of sublayers minus a constant value based on information stored in the sequence parameter set. Furthermore, the video encoder generates the video data stream such that the decoding parameter set includes a second indication value, the second indication value indicating the maximum number of sublayers or the maximum number of sublayers minus the constant value based on information stored in the decoding parameter set. The second indication value of the decoding parameter set is an upper boundary prior to the first indication value of the sequence parameter set.
[0277] In an embodiment, the constant value may, for example, be 1. The video encoder may be configured, for example, to generate the video data stream such that the sequence parameter set may, for example, include the first indication value, which indicates the maximum number of sublayers or the maximum number of sublayers minus 1 based on information stored in the sequence parameter set. The video encoder may, for example, be configured to generate the video data stream such that the decoding parameter set may, for example, include the second indication value, which indicates the maximum number of sublayers or the maximum number of sublayers minus 1 based on information stored in the decoding parameter set.
[0278] Furthermore, according to an embodiment, an apparatus is provided for receiving a video data stream as an input bit stream, wherein the video data stream has video encoded therein. The video data stream includes a sequence parameter set and a decoding parameter set. The sequence parameter set includes a first indication value, which indicates a maximum number of sub-layers or the maximum number of sub-layers minus a constant value based on information stored in the sequence parameter set. The decoding parameter set includes a second indication value, which indicates the maximum number of sub-layers or the maximum number of sub-layers minus the constant value based on information stored in the decoding parameter set. If the first indication value is greater than the second indication value, the apparatus is used to process the input bit stream to generate an output bit stream, such that the output bit stream includes the second indication value as an indication of the maximum number of sub-layers or the maximum number of sub-layers minus the constant value.
[0279] In an embodiment, the constant value may, for example, be 1. The sequence parameter set may, for example, include the first indication value, which indicates the maximum number of sublayers or the maximum number of sublayers minus 1 based on information stored in the sequence parameter set. The decoding parameter set may, for example, include a second indication value, which indicates the maximum number of sublayers or the maximum number of sublayers minus 1 based on information stored in the decoding parameter set. If the first indication value may, for example, be greater than the second indication value, the apparatus may, for example, be configured to process the input bitstream to generate an output bitstream such that the output bitstream may, for example, include the second indication value as the maximum number of sublayers minus 1.
[0280] When performing bitstream pruning (also known as sublayer extraction) and discarding NAL cells belonging to a specific temporal ID from the bitstream, the parameter set is usually not modified, although some high-level parameter sets, such as DPS, can be rewritten.
[0281] This leads to inconsistencies within the bitstream, making the decoding process unclear or more complex. For example, if the DPS is rewritten as described in the previous section, and the actual number of sublayers is indicated as dps_max_sublayers_minus1 (excluding discarded temporal IDs), but the value at the SPS is not modified (which is not expected), sps_max_sublayers_minus1 will have a higher value than the value at the DPS. The decoder will not really know which value to trust.
[0282] In one embodiment, the bitstream constraint is that `sps_max_sublayers_minus1` must be less than or equal to `dps_max_sublayers_minus1`. Therefore, some type of error resilience can be built if such a "problem" is encountered. With this constraint, it's clear that when parsing a "bitstream" that violates this condition, a higher value of `SPS` will simply correspond to the new bitstream (e.g., one that wasn't detected due to a missing end-of-bitstream (EOB) NAL cell). This means that when performing bitstream pruning, `dps_max_sublayers_minus1` cannot be changed unless `sps_max_sublayers_minus1` is also changed.
[0283] In another embodiment, this constraint is not required, and DPS signaling is used to set the operation point or output layer set of the bitstream sent to the decoder. That is, when dps_max_sublayers_minus1 is less than sps_max_sublayers_minus1, the value signaled in the DPS takes precedence over the value indicated in the SPS that the bitstream sent to the decoder has at most dps_max_sublayers_minus1.
[0284] Since the number of sublayers may differ between layers when there are more than one layer, dps_max_sublayers_minus1 applies to all output layers, while all non-output layers have dps_max_sublayers_minus1 or its corresponding sps_max_sublayers_minus1 (if the latter is less than dps_max_sublayers_minus1).
[0285] Although some aspects have been described in the context of the apparatus, it is clear that these aspects also represent a description of the corresponding method, where a block or device corresponds to a method step or a feature of a method step. Similarly, aspects described in the context of method steps also represent a description of a corresponding block or item or feature of the corresponding apparatus. Some or all of the method steps can be performed by (or using) hardware devices such as, for example, microprocessors, programmable computers, or electronic circuits. In some embodiments, one or more of the most important method steps can be performed by such devices.
[0286] Depending on certain implementation requirements, embodiments of the present invention can be implemented in hardware, software, or at least partially in hardware or at least partially in software. Implementations can be carried out using digital storage media (e.g., floppy disks, DVDs, Blu-ray discs, CDs, ROMs, PROMs, EPROMs, EEPROMs, or flash memory) having electronically readable control signals stored thereon that cooperate with (or are capable of cooperating with) a programmable computer system to perform the corresponding methods. Therefore, the digital storage medium can be computer-readable.
[0287] Some embodiments of the invention include a data carrier having electronically readable control signals that are capable of cooperating with a programmable computer system to perform one of the methods described herein.
[0288] Generally, embodiments of the present invention can be implemented as a computer program product having program code that is operational so that, when the computer program product is run on a computer, one of the methods is performed. The program code may, for example, be stored on a machine-readable medium.
[0289] Other embodiments include a computer program stored on a machine-readable medium for performing one of the methods described herein.
[0290] In other words, therefore, an embodiment of the inventive method is a computer program having program code for performing one of the methods described herein when the computer program is run on a computer.
[0291] Therefore, another embodiment of the inventive method is a data carrier (or digital storage medium or computer-readable medium) comprising a computer program recorded thereon for performing one of the methods described herein. The data carrier, digital storage medium, or recorded medium is typically tangible and / or non-transitory.
[0292] Therefore, another embodiment of the inventive method is a data stream or signal sequence representing a computer program for performing one of the methods described herein. This data stream or signal sequence can, for example, be configured to be transferred via a data communication connection (e.g., via the Internet).
[0293] Other embodiments include processing components, such as computers or programmable logic devices, configured or adapted to perform one of the methods described herein.
[0294] Another embodiment includes a computer having a computer program installed thereon for performing one of the methods described herein.
[0295] Further embodiments of the invention include an apparatus or system configured to transfer (e.g., electronically or optically) a computer program for carrying out one of the methods described herein to a receiver. The receiver may be, for example, a computer, a mobile device, a memory device, etc. The apparatus or system may, for example, include a file server for transferring the computer program to the receiver.
[0296] In some embodiments, a programmable logic device (e.g., a field-programmable gate array) may be used to implement some or all of the functionality of the methods described herein. In some embodiments, the field-programmable gate array may cooperate with a microprocessor to implement one of the methods described herein. Generally, the method is preferably implemented by any hardware device.
[0297] The device described in this article can be implemented using hardware devices, a computer, or a combination of hardware devices and a computer.
[0298] The methods described in this article can be implemented using hardware devices, computers, or a combination of hardware devices and computers.
[0299] The above embodiments are merely illustrative of the principles of the invention. It is understood that modifications and variations of the arrangements and details described herein will be readily apparent to those skilled in the art. Therefore, the intent of the invention is limited only by the scope of the forthcoming patent claims, and not by the specific details presented in the manner of describing and interpreting the embodiments herein.
[0300] References [1] ISO / IEC, ITU-T. High-efficiency video decoding. ITU-T recommends H.265 | ISO / IEC 23008 10 (HEVC), 1st edition, 2013; 2nd edition, 2014.
Claims
1. A method for receiving a video data stream as an input bit stream, wherein, Video data streams contain video encoded into them. Specifically, for the playback speed modification factor (Nx), the method includes determining the decoding capability requirement based on decoding capability requirement constraints. The playback speed modification factor (Nx) is either a forward playback speed modification factor or a backward playback speed modification factor.
2. A method for encoding video into a video data stream, such that the video data stream has the video encoded therein. in, The method includes generating the video data stream such that the video data stream includes information about sampling rate limits, and / or The method includes generating the video data stream such that the video data stream includes information about bitrate limits.
3. A video data stream having video encoded therein, in, The video data stream includes multiple access units, wherein the access unit group of the video data stream includes the multiple access units. The video data stream includes an indication for each access unit in the subset of access units, indicating the existence of a seamless stitching point after the access unit. The access unit subset is an appropriate subset of the access unit group of the video data stream.
4. A video encoder for encoding video into a video data stream, such that the video data stream has the video encoded therein. in, The video encoder is used to generate the video data stream, such that the video data stream includes multiple access units, wherein the access unit group of the video data stream includes the multiple access units. The video encoder is used to generate the video data stream, such that the video data stream includes an indication for each access unit of the subset of access units, indicating the existence of a seamless stitching point after the access unit. The video encoder is used to generate the video data stream such that the subset of access units is an appropriate subset of the group of access units in the video data stream.
5. An apparatus for receiving a video data stream as an input bit stream, wherein, The video data stream contains video encoded within it. The device is used to process the video data stream. The video data stream includes multiple access units, and the access unit group of the video data stream includes the multiple access units. The video data stream includes an indication for each access unit in the subset of access units, indicating the existence of a seamless stitching point after the access unit. Wherein, the subset of access units is an appropriate subset of the group of access units in the video data stream. The device is used to process the instruction.
6. A method for encoding video into a video data stream, such that the video data stream has the video encoded therein. in, The method includes generating the video data stream such that the video data stream includes a plurality of access units, wherein the access unit group of the video data stream includes the plurality of access units. The method includes generating the video data stream such that the video data stream includes an indication for each access unit of the access unit subset, indicating the existence of a seamless stitching point after the access unit. The method includes generating the video data stream such that the subset of access units is an appropriate subset of the group of access units in the video data stream.
7. A method for receiving a video data stream as an input bit stream, wherein, The video data stream contains video encoded within it. The method includes processing the video data stream. The video data stream includes multiple access units, and the access unit group of the video data stream includes the multiple access units. The video data stream includes an indication for each access unit in the subset of access units, indicating the existence of a seamless stitching point after the access unit. Wherein, the subset of access units is an appropriate subset of the group of access units in the video data stream. The method includes processing the instruction.
8. An apparatus for receiving a video data stream as an input bit stream, wherein, Video data streams contain video encoded into them. The device is used to process the input bitstream in relation to an output layer set to obtain a sub-bitstream. Wherein, if the output layer set does not include a predefined layer (vps_layer_id[0]) among the multiple layers existing in the video data stream, the device is used to remove at least one non-expandable nested supplementary enhancement information message allocated together with the predefined layer (vps_layer_id[0]), and If the output layer set includes the predefined layer (vps_layer_id[0]), the device is configured not to remove any non-expandable nested supplementary enhancement information messages allocated together with the predefined layer (vps_layer_id[0]).
9. A video data stream having video encoded therein, in, The video data stream includes a set of video parameters. The video parameter set includes multiple output layer sets. If none of the plurality of output layer sets includes all layers present in the plurality of layers in the video data stream, then the video data stream does not include any non-scalable nested supplementary enhancement information messages used for assuming a reference decoder.
10. A video encoder for encoding video into a video data stream, such that the video data stream has the video encoded therein. in, The video encoder is used to generate the video data stream, such that the video data stream includes a set of video parameters. The video encoder is used to generate the video data stream, such that the video parameter set includes multiple output layer sets. If none of the plurality of output layer sets includes all layers present in the plurality of layers in the video data stream, then the video encoder generates the video data stream such that the video data stream does not include any non-expandable nested supplementary enhancement information messages used for assuming a reference decoder.