Video data stream, video encoder, apparatus and method for hypothetical reference decoder and for output layer set

By introducing the playback speed modification factor Nx and code rate/sampling rate limit in video encoding and decoding, a video data stream containing seamless splicing points and scalable nestable supplementary enhancement information is solved, and the shortcomings of parallel processing and data stream management in the prior art are achieved, and more efficient video encoding and decoding capabilities are achieved.

CN114787921BActive Publication Date: 2025-08-29FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202080088769.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-12-20
Filing Date
2020-12-18
Publication Date
2025-08-29
Estimated Expiration
2040-12-18

AI Technical Summary

Technical Problem

The existing video encoding and decoding technologies have shortcomings in parallel processing capabilities and the management of video data streams, especially in the H.265/HEVC standard, the CTU processing method limits the flexibility of more efficient parallel processing and video data streams.

Method used

By introducing playback speed modification factor Nx and code rate/sampling rate limiting information, a video data stream is generated, and seamless splicing points and extensible nestable supplementary enhancement information are included in the data stream, the processing methods of video encoder and decoder are optimized, and multi-level operation points and operation mappings are supported to achieve flexible management of video data streams.

Benefits of technology

It improves the parallel processing capabilities of video encoding and decoding, enhances the flexibility and management efficiency of video data streams, supports multi-level operation and seamless splicing, and meets different decoding capabilities requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114787921B_ABST
    Figure CN114787921B_ABST
Patent Text Reader

Abstract

A video encoder is provided, which is configured to: encode multiple access units into a video data stream, the multiple access units including: a first access unit and a second access unit; and encode a delay_for_concatenation_ensured_flag corresponding to the first access unit into the video data stream, and in response to the value of delay_for_concatenation_ensured_flag being 1, determine that the first access unit is located at a position in the video data stream, which makes the splicing of the first access unit with the second access unit having the following characteristics seamless: concatenation_flag is equal to 1, and the selected InitCpbRemovalDelay is less than or equal to the value of max_initial_removal_delay_for_concatination.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to video encoding and video decoding, and in particular to a video encoder, to a video decoder, to methods for encoding and decoding, and to a video data stream for implementing advanced video coding concepts. Background Art

[0002] H.265 / HEVC (HEVC = High Efficiency Video Coding) is a video codec that already provides tools for improving or even enabling parallel processing at the encoder and / or decoder. For example, HEVC supports the subdivision of a picture into an array of tiles that are coded independently of each other. Another concept supported by HEVC involves WPP, according to which CTU rows or CTU lines of a picture can be processed in parallel (e.g., in the form of slices) from left to right, as long as a certain minimum CTU offset (CTU = Coding Tree Unit) is observed in the processing of consecutive CTU rows. However, it would be advantageous to have a video codec that has the ability to support parallel processing of video encoders and / or video decoders even more efficiently.

[0003] Typically, in video coding, the decoding process of picture samples requires smaller partitions, where samples are divided into rectangular areas for joint processing, such as prediction or transform coding. Therefore, pictures are divided into blocks of a specific size that is constant during the encoding of a video sequence. In the H.264 / AVC standard, fixed-size blocks of 16x16 samples, so-called macroblocks, are used (AVC = Advanced Video Coding).

[0004] In the state-of-the-art HEVC standard (see [1]), there is a maximum size of Coding Tree Block (CTB) or Coding Tree Unit (CTU) of 64×64 samples. In the further description of HEVC, the more general term CTU is used for this type of block.

[0005] CTUs are processed in raster scan order, starting with the top-left CTU and processing the CTUs in the image row by row, working down to the bottom-right CTU.

[0006] Decoded CTU data is organized into containers called slices. Originally, in previous video coding standards, a slice referred to a segment consisting of one or more consecutive CTUs of a picture. Slices were adopted for segmenting decoded data. From another perspective, a complete picture can also be defined as a large segment, and therefore, historically, the term slice still applies. In addition to the decoded picture samples, a slice also includes additional information related to the decoding process of the slice itself, which is placed in the so-called slice header.

[0007] According to the current state of the art, the VCL (Video Coding Layer) also includes techniques for fragmentation and spatial segmentation. For example, such segmentation can be applied in video coding for various reasons, such as processing load balancing in parallelization, CTU size matching in network transmission, and error mitigation.

[0008] As specified in the video coding standards, the bitstream includes information associated with HRD conformance. This conformance consists of a hypothetical reference decoder (HRD) that includes a buffer model that assumes NAL units enter the coded picture buffer (CPB) before the decoder and are removed from it at specific timings to ensure that the CPB size is not exceeded (buffer overrun) or that NAL units arrive no later than they need to be removed (buffer underrun). Furthermore, the model includes a decoded picture buffer (DPB), from which decoded pictures are output when they are no longer needed for prediction. In many implementations, the size of decoded pictures is also constrained. Timing information for the HRD is conveyed in the bitstream via so-called SEI messages. Specifically, the buffering period (BP) SEI message defines the specific timing information for the buffering period (a number of access units, or AUs); the picture timing (PT) SEI message conveys timing information for a single associated AU; and the decoding unit information (DUI) SEI message conveys timing information for an associated subset of AUs (i.e., decoding units, or DUs). Summary of the Invention

[0009] It is an object of the present invention to provide improved concepts for video encoding and video decoding.

[0010] The purpose of the present invention is solved by the embodiments disclosed in this application.

[0011] This application also provides preferred embodiments.

[0012] According to an embodiment, an apparatus is provided for receiving a video data stream as an input bitstream, wherein the video data stream has video encoded therein. The apparatus is configured to determine a decoding capability requirement based on a decoding capability requirement constraint for a playback speed modification factor (Nx). The playback speed modification factor (Nx) is either a forward playback speed modification factor or a backward playback speed modification factor.

[0013] Furthermore, according to an embodiment, a video data stream having video encoded therein is provided, wherein the video data stream includes information about sampling rate limitations, and / or the video data stream includes information about bit rate limitations.

[0014] Furthermore, according to an embodiment, a video encoder for encoding a video into a video data stream is provided, such that the video data stream has the video encoded therein. The video encoder is configured to generate the video data stream such that the video data stream includes information about a sampling rate limit, and / or the video encoder is configured to generate the video data stream such that the video data stream includes information about a bit rate limit.

[0015] Furthermore, according to an embodiment, a method is provided for receiving a video data stream as an input bitstream, wherein the video data stream has video encoded therein. For a playback speed modification factor (Nx), the method includes determining a decoding capability requirement based on a decoding capability requirement constraint. The playback speed modification factor (Nx) is either a forward playback speed modification factor or a backward playback speed modification factor.

[0016] Furthermore, according to an embodiment, a method for encoding a video into a video data stream is provided, such that the video data stream has the video encoded therein. The method includes generating the video data stream such that the video data stream includes information about a sampling rate limit, and / or the method includes generating the video data stream such that the video data stream includes information about a bit rate limit.

[0017] Furthermore, a computer program is provided for implementing one of the methods as described above when executed on a computer or a signal processor.

[0018] Furthermore, according to an embodiment, a video data stream having video encoded therein is provided. The video data stream comprises a plurality of access units, wherein an access unit group of the video data stream comprises the plurality of access units. Furthermore, the video data stream comprises, for each access unit of a subset of access units, an indication indicating that a seamless splice point exists after the access unit. The subset of access units is a proper subset of the access unit group of the video data stream.

[0019] Furthermore, according to an embodiment, a video encoder for encoding a video into a video data stream is provided, such that the video data stream has the video encoded therein. The video encoder is configured to generate the video data stream such that the video data stream comprises a plurality of access units, wherein an access unit group of the video data stream comprises the plurality of access units, and furthermore, the video encoder is configured to generate the video data stream such that the video data stream comprises, for each access unit of a subset of access units, an indication indicating that a seamless splice point exists after the access unit. Furthermore, the video encoder is configured to generate the video data stream such that the subset of access units is a proper subset of the access unit group of the video data stream.

[0020] Furthermore, according to an embodiment, an apparatus is provided for receiving a video data stream as an input bitstream, wherein the video data stream has video encoded therein. The apparatus is configured to process the video data stream. The video data stream comprises a plurality of access units, wherein an access unit group of the video data stream comprises the plurality of access units. Furthermore, the video data stream comprises, for each access unit of a subset of access units, an indication indicating that a seamless splice point exists after the access unit. The subset of access units is a suitable subset of the access unit group of the video data stream. The apparatus is configured to process the indication.

[0021] Furthermore, according to an embodiment, a method for encoding a video into a video data stream is provided, such that the video data stream has the video encoded therein. The method comprises generating the video data stream such that the video data stream comprises a plurality of access units, wherein an access unit group of the video data stream comprises the plurality of access units. Furthermore, the method comprises generating the video data stream such that the video data stream comprises, for each access unit of a subset of access units, an indication indicating that a seamless splice point exists after the access unit, furthermore, the method comprises generating the video data stream such that the subset of access units is a proper subset of the access unit group of the video data stream.

[0022] Furthermore, according to an embodiment, a method for receiving a video data stream as an input bitstream is provided, wherein the video data stream has video encoded therein. The method comprises processing the video data stream, the video data stream comprising a plurality of access units, wherein an access unit group of the video data stream comprises the plurality of access units. Furthermore, the video data stream comprises, for each access unit of a subset of access units, an indication indicating that a seamless splice point exists after the access unit. The subset of access units is a proper subset of the access unit group of the video data stream. The method comprises processing the indication.

[0023] Furthermore, a computer program is provided for implementing one of the methods as described above when executed on a computer or a signal processor.

[0024] Furthermore, according to an embodiment, an apparatus is provided for receiving a video data stream as an input bitstream, wherein the video data stream has video encoded therein. The apparatus is configured to process the input bitstream to obtain sub-bits depending on an output layer set. If the output layer set does not include a predefined layer among a plurality of layers present in the video data stream, the apparatus is configured to remove at least one non-scalable nested supplemental enhancement information message allocated with the predefined layer. If the output layer set includes the predefined layer, the apparatus is configured not to remove any non-scalable nested supplemental enhancement information message allocated with the predefined layer.

[0025] Furthermore, according to an embodiment, a video data stream having video encoded therein is provided. The video data stream includes a video parameter set. The video parameter set includes a plurality of output layer sets. If none of the plurality of output layer sets includes all of a plurality of layers present in the video data stream, then the video data stream does not include any non-scalable nested supplemental enhancement information message for a hypothetical reference decoder.

[0026] Furthermore, according to an embodiment, a video encoder is provided for encoding video into a video data stream, such that the video data stream has the video encoded therein. The video encoder is configured to generate the video data stream such that the video data stream includes a video parameter set. Furthermore, the video encoder is configured to generate the video data stream such that the video parameter set includes a plurality of output layer sets. If none of the plurality of output layer sets includes all of a plurality of layers present in the video data stream, the video encoder is configured to generate the video data stream such that the video data stream does not include any non-scalable nested supplemental enhancement information messages for a hypothetical reference decoder.

[0027] Furthermore, according to an embodiment, a method for receiving a video data stream as an input bitstream is provided, wherein the video data stream has video encoded therein. The method comprises processing the input bitstream to obtain a sub-bitstream in dependence on an output layer set. If the output layer set does not include a predefined layer (vps_layer_id[0]) among a plurality of layers present in the video data stream, the method comprises removing at least one non-scalable nested supplemental enhancement information message assigned with the predefined layer (vps_layer_id[0]). If the output layer set includes the predefined layer (vps_layer_id[0]), the method comprises not removing any non-scalable nested supplemental enhancement information message assigned with the predefined layer (vps_layer_id[0]).

[0028] Furthermore, according to an embodiment, a method for encoding video into a video data stream is provided, such that the video data stream has the video encoded therein. The method includes generating the video data stream such that the video data stream includes a video parameter set. Furthermore, the method includes generating the video data stream such that the video parameter set includes a plurality of output layer sets. If none of the plurality of output layer sets includes all of a plurality of layers present in the video data stream, the method includes generating the video data stream such that the video data stream does not include any non-scalable nested supplemental enhancement information messages for a hypothetical reference decoder.

[0029] Furthermore, according to an embodiment, an apparatus is provided for receiving a video data stream as an input bitstream, wherein the video data stream has video encoded therein. The video data stream comprises initial profile level information. Furthermore, the video data stream comprises a first set of two or more layers. The apparatus is configured to process the input bitstream in dependence on an output layer set of a plurality of output layer sets to generate an output bitstream by removing at least one layer from the first set of two or more layers to obtain a second set of one or more layers, such that the output bitstream comprises the second set of one or more layers without the at least one layer removed from the first set. Furthermore, the apparatus is configured to remove from the initial profile level information those parts of the initial profile level information that are not relevant to at least one layer of the second set of one or more layers.

[0030] Furthermore, according to an embodiment, an apparatus is provided for receiving a video data stream as an input bitstream, wherein the video data stream has video encoded therein. The video data stream includes scalable nested supplemental enhancement information, the scalable nested supplemental enhancement information including decoding parameter set information for each output layer set of a plurality of output layer sets. The apparatus is configured to process the input bitstream in dependence on an output layer set of the plurality of output layer sets to obtain an output bitstream, such that the output bitstream includes the decoding parameter set information for the output layer set and such that the output bitstream does not include the decoding parameter set information for any other output layer set of the plurality of output layer sets.

[0031] Furthermore, according to an embodiment, a video data stream having video encoded therein is provided, wherein the video data stream includes scalable nested supplemental enhancement information including decoding parameter set information for each of a plurality of output layer sets.

[0032] Furthermore, according to an embodiment, a video encoder for encoding a video into a video data stream is provided, such that the video data stream has the video encoded therein. The video encoder is configured to generate the video data stream such that the video data stream includes scalable nested supplemental enhancement information, the scalable nested supplemental enhancement information including decoding parameter set information for each output layer set in a plurality of output layer sets.

[0033] Furthermore, according to an embodiment, a video data stream having video encoded therein is provided. The video data stream includes an indication of a plurality of operation points. Furthermore, the video data stream includes a first mapping for each of the plurality of operation points, the first mapping assigning one or more of a plurality of profile levels to the operation point. Furthermore, the video data stream includes a plurality of second mappings, wherein each of the plurality of mappings assigns one of the plurality of operation points to one of a plurality of output level sets.

[0034] Furthermore, according to an embodiment, a video encoder for encoding a video into a video data stream is provided, such that the video data stream has the video encoded therein. The video encoder is configured to generate the video data stream such that the video data stream includes an indication of a plurality of operation points. Furthermore, the video encoder is configured to generate the video data stream such that the video data stream includes a first mapping for each of the plurality of operation points, the first mapping assigning one or more of a plurality of profile levels to the operation point. Furthermore, the video encoder is configured to generate the video data stream such that the video data stream includes a plurality of second mappings, wherein each of the plurality of mappings assigns one of the plurality of operation points to one of a plurality of output layer sets.

[0035] Furthermore, according to an embodiment, an apparatus is provided for receiving a video data stream as an input bitstream, wherein the video data stream has video encoded therein. The video data stream includes an indication of a plurality of operation points. Furthermore, the video data stream includes a first mapping for each of the plurality of operation points, the first mapping mapping one or more profile levels of a plurality of profile levels to the operation point. Furthermore, the video data stream includes a plurality of second mappings, wherein each of the plurality of mappings maps one of the plurality of operation points to an output level in a plurality of output level sets. The apparatus is configured to process the input bitstream in dependence on an output layer set in a plurality of output layer sets and in dependence on at least one of the second mappings mapping one of the plurality of operation points in the output level set to produce an output bitstream.

[0036] Furthermore, according to an embodiment, a method for receiving a video data stream as an input bitstream is provided, wherein the video data stream has video encoded therein. The video data stream comprises initial profile level information. Furthermore, the video data stream comprises a first set of two or more layers. The method comprises processing the input bitstream in dependence on an output layer set of a plurality of output layer sets to generate an output bitstream by removing at least one layer from the first set of two or more layers to obtain a second set of one or more layers, such that the output bitstream comprises the second set of one or more layers without the at least one layer that has been removed from the first set. Furthermore, the method comprises removing from the initial profile level information those parts of the initial profile level information that are not relevant to at least one layer of the second set of one or more layers.

[0037] Furthermore, according to an embodiment, a method is provided for receiving a video data stream as an input bitstream, wherein the video data stream has video encoded therein. The video data stream includes scalable nested supplemental enhancement information, the scalable nested supplemental enhancement information including decoding parameter set information for each output layer set of a plurality of output layer sets. The method includes processing the input bitstream to obtain an output bitstream in dependence on an output layer set of the plurality of output layer sets, such that the output bitstream includes the decoding parameter set information for the output layer set and such that the output bitstream does not include the decoding parameter set information for any other output layer set of the plurality of output layer sets.

[0038] Furthermore, according to an embodiment, a method for encoding a video into a video data stream is provided, wherein the video data stream has the video encoded therein. The method includes generating the video data stream such that the video data stream includes scalable nested supplemental enhancement information, wherein the scalable nested supplemental enhancement information includes decoding parameter set information for each output layer set in a plurality of output layer sets.

[0039] Furthermore, according to an embodiment, a method for encoding video into a video data stream is provided, such that the video data stream has video encoded therein. The method includes generating the video data stream such that the video data stream includes an indication of a plurality of operation points. Furthermore, the method includes generating the video data stream such that the video data stream includes a first mapping for each of the plurality of operation points, the first mapping assigning one or more of a plurality of profile levels to the operation point. Furthermore, the method includes generating the video data stream such that the video data stream includes a plurality of second mappings, wherein each of the plurality of mappings assigns one of the plurality of operation points to one of a plurality of output layer sets.

[0040] Furthermore, according to an embodiment, a method is provided for receiving a video data stream as an input bitstream, wherein the video data stream has video encoded therein. The video data stream includes an indication of a plurality of operation points. Furthermore, the video data stream includes a first mapping for each of the plurality of operation points, the first mapping mapping one or more of a plurality of profile levels to the operation point. Furthermore, the video data stream includes a plurality of second mappings, wherein each of the plurality of mappings maps one of the plurality of operation points to one of a plurality of output level sets. The method includes processing the input bitstream dependent on an output layer set from a plurality of output layer sets and dependent on at least one of the second mappings mapping one of the plurality of operation points of the output level set to produce an output bitstream.

[0041] Furthermore, a computer program is provided for implementing one of the methods as described above when executed on a computer or a signal processor.

[0042] According to an embodiment, a video data stream having video encoded therein is provided. The video data stream includes a sequence parameter set and a decoding parameter set. The sequence parameter set includes a first indicator value that indicates a maximum number of sublayers or a constant value less than the maximum number of sublayers based on information stored in the sequence parameter set. The decoding parameter set includes a second indicator value that indicates the maximum number of sublayers or the constant value less than the maximum number of sublayers based on information stored in the decoding parameter set. The first indicator value is less than or equal to the second indicator value.

[0043] Furthermore, according to an embodiment, a video encoder for encoding video into a video data stream is provided, such that the video data stream has the video encoded therein. The video encoder is configured to generate the video data stream such that the video data stream includes a sequence parameter set and a decoding parameter set. The video encoder is configured to generate the video data stream such that the sequence parameter set includes a first indicator value, the first indicator value indicating a maximum number of sublayers or the maximum number of sublayers minus a constant value based on information stored in the sequence parameter set. Furthermore, the video encoder generates the video data stream such that the decoding parameter set includes a second indicator value indicating the maximum number of sublayers or the maximum number of sublayers minus the constant value based on information stored in the decoding parameter set. Furthermore, the video encoder is configured to generate the video data stream such that the first indicator value is less than or equal to the second indicator value.

[0044] Furthermore, according to an embodiment, a video encoder is provided for encoding video into a video data stream, such that the video data stream has the video encoded therein. The video encoder is configured to generate the video data stream such that the video data stream includes a sequence parameter set and a decoding parameter set. Furthermore, the video encoder is configured to generate the video data stream such that the sequence parameter set includes a first indicator value, the first indicator value indicating a maximum number of sublayers or the maximum number of sublayers minus a constant value based on information stored in the sequence parameter set. Furthermore, the video encoder generates the video data stream such that the decoding parameter set includes a second indicator value indicating the maximum number of sublayers or the maximum number of sublayers minus the constant value based on information stored in the decoding parameter set. The second indicator value of the decoding parameter set is an upper bound that takes precedence over the first indicator value of the sequence parameter set.

[0045] Furthermore, according to an embodiment, an apparatus is provided for receiving a video data stream as an input bitstream, wherein the video data stream has video encoded therein. The video data stream includes a sequence parameter set and a decoding parameter set. The sequence parameter set includes a first indicator value, the first indicator value indicating a maximum number of sublayers or the maximum number of sublayers minus a constant value based on information stored in the sequence parameter set. The decoding parameter set includes a second indicator value, the second indicator value indicating the maximum number of sublayers or the maximum number of sublayers minus the constant value based on information stored in the decoding parameter set. If the first indicator value is greater than the second indicator value, the apparatus is configured to process the input bitstream to generate an output bitstream, such that the output bitstream includes the second indicator value as an indication of the maximum number of sublayers or the maximum number of sublayers minus the constant value.

[0046] Furthermore, according to an embodiment, a method for encoding video into a video data stream is provided, such that the video data stream has the video encoded therein. The method includes generating the video data stream such that the video data stream includes a sequence parameter set and a decoding parameter set. Furthermore, the method includes generating the video data stream such that the sequence parameter set includes a first indicator value, the first indicator value indicating a maximum number of sublayers or the maximum number of sublayers minus a constant value based on information stored in the sequence parameter set. Furthermore, the method includes generating the video data stream such that the decoding parameter set includes a second indicator value indicating the maximum number of sublayers or the maximum number of sublayers minus the constant value based on information stored in the decoding parameter set. The method includes generating the video data stream such that the first indicator value is less than or equal to the second indicator value.

[0047] Furthermore, according to an embodiment, a method for encoding video into a video data stream is provided, such that the video data stream has the video encoded therein. The method includes generating the video data stream such that the video data stream includes a sequence parameter set and a decoding parameter set. Furthermore, the method includes generating the video data stream such that the sequence parameter set includes a first indicator value, the first indicator value indicating a maximum number of sublayers or the maximum number of sublayers minus a constant value based on information stored in the sequence parameter set. Furthermore, the method includes generating the video data stream such that the decoding parameter set includes a second indicator value, the second indicator value indicating the maximum number of sublayers or the maximum number of sublayers minus the constant value based on information stored in the decoding parameter set. The second indicator value of the decoding parameter set is an upper bound that takes precedence over the first indicator value of the sequence parameter set.

[0048] Furthermore, according to an embodiment, a method for receiving a video data stream as an input bitstream is provided, wherein the video data stream has video encoded therein. The video data stream includes a sequence parameter set and a decoding parameter set. The sequence parameter set includes a first indicator value, the first indicator value indicating a maximum number of sublayers or the maximum number of sublayers minus a constant value based on information stored in the sequence parameter set. The decoding parameter set includes a second indicator value, the second indicator value indicating the maximum number of sublayers or the maximum number of sublayers minus the constant value based on information stored in the decoding parameter set, and if the first indicator value is greater than the second indicator value, the method includes processing the input bitstream to generate an output bitstream, such that the output bitstream includes the second indicator value as an indication of the maximum number of sublayers or the maximum number of sublayers minus the constant value.

[0049] Furthermore, a computer program is provided for implementing the method as described above when executed on a computer or a signal processor. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] Figure 1 A video encoder for encoding video into a video data stream according to an embodiment is described.

[0051] Figure 2 An apparatus for receiving an input video data stream according to an embodiment is described.

[0052] Figure 3 A video decoder for receiving a video data stream having video stored therein according to an embodiment is described.

[0053] Figure 4 Examples for using reference picture resampling in conjunction with temporal scalability are described.

[0054] Figure 5 Two splice point options are illustrated, where Option I (top) maintains frame rate throughout the example, and Option II (bottom) has a frame rate drop in the last group of pictures in the bitstream.

[0055] Figure 6 A video encoder is described.

[0056] Figure 7 A video decoder is described.

[0057] Figure 8 The relationship between the reconstruction signal (eg the reconstructed picture) on the one hand and the combination of the prediction residual signal and the prediction signal as signaled in the data stream on the other hand is explained. DETAILED DESCRIPTION

[0058] The following description of the drawings begins with the presentation of a description of an encoder and decoder of a block-based prediction codec for decoding pictures of a video, in order to form an example of a decoding framework into which embodiments of the present invention may be built. Figures 6 to 8 The corresponding encoders and decoders are described. Thereafter, a description of embodiments of the concepts of the invention is given together with information on how these concepts can be implemented into Figure 6 and Figure 7 The descriptions in the encoder and decoder are presented together, although with Figures 1 to 3 The embodiments described below can also be used to form Figure 6 and Figure 7 The encoder and decoder are based on the decoding framework that operates the encoder and decoder.

[0059] Figure 6 A video encoder is shown, an apparatus for predictively decoding a picture 12 into a data stream 14, exemplarily using transform-based residual coding. The apparatus or encoder is denoted by reference numeral 10. Figure 7 A corresponding video decoder 20 is shown, for example, a decoder 20 configured to predictively decode a picture 12′ from the data stream 14 also using transform-based residual decoding, wherein a prime character has been used to indicate that the picture 12′ reconstructed by the decoder 20 deviates from the picture 12 originally encoded by the encoder 10 with respect to decoding losses introduced by quantization of the prediction residual signal. Figure 6 and Figure 7 Transform-based prediction residual coding is used illustratively, although embodiments of the present application are not limited to this type of prediction residual coding. Figure 6 and Figure 7 The same is true for other details of the description, as will be outlined below.

[0060] The encoder 10 is configured to perform a spatial-to-spectral transform on the prediction residual signal and encode the prediction residual signal obtained thereby into the data stream 14. Similarly, the decoder 20 is configured to decode the prediction residual signal from the data stream 14 and perform a spectral-to-spatial transform on the prediction residual signal obtained thereby.

[0061] Internally, the encoder 10 may include a prediction residual signal generator 22 that generates a prediction residual signal 24 to measure the deviation of the prediction signal 26 from the original signal, such as the deviation from the picture 12. The prediction residual signal generator 22 may, for example, be a subtractor that subtracts the prediction signal from the original signal (e.g., from the picture 12). The encoder 10 then further includes a transformer 28 that performs a spatial-to-spectral transform on the prediction residual signal 24 to obtain a spectral-domain prediction residual signal 24'. This spectral-domain prediction residual signal 24' is then quantized by a quantizer 32, also included in the encoder 10. The quantized prediction residual signal 24'' is then decoded into the data stream 14. To this end, the encoder 10 may optionally include an entropy decoder 34 that entropy decodes the transformed and quantized prediction residual signal into the data stream 14. The prediction signal 26 is generated by a prediction stage 36 of the encoder 10 based on the prediction residual signal 24 that is encoded into the data stream 14 and can be decoded from the data stream 14. For this reason, Figure 6 As shown in FIG, the prediction stage 36 may internally include a dequantizer 38 that dequantizes the prediction residual signal 24 ″ to obtain a spectral domain prediction residual signal 24 ′″ corresponding to the prediction residual signal 24 ′ minus quantization losses, followed by an inverse transformer 40 that inversely transforms (e.g., spectral to spatial transformation) the latter prediction residual signal 24 ′″ to obtain a prediction residual signal 24 ′″ corresponding to the original prediction residual signal 24 minus quantization losses. A combiner 42 of the prediction stage 36 then recombines the prediction signal 26 and the prediction residual signal 24 ′″, such as by addition, to obtain a reconstructed signal 46 , e.g., a reconstruction of the original signal. The reconstructed signal 46 may correspond to the signal 26 . The prediction module 44 of the prediction stage 36 then generates the prediction signal 26 based on the signal 26 by using, for example, spatial prediction (e.g., intra-picture prediction) and / or temporal prediction (e.g., inter-picture prediction).

[0062] Likewise, Figure 7As shown in FIG, the decoder 20 may be internally composed of components corresponding to the prediction stage 36 and interconnected in a manner corresponding to the prediction stage 36. In particular, the entropy decoder 50 of the decoder 20 may entropy decode the quantized spectral domain prediction residual signal 24″ from the data stream, whereupon the dequantizer 52, inverse transformer 54, combiner 56, and prediction module 58, which are interconnected and cooperate in the manner described above with respect to the modules of the prediction stage 36, recover the reconstructed signal based on the prediction residual signal 24″, such that Figure 7 As shown in , the output of the combiner 56 results in a reconstructed signal, picture 12 ′.

[0063] Although not specifically described above, it is readily apparent that the encoder 10 can set certain decoding parameters, including, for example, prediction modes, motion parameters, etc., according to some optimization scheme, such as, for example, in a manner that optimizes certain rate- and distortion-related criteria (e.g., decoding cost). For example, the encoder 10 and decoder 20, and corresponding modules 44 and 58, can each support different prediction modes, such as intra-frame decoding mode and inter-frame decoding mode. The granularity at which the encoder and decoder switch between these prediction mode types can correspond to subdividing pictures 12 and 12', respectively, into decoding segments or decoding blocks. For example, in units of these decoding segments, a picture can be subdivided into blocks that are decoded intra-frame and blocks that are decoded inter-frame. As outlined in more detail below, intra-frame decoded blocks are predicted based on the spatial, already decoded / decoded neighborhood of the corresponding block. There may be several intra-coding modes, selected for respective intra-coding segments including directional or angular intra-coding modes, according to which the respective segments are filled by extrapolating sample values ​​of a neighborhood into the respective intra-coding segment along a specific direction specific to the respective directional intra-coding mode. The intra-coding modes may, for example, also include one or more other modes, such as a DC coding mode (according to which the prediction of the respective intra-coding block assigns a DC value to all samples within the respective intra-coding segment), and / or a planar intra-coding mode (according to which the prediction of the respective block is approximated or determined as a spatial distribution of sample values ​​described by a two-dimensional linear function at the sample positions of the respective intra-coding block, with a driving tilt and offset of the plane defined by the two-dimensional linear function based on neighboring samples). In contrast, inter-coding blocks may, for example, be predicted temporally. For inter-coded blocks, a motion vector can be signaled within the data stream, indicating the spatial displacement of a portion of a previously coded picture of the video to which picture 12 belongs, at which portion the previously coded / decoded picture was sampled in order to obtain the prediction signal for the corresponding inter-coded block. This means that, in addition to the residual signal coding included by the data stream 14 (such as the entropy-coded transform coefficient levels representing the quantized spectral domain prediction residual signal 24 ″), the data stream 14 can also encode the following into it: decoding mode parameters for assigning a decoding mode to the individual blocks; prediction parameters for some of the blocks (such as motion parameters for inter-coded segments) and optionally other parameters (such as parameters for controlling and signaling the subdivision of pictures 12 and 12 ′, respectively, into segments). The decoder 20 uses these parameters to subdivide the pictures in the same way as the encoder did, assigning the same prediction mode to the segments and performing the same prediction, resulting in the same prediction signal.

[0064] Figure 8The relationship between the reconstruction signal (e.g. the reconstructed picture 12') on the one hand and the combination of the prediction residual signal 24'''' and the prediction signal 26 as signaled in the data stream 14 on the other hand is illustrated. As already indicated above, the combination can be an addition. The prediction signal 26 is Figure 8 The picture region is illustrated as being subdivided into intra-coded blocks illustratively indicated using shading and inter-coded blocks illustratively indicated using non-shading. The subdivision may be any subdivision, such as a regular subdivision of the picture region into rows and columns of square or non-square blocks, or a subdivision of the picture 12 from a root block multi-tree into a plurality of leaf blocks of different sizes, such as a quadtree subdivision, wherein in Figure 8 Their hybrid is illustrated in

[15] , where the picture region is first subdivided into rows and columns of root blocks and then further subdivided into one or more leaf blocks according to a recursive multitree subdivision.

[0065] Similarly, for intra-coded blocks 80, the data stream 14 may have an intra-coding mode encoded therein that assigns one of several supported intra-coding modes to the corresponding intra-coded block 80. For inter-coded blocks 82, the data stream 14 may have one or more motion parameters encoded therein. In general, inter-coded blocks 82 are not limited to being temporally coded. Alternatively, inter-coded blocks 82 may be any block predicted from a previously coded portion other than the current picture 12 itself, such as a previously coded picture of the video to which picture 12 belongs, or another view or hierarchically lower-level picture if the encoder and decoder are scalable encoders and decoders, respectively.

[0066] Figure 8 The prediction residual signal 24''' in is also illustrated as subdividing the picture area into blocks 84. These blocks may be referred to as transform blocks to distinguish them from the decoded blocks 80 and 82. In practice, Figure 8 It is illustrated that the encoder 10 and the decoder 20 may use two different subdivisions of the pictures 12 and 12', respectively, into blocks, namely one into coded blocks 80 and 82, respectively, and another into transform blocks 84, respectively. The two subdivisions may be identical, for example, each coded block 80 and 82 may simultaneously form a transform block 84, but Figure 8A case is described in which, for example, the subdivision into transform blocks 84 forms an extension of the subdivision into coded blocks 80 and 82, such that any block boundary between blocks 80 and 82 overlaps a boundary between two blocks 84, or alternatively, each block 80 and 82 coincides with either one of transform blocks 84 or a cluster of transform blocks 84. However, the subdivision can also be determined or selected independently of one another, such that transform block 84 can alternatively straddle a block boundary between blocks 80 and 82. With respect to the subdivision into transform blocks 84, similar statements hold true as those made regarding the subdivision into blocks 80 and 82. For example, block 84 can be the result of a regular subdivision of a picture region into blocks (with or without arrangement into rows and columns), the result of a recursive multitree subdivision of a picture region, or a combination thereof, or any other type of blockation. Note, in passing, that blocks 80, 82, and 84 are not limited to quadratic, rectangular, or any other shape.

[0067] Figure 8 It is further illustrated that the combination of the prediction signal 26 and the prediction residual signal 24''' directly results in the reconstructed picture 12'. However, it should be noted that according to alternative embodiments more than one prediction signal 26 may be combined with the prediction residual signal 24''' to result in the picture 12'.

[0068] exist Figure 8 In the embodiment described below, the transform blocks 84 shall have the following meanings. The transformer 28 and the inverse transformer 54 perform their transforms in units of these transform blocks 84. For example, many codecs use some kind of DST or DCT for all transform blocks 84. Some codecs allow skipping of transforms so that the prediction residual signal is decoded directly in the spatial domain for some of the transform blocks 84. However, according to the embodiments described below, the encoder 10 and the decoder 20 are configured in such a way that they support several transforms. For example, the transforms supported by the encoder 10 and the decoder 20 may include:

[0069] o DCT-II (or DCT-III), where DCT stands for Discrete Cosine Transform

[0070] o DST-IV, where DST stands for discrete sine transform

[0071] o DCT-IV

[0072] o DST-VII

[0073] o Identity Transformation (IT)

[0074] Naturally, while the transformer 28 will support all forward transformed versions of these transforms, the decoder 20 or inverse transformer 54 will support their corresponding backward or inverse versions:

[0075] o Inverse DCT-II (or Inverse DCT-III)

[0076] o Reverse DST-IV

[0077] o Inverse DCT-IV

[0078] o Inverse DST-VII

[0079] o Identity Transformation (IT)

[0080] The subsequent description provides more details on which transforms can be supported by the encoder 10 and decoder 20. In any case, it should be noted that the set of supported transforms may include only one transform, such as a spectral to spatial or spatial to spectral transform.

[0081] As already outlined above, Figures 6 to 8

[0045] The examples have been presented as examples in which the inventive concepts described further below may be implemented in order to form specific examples for encoders and decoders according to the present application. Figure 6 and Figure 7 The encoder and decoder of may represent possible implementations of the encoder and decoder described below, respectively. However, Figure 6 and Figure 7 However, an encoder according to an embodiment of the present application may use the concepts outlined in more detail below to perform block-based encoding of picture 12, and with Figure 6 The difference between the encoder and the video encoder is that it is not a video encoder but a still picture encoder, that it does not support inter-frame prediction, or that it uses a different Figure 8 Likewise, a decoder according to an embodiment of the present application may perform block-based decoding of a picture 12' from a data stream 14 using the coding concepts outlined further below, but may be similar to Figure 7 The decoder 20 of the embodiment of the present invention differs in that it is not a video decoder but a still picture decoder, that it does not support intra-frame prediction, or that it decodes in a different manner than that described with respect to FIG. Figure 8 The picture 12 ′ is divided into blocks in the manner described and / or its prediction residual is derived from the data stream 14 , for example not in the transform domain but in the spatial domain.

[0082] Figure 1 A video encoder 100 for encoding a video into a video data stream according to an embodiment is illustrated. The video encoder 100 is configured to generate a video data stream.

[0083] Figure 2An apparatus 200 for receiving an input video data stream according to an embodiment is illustrated. The input video data stream has video encoded therein. The apparatus 200 is configured to generate an output video data stream from the input video data stream.

[0084] Figure 3 A video decoder 300 for receiving a video data stream having video stored therein according to an embodiment is illustrated. The video decoder 300 is configured to decode video from the video data stream.

[0085] Furthermore, a system according to an embodiment is provided. The system comprises Figure 2 The device and Figure 3 Video decoder. Figure 3 The video decoder (300) is configured to receive Figure 2 The device (200) outputs a video data stream. Figure 3 The video decoder 300 is configured to Figure 2 The output video data stream of the apparatus 200 is decoded video.

[0086] In an embodiment, the system may further include, for example Figure 1 The video encoder 100. For example, Figure 2 The apparatus 200 may be configured to Figure 1 The video encoder 100 receives a video data stream as an input video data stream.

[0087] The (optional) intermediate device 210 of the apparatus 200 may, for example, be configured to receive a video data stream as an input video data stream from the video encoder 100 and to generate an output video data stream from the input video data stream. For example, the intermediate device may, for example, be configured to modify (header / metadata) information of the input video data stream and / or may, for example, be configured to delete pictures from the input video data stream and / or may be configured to mix / splice the input video data stream with an additional second bitstream having a second video encoded therein.

[0088] The (optional) video decoder 221 may, for example, be configured to decode video from the output video data stream.

[0089] The (optional) hypothetical reference decoder 222 may, for example, be configured to determine timing information of the video from the output video data stream, or may, for example, be configured to determine buffer information of a buffer in which the video or a portion of the video is to be stored.

[0090] The system includes Figure 1 The video encoder 101 and Figure 2 Video decoder 151.

[0091] The video encoder 101 is configured to generate an encoded video signal. The video decoder 151 is configured to decode the encoded video signal to reconstruct pictures of the video.

[0092] Hereinafter, trick mode and fast playback are described.

[0093] According to an embodiment, an apparatus is provided for receiving a video data stream as an input bitstream, wherein the video data stream has video encoded therein. The apparatus is configured to determine a decoding capability requirement based on a decoding capability requirement constraint for a playback speed modification factor (Nx). The playback speed modification factor (Nx) is either a forward playback speed modification factor or a backward playback speed modification factor.

[0094] In an embodiment, the decoding capability requirement may be, for example, a modified sampling rate.The decoding capability requirement limit may be, for example, a sampling rate limit.

[0095] According to an embodiment, the video data stream may, for example, include information about sampling rate limitations.

[0096] In an embodiment, the apparatus may, for example, be configured to determine the modified sampling rate in dependence on a picture size of the pictures of the video.

[0097] According to an embodiment, the apparatus may eg be configured to determine the modified sampling rate also in dependence on the frame rate and in dependence on the playback speed modification factor (Nx).The video data stream may eg comprise information about the frame rate.

[0098] In an embodiment, the frame rate depends on a sub-layer of the one or more sub-layers. The video data stream may, for example, include sub-layer-specific frame rate information for each of the one or more sub-layers. The apparatus may, for example, be configured to obtain the frame rate from the sub-layer-specific frame rate information for one of the one or more sub-layers.

[0099] According to an embodiment, the apparatus may, for example, be configured to determine that the modified sampling rate is a highest picture size among a plurality of pictures of the video in dependence on a picture size of the pictures of the video.

[0100] In an embodiment, the video data stream may, for example, comprise an application factor.The apparatus may, for example, be configured to calculate the modified sampling rate using the top layer picture size and the application factor.

[0101] According to an embodiment, the maximum picture size depends on a sublayer in the one or more sublayers. The video data stream may, for example, include sublayer-specific maximum picture size information for each of the one or more sublayers. The apparatus may, for example, be configured to obtain the maximum picture size from the sublayer-specific maximum picture size information for one of the one or more sublayers.

[0102] In an embodiment, the video data stream may for example comprise an application factor.The apparatus may for example be configured to use the sampling rate of the video and to determine a modified sampling rate using the application factor.

[0103] According to an embodiment, the decoding capability requirement may be, for example, a modified bit rate. The decoding capability requirement restriction may be, for example, a bit rate restriction.

[0104] According to an embodiment, the video data stream may include information about bit rate limitation, for example.

[0105] According to an embodiment, the apparatus may, for example, be configured to determine the modified bit rate in dependence on an initial bit rate of the video data stream and in dependence on a playback speed modification factor (Nx).

[0106] In an embodiment, the initial bit rate is dependent on the sub-layer video data stream in the one or more sub-layers including sub-layer-specific initial bit rate information that can be configured, for example, as each sub-layer in the one or more sub-layers. The apparatus can be configured, for example, to obtain the initial bit rate from sub-layer-specific frame rate information for one of the one or more sub-layers.

[0107] According to an embodiment, the video data stream may include, for example, an IRAP-only flag (IRAP = Intra Random Access Point), which indicates whether decoding of only IRAP pictures is possible when the initial playback speed is increased. Depending on the IRAP-only flag, the device may, for example, be configured to read decoding capability requirements from the video data stream.

[0108] In an embodiment, the video data stream may include a reference picture only flag, the IRAP flag indicating whether only reference picture decoding is possible when the initial playback speed is increased. Depending on the reference picture only flag, the device may be configured to read the decoding capability requirement from the video data stream.

[0109] Furthermore, according to an embodiment, a video data stream having video encoded therein is provided. The video data stream may, for example, include information about sampling rate limitations, and / or the video data stream may, for example, include information about bit rate limitations.

[0110] According to an embodiment, the video data stream may, for example, include information about the frame rate.

[0111] In an embodiment, the video data stream may, for example, include sub-layer specific frame rate information for each of the one or more sub-layers.

[0112] According to an embodiment, the video data stream may, for example, comprise an application factor for determining the modified sampling rate.

[0113] In an embodiment, the video data stream may, for example, include sub-layer specific maximum picture size information for each of the one or more sub-layers.

[0114] According to an embodiment, the video data stream may, for example, include sub-layer specific initial bit rate information for each of the one or more sub-layers.

[0115] In an embodiment, the video data stream may, for example, include an IRAP-only flag indicating whether decoding of only IRAP pictures is possible when the initial playback speed is increased.

[0116] According to an embodiment, the video data stream may, for example, include a reference picture only flag, the IRAP only flag indicating whether reference picture only decoding is possible when the initial playback speed is increased.

[0117] In an embodiment, the video data stream may, for example, comprise level information depending on the IRAP-only flag; and / or the video data stream may, for example, comprise level information depending on the reference picture-only flag.

[0118] Furthermore, according to an embodiment, a video encoder for encoding a video into a video data stream is provided, such that the video data stream has the video encoded therein. The video encoder is configured to generate the video data stream such that the video data stream includes information about a sampling rate limit, and / or the video encoder is configured to generate the video data stream such that the video data stream includes information about a bit rate limit.

[0119] In an embodiment, the video encoder may, for example, be configured to generate the video data stream such that the video data stream may, for example, include information about the frame rate limitation.

[0120] According to an embodiment, the video encoder may, for example, be configured to generate the video data stream such that the video data stream may, for example, include sub-layer specific frame rate information for each of the one or more sub-layers.

[0121] In an embodiment, the video encoder may, for example, be configured to generate the video data stream such that the video data stream may, for example, comprise an application factor for determining the modified sampling rate.

[0122] According to an embodiment, the video encoder may, for example, be configured to generate a video data stream such that the video data stream may, for example, include sub-layer specific maximum picture size information for each of the one or more sub-layers.

[0123] In an embodiment, the video encoder may, for example, be configured to generate a video data stream such that the video data stream may, for example, include sub-layer specific initial bit rate information for each of the one or more sub-layers.

[0124] According to an embodiment, the video encoder may for example be configured to generate a video data stream such that the video data stream may for example include an IRAP-only flag indicating whether decoding of only IRAP pictures is possible when the initial playback speed may for example be increased.

[0125] In an embodiment, the video encoder may eg be configured to generate the video data stream such that the video data stream may eg include a reference only picture flag indicating whether reference only picture decoding is possible when the initial playback speed may eg be increased.

[0126] According to an embodiment, the video encoder may, for example, be configured to generate a video data stream such that depending on only the IRAP flag, the video data stream may, for example, include level information; and / or the video encoder may, for example, be configured to generate a video data stream such that depending on only the reference picture flag, the video data stream may, for example, include level information.

[0127] In an embodiment, the apparatus may, for example, be configured to decode an input bitstream to decode video.

[0128] Furthermore, according to an embodiment, a system for encoding a video into a video data stream and decoding the video is provided. The system includes a video encoder as described above and an apparatus as described above. The video encoder is configured to encode the video into a video data stream, such that the video data stream has the video encoded therein. The apparatus is configured to receive the video data stream as an input bitstream. Furthermore, the apparatus is configured to decode the input bitstream to decode the video.

[0129] Fast forward operations or reverse operations (fast backward playback) are typical operations performed in video applications. Typically, these operations involve decoding only a subset of the bitstream and decoding it at a higher speed (e.g., frame rate) than the speed indicated by the bitstream. Moreover, even for non-trick mode operation, in some scenarios it may be interesting to play the video at a higher speed. Although there is not a huge difference in fast forward and fast playback, the distinction considered in this description may be that the second goal is continuous and smooth playback (potentially even with audio playback and synchronization), while the first goal is not, for example, fast forward may be a discontinuous playback with many "jumps" in the content. In any case, the result is playback of the content at Nx speed.

[0130] Obviously, playing back content at a higher speed also requires decoding the bitstream at a higher speed, which results in a higher decoding capability requirement (ie, a higher level) as a result of that higher sampling rate and higher bit rate as indicated in the HRD parameters.

[0131] When reference picture resampling (RPR) is not used, i.e., all pictures within a CVS or bitstream have the same size, the sampling rate can be easily calculated as the picture size in samples divided by the frame rate and multiplied by the acceleration factor Nx. Similarly, the bit rate will be multiplied by the acceleration factor Nx.

[0132] In a first embodiment, the two values ​​described, derived as discussed above, are checked against the level limits in VVC, and the level to which Nx Speedup belongs is calculated as the minimum value for which the derived value is less than the level limit.

[0133] Note that the frame rate or bit rate is sub-layer specific, which means that no matter what sub-layer is being played back (e.g., fast forwarding is performed by only decoding temporal level 0), the sub-layer specific signaled frame rate or bit rate is used to calculate the corresponding value.

[0134] However, when using reference picture resampling (RPR), the sampling rate cannot be easily derived, ie only the worst case of the highest picture size can be used, which will result in the derivation of the worst case sampling rate. Figure 4 An example illustrating how to use RPR with temporal scalability.

[0135] Figure 4 An example of using reference picture resampling in conjunction with temporal scalability is illustrated.

[0136] In this example, the image size of the lowest sub-layer might be 1920x1080 and the image size of the highest sub-layer might be 960x540. By using only the largest image size and assuming a 60fps bitstream, approximately 124*10 pixels per second will be exported. 6 The actual value will be approximately 78*10 6 That is, the worst case will be 1.6 times the actual value.

[0137] In one embodiment, the bitstream includes an indication of a factor that needs to be applied to a sampling rate derived using a worst-case value. Using the signaled factor, a worst-case sampling rate can be derived, divided by the signaled factor, and multiplied by the speed decoding factor to calculate the actual sampling rate by accelerating the decoding operation.

[0138] In another embodiment, a maximum picture size value is indicated per sub-layer, which will lead to a better approximation or even a correct value derivation if the picture size does not change within the sub-layer.

[0139] Furthermore, there are cases where sub-layers are not fully utilized. Sometimes, fast forwarding is performed by decoding and replaying only IRAP pictures (e.g., IDR) or pictures that are not non-reference pictures. However, it is unclear what level is required to decode such sub-bitstreams when decoding at Nx speed. In another embodiment, additional level signaling is indicated for either of the following two cases:

[0140] IRAP only: Only IRAP is decoded.

[0141] Reference pictures only: non-reference pictures are discarded and the rest are decoded.

[0142] The syntax of such an embodiment could be as follows.

[0143] For profile_tier_level:

[0144]

[0145]

[0146] Similar to the case of irap, only reference pictures can be added to the grammar.

[0147] The code rate and CPB size can be similarly signaled as:

[0148]

[0149] Hereinafter, protection of delay_for_concatenation_ensured_flag is described.

[0150] According to an embodiment, a video data stream having video encoded therein is provided. The video data stream includes a plurality of access units, wherein an access unit group of the video data stream includes the plurality of access units. The video data stream also includes, for each access unit of a subset of access units, an indication indicating that a seamless splice point exists after the access unit. The subset of access units is a proper subset of the access unit group of the video data stream.

[0151] In an embodiment, a video data stream may, for example, include multiple temporal sub-layers. An indication of the presence of a seamless splicing point after an access unit is present only for access units of a first group of temporal sub-layers among the multiple temporal sub-layers, but an indication of the presence of a seamless splicing point after the access unit is not present for access units of a second group of temporal sub-layers among the multiple temporal sub-layers.

[0152] According to an embodiment, the video data stream may include, for example, buffering period supplemental enhancement information. The buffering period supplemental enhancement information indicates a subset of access units among a plurality of access units for which a seamless splicing point exists after the access unit.

[0153] In an embodiment, a video data stream may include, for example, a divisor. A delay value may be assigned to each of a plurality of access units in the video data stream. The indication that a seamless splice point exists in an access unit is present only for those access units in the plurality of access units in the video data stream that have an assigned delay value, which may, for example, be equal to a remainder value when divided by the divisor.

[0154] According to an embodiment, the remainder value may be, for example, 0.

[0155] In an embodiment, the video data stream may, for example, include remainder values.

[0156] According to an embodiment, the indication for each access unit of the subset of access units that there is a seamless splicing point after the access unit may be, for example, a flag.

[0157] In an embodiment, the flag may be, for example, delay_for_concatenation_ensured_flag.

[0158] According to an embodiment, if, for example, a constant frame rate for output pictures obtained by decoding the video data stream cannot be guaranteed, the flag exhibits a first value. If, for example, a constant frame rate for output pictures obtained by decoding the video data stream can be guaranteed, the flag may, for example, exhibit a second value different from the first value.

[0159] In an embodiment, the flag exhibits the second value if the output pictures resulting from the decoded video data stream have equidistant output times.

[0160] Furthermore, according to an embodiment, a video encoder for encoding a video into a video data stream is provided, such that the video data stream has the video encoded therein. The video encoder is configured to generate the video data stream such that the video data stream includes a plurality of access units, wherein an access unit group of the video data stream includes the plurality of access units, and further, the video encoder is configured to generate the video data stream such that the video data stream includes, for each access unit of a subset of access units, an indication indicating that a seamless splice point exists after the access unit. Furthermore, the video encoder is configured to generate the video data stream such that the subset of access units is a proper subset of the access unit group of the video data stream.

[0161] In an embodiment, the video encoder may be configured to generate a video data stream such that the video data stream may include information about a plurality of temporal sub-layers. The video encoder may be configured to generate the video data stream such that an indication of the presence of a seamless splicing point after an access unit exists only for access units of a first group of temporal sub-layers among the plurality of temporal sub-layers, but the video encoder may be configured to generate the video data stream such that the indication of the presence of a seamless splicing point after the access unit does not exist for access units of a second group of temporal sub-layers among the plurality of temporal sub-layers.

[0162] According to an embodiment, the video encoder may be configured to generate a video data stream such that the video data stream may include, for example, buffering period supplemental enhancement information. The video encoder may be configured to generate a video data stream such that the buffering period supplemental enhancement information indicates a subset of access units among a plurality of access units for which a seamless splicing point exists after the access unit.

[0163] In an embodiment, a video encoder may be configured to generate a video data stream, for example, such that the video data stream may include a divisor. The video encoder may be configured to generate the video data stream, for example, such that a delay value may be assigned to each of a plurality of access units of the video data stream. The video encoder may be configured to generate the video data stream, for example, such that an indication that a seamless splice point exists in an access unit is present only for those access units of the plurality of access units of the video data stream that have an assigned delay value, which may be, for example, equal to a remainder value when divided by the divisor.

[0164] According to an embodiment, the video encoder may, for example, be configured to generate the video data stream such that the remainder value may, for example, be zero.

[0165] In an embodiment, the video encoder may, for example, be configured to generate the video data stream such that the video data stream may, for example, include the remainder value.

[0166] According to an embodiment, the video encoder may eg be configured to generate the video data stream such that the indication for each access unit of the subset of access units that there is a seamless splicing point after the access unit may eg be a flag.

[0167] In an embodiment, the video encoder may eg be configured to generate the video data stream such that the indication for each access unit of the subset of access units that there is a seamless splice point after the access unit may eg be a flag.

[0168] According to an embodiment, if a constant frame rate for output pictures obtained by decoding the video data stream cannot be guaranteed, for example, the video encoder is configured to generate the video data stream so that the flag exhibits a first value. If a constant frame rate for output pictures obtained by decoding the video data stream can be guaranteed, for example, the video encoder can be configured to generate the video data stream so that the flag exhibits a second value different from the first value.

[0169] In an embodiment, if output pictures resulting from the decoded video data stream have equidistant output times, the video encoder may, for example, be configured to generate the video data stream such that the flag exhibits the second value.

[0170] Furthermore, according to an embodiment, an apparatus is provided for receiving a video data stream as an input bitstream, wherein the video data stream has video encoded therein. The apparatus is configured to process the video data stream. The video data stream includes a plurality of access units, wherein an access unit group of the video data stream includes the plurality of access units. Furthermore, the video data stream includes, for each access unit of a subset of access units, an indication indicating that a seamless splice point exists after the access unit. The access unit subset is a suitable subset of the access unit group of the video data stream. The apparatus is configured to process the indication.

[0171] In an embodiment, a video data stream may, for example, include multiple temporal sub-layers. An indication of the presence of a seamless splicing point after an access unit is present only for access units of a first group of temporal sub-layers among the multiple temporal sub-layers, but an indication of the presence of a seamless splicing point after the access unit is not present for access units of a second group of temporal sub-layers among the multiple temporal sub-layers.

[0172] According to an embodiment, a video data stream may include, for example, buffering period supplemental enhancement information. The buffering period supplemental enhancement information indicates a subset of access units among a plurality of access units for which a seamless splicing point exists after the access unit. The apparatus may, for example, be configured to process the buffering period supplemental enhancement information.

[0173] In an embodiment, a video data stream may, for example, include a divisor. The apparatus may, for example, be configured to receive the divisor. A delay value may, for example, be assigned to each of a plurality of access units in the video data stream. The indication that a seamless splice point exists in an access unit is present only for those access units in the plurality of access units in the video data stream that have an assigned delay value, which may, for example, be equal to a remainder value when divided by the divisor.

[0174] According to an embodiment, the remainder value may be, for example, 0.

[0175] In an embodiment, the video data stream may, for example, include remainder values.

[0176] The apparatus may, for example, be configured to receive a remainder value.

[0177] According to an embodiment, the indication for each access unit of the subset of access units that there is a seamless splicing point after the access unit may eg be a flag.The apparatus may eg be configured to process the flag.

[0178] In an embodiment, the flag may be, for example, delay_for_concatenation_ensured_flag.

[0179] According to an embodiment, if a constant frame rate for output pictures obtained by decoding the video data stream cannot be guaranteed, for example, the flag exhibits a first value. If a constant frame rate for output pictures obtained by decoding the video data stream can be guaranteed, for example, the flag exhibits a second value different from the first value.

[0180] In an embodiment, the flag exhibits the second value if the output pictures resulting from the decoded video data stream have equidistant output times.

[0181] According to an embodiment, the apparatus may, for example, be configured to decode an input bitstream to decode a video.

[0182] Furthermore, according to an embodiment, a system for encoding a video into a video data stream and decoding the video is provided. The system includes a video encoder as described above and an apparatus as described above. The video encoder is configured to encode the video into a video data stream, such that the video data stream has the video encoded therein. The apparatus is configured to receive the video data stream as an input bitstream. Furthermore, the apparatus is configured to decode the input bitstream to decode the video.

[0183] The current VVC draft specification includes instructions for devices that allow splicing, that is, devices that concatenate portions of bitstream A and bitstream B to find AUs in bitstream A that allow seamless splicing of bitstream B. Seamless splicing means that the CPB removal time distance between the last AU of bitstream A and the first AU of bitstream B is the same as the CPB removal time of two consecutive AUs in bitstream A. Conversely, non-seamless splicing means that because the CPB of the first AU of bitstream B cannot be sent to the decoder at the correct time (i.e., the last AU of bitstream A takes longer than expected to be sent), the decoding (CPB removal) of the first AU of bitstream B is delayed, and the AU cannot be decoded as quickly as desired. The current picture timing SEI message in the VVC draft specification contains a flag called delay_for_concatenation_ensured_flag, which, when set to 1, indicates that for a particular associated AU, as long as the following AU (i.e., the start of bitstream B) has a BP SEI message with delay_for_concatenation_ensured_flag equal to 1, and the selected InitCpbRemovalDelay is less than or equal to the value of max_initial_removal_delay_for_concatination indicated in bitstream A, then the splicing at that bitstream position will be seamless. Therefore, for all these AUs of bitstream A, delay_for_concatenation_ensured_flag indicates a variable option for seamless splicing.

[0184] The problem is that not all access units given the above characteristics will also represent meaningful splice points, and signaling the delay_for_concatenation_ensured_flag even for insignificant splice points can cause some poorly implemented splicing devices to splice at a less-than-optimal position. Furthermore, indicating the delay_for_concatenation_ensured_flag for all access units, even those that are not meaningful splice points, incurs rate and processing overhead on the encoder side to correctly set the flag value, even when the flag value is of no use to the splicing device. In this context, an insignificant splice point means that even if the frame rate at the actual splice point is maintained, it is possible that some of these splice points will omit a portion of bitstream A, resulting in non-continuous playback of bitstream A before the splice point, for example, when splicing in the middle of a hierarchical GOP. Figure 5An example is shown in Figure 1, where two bitstreams (A and B) are spliced ​​at two positions (I and II). The first option (top) leads to a constant frame rate throughout the example, while the second option (bottom) causes the frame rate to drop due to the wrong splicing point chosen.

[0185] Figure 5 Two splice point options are illustrated, where Option I (top) maintains the frame rate throughout the example, while Option II (bottom) has a frame rate drop in the last GOP of bitstream A.

[0186] Therefore, the present invention reduces signaling overhead and avoids misleading the splicing device by selectively signaling delay_for_concatenation_ensured_flag for a subset of access units.

[0187] In one embodiment, the delay_for_concatenation_ensured_flag is signaled only for AUs of a special timing inter-layer, where the special timing inter-layer is indicated in the related BP SEI message.

[0188] In another embodiment, the delay_for_concatenation_ensured_flag is signaled only at a subset of all AUs in the buffering period, where the AU subset is identified by information from the associated BP SEI message.

[0189] In one embodiment, the AU subset is indicated by signaling a divisor and an optional remainder value. The delay_for_concatenation_ensured_flag in the PT SEI message for the buffering period is signaled only if the value of the associated AU (e.g., cpb_removal_delay or dpb_output_delay in the PT SEI message) divided by the divisor is equal to the optional remainder, or equal to zero in the case of no divisor.

[0190] In another embodiment, if the output picture cannot ensure a constant frame rate as shown in the figure, the value of delay_for_concatenation_ensured_ is set to 0. Note that the current specification currently only focuses on decoding time and ensures that the difference between the last arrival time (the AU has fully arrived at the CPB) and the CPB removal delay of the AU with this flag set to 1 meets the condition (i.e., > the threshold). Therefore, the decoding time of consecutive AUs is equidistant. Here, the present invention will add that if this flag is set to 1, the output times of consecutive AUs in bitstream A also have equidistant output times.

[0191] In the following, output layer set specific HRD is described.

[0192] According to an embodiment, an apparatus for receiving a video data stream as an input bitstream is provided, wherein the video data stream has video encoded therein. The apparatus is configured to process the input bitstream to obtain sub-bits in dependence on an output layer set. If the output layer set does not include a predefined layer (vps_layer_id[0]) among a plurality of layers present in the video data stream, the apparatus is configured to remove at least one non-scalable nested supplemental enhancement information message assigned with the predefined layer (vps_layer_id[0]). If the output layer set includes the predefined layer (vps_layer_id[0]), the apparatus is configured not to remove any non-scalable nested supplemental enhancement information message assigned with the predefined layer (vps_layer_id[0]).

[0193] Wherein, if there is no output layer set consisting of all layers of the multiple layers in the current decoded video sequence of the video data stream, there may be, for example, no non-extensible nested supplementary enhancement information message having a buffer period payload, a picture timing payload, and at least one of a decoding unit information payload.

[0194] According to an embodiment, if there is no output layer set consisting of all layers of the multiple layers in the current decoded video sequence of the video data stream, there may be, for example, no non-extensible nested supplementary enhancement information message having the buffering period payload and the picture timing payload and at least one of the decoding unit information payload and the sub-picture level information payload.

[0195] In an embodiment, if there is no output layer set consisting of all layers of the multiple layers in the current decoded video sequence of the video data stream, there may be, for example, no non-extensible nested supplementary enhancement information message whose payload type is equal to the first value indicating the buffering period payload, or equal to the second value indicating the picture timing payload, or equal to the third value indicating the decoding unit information payload, or equal to the fourth value indicating the sub-picture level information payload.

[0196] According to an embodiment, for the non-scalable nested supplemental enhancement information message whose payload type is equal to the first value, or equal to the second value, or equal to the third value, or equal to the fourth value, when present, the non-scalable nested supplemental enhancement information message is applied to all the output layer sets consisting of all layers in the current decoded video sequence in the entire bitstream.

[0197] In an embodiment, if there is no output layer set consisting of all layers of the multiple layers in the current decoded video sequence of the video data stream, there may be, for example, no non-scalable nested supplementary enhancement information message whose payload type is equal to 0 indicating the buffering period payload, or equal to 1 indicating the picture timing payload, or equal to 130 indicating the decoding unit payload, or equal to 203 indicating the sub-picture level information payload.

[0198] According to an embodiment, for the non-scalable nested supplemental enhancement information message with payload type equal to 0 or equal to 1 or equal to 130 or equal to 203, when present, the non-scalable nested supplemental enhancement information message applies to all the output layer sets consisting of all layers in the current decoded video sequence in the entire bitstream.

[0199] In an embodiment, if the output layer set does not include the predefined layer (vps_layer_id[0]), the apparatus may, for example, be configured to remove all nested supplemental enhancement information messages allocated together with the predefined layer (vps_layer_id[0]).

[0200] Furthermore, according to an embodiment, a video data stream having video encoded therein is provided. The video data stream includes a video parameter set. The video parameter set includes a plurality of output layer sets. If none of the plurality of output layer sets includes all of a plurality of layers present in the video data stream, then the video data stream does not include any non-scalable nested supplemental enhancement information message for a hypothetical reference decoder.

[0201] Wherein, if there is no output layer set consisting of all layers of the multiple layers in the current decoded video sequence of the video data stream, there may be, for example, no non-extensible nested supplementary enhancement information message having a buffer period payload, a picture timing payload, and at least one of a decoding unit information payload.

[0202] According to an embodiment, if there is no output layer set consisting of all layers of the multiple layers in the current decoded video sequence of the video data stream, there may be, for example, no non-extensible nested supplementary enhancement information message having the buffering period payload and the picture timing payload and at least one of the decoding unit information payload and the sub-picture level information payload.

[0203] In an embodiment, if there is no output layer set consisting of all layers of the multiple layers in the current decoded video sequence of the video data stream, there may be, for example, no non-extensible nested supplementary enhancement information message whose payload type is equal to the first value indicating the buffering period payload, or equal to the second value indicating the picture timing payload, or equal to the third value indicating the decoding unit information payload, or equal to the fourth value indicating the sub-picture level information payload.

[0204] According to an embodiment, for the non-scalable nested supplemental enhancement information message whose payload type is equal to the first value, or equal to the second value, or equal to the third value, or equal to the fourth value, when present, the non-scalable nested supplemental enhancement information message is applied to all the output layer sets consisting of all layers in the current decoded video sequence in the entire bitstream.

[0205] In an embodiment, if there is no output layer set consisting of all layers of the multiple layers in the current decoded video sequence of the video data stream, there may be, for example, no non-scalable nested supplementary enhancement information message whose payload type is equal to 0 indicating the buffering period payload, or equal to 1 indicating the picture timing payload, or equal to 130 indicating the decoding unit payload, or equal to 203 indicating the sub-picture level information payload.

[0206] According to an embodiment, for the non-scalable nested supplemental enhancement information message with payload type equal to 0 or equal to 1 or equal to 130 or equal to 203, when present, the non-scalable nested supplemental enhancement information message applies to all the output layer sets consisting of all layers in the current decoded video sequence in the entire bitstream.

[0207] In an embodiment, if said video data stream comprises at least one non-scalable nested supplemental enhancement information message for said hypothetical reference decoder, at least one of said plurality of output layer sets may, for example, comprise all layers of said plurality of layers present in said video data stream.

[0208] According to an embodiment, all supplementary enhancement information messages for the hypothetical reference decoder are extensible nested supplementary enhancement information messages, and all supplementary enhancement information messages are independent of the output layer sets of the multiple output layer sets, and the output layer sets of the multiple output layer sets may, for example, include all layers of the multiple layers present in the video data stream.

[0209] In an embodiment, all non-scalable nested supplemental enhancement information messages for said hypothetical reference decoder are associated with an output layer set of said plurality of output layer sets, which may for example comprise all layers of said plurality of layers present in said video data stream.

[0210] Furthermore, according to an embodiment, a video encoder is provided for encoding video into a video data stream, such that the video data stream has the video encoded therein. The video encoder is configured to generate the video data stream such that the video data stream includes a video parameter set. Furthermore, the video encoder is configured to generate the video data stream such that the video parameter set includes a plurality of output layer sets. If none of the plurality of output layer sets includes all of a plurality of layers present in the video data stream, the video encoder is configured to generate the video data stream such that the video data stream does not include any non-scalable nested supplemental enhancement information messages for a hypothetical reference decoder.

[0211] Wherein, if, for example, there may not be an output layer set consisting of all layers of the multiple layers in the current decoded video sequence of the video data stream, then, for example, there may not be a non-extensible nested supplementary enhancement information message having a buffer period payload, a picture timing payload, and at least one of a decoding unit information payload.

[0212] According to an embodiment, if, for example, there may not be an output layer set consisting of all layers of the multiple layers in the current decoded video sequence of the video data stream, then, for example, there may not be a non-extensible nested supplementary enhancement information message having the buffering period payload and the picture timing payload and at least one of the decoding unit information payload and the sub-picture level information payload.

[0213] In an embodiment, if there is no output layer set consisting of all layers of the multiple layers in the current decoded video sequence of the video data stream, there may be, for example, no non-extensible nested supplementary enhancement information message whose payload type is equal to the first value indicating the buffering period payload, or equal to the second value indicating the picture timing payload, or equal to the third value indicating the decoding unit information payload, or equal to the fourth value indicating the sub-picture level information payload.

[0214] According to an embodiment, for the non-scalable nested supplemental enhancement information message whose payload type is equal to the first value, or equal to the second value, or equal to the third value, or equal to the fourth value, when present, the non-scalable nested supplemental enhancement information message is applied to all the output layer sets consisting of all layers in the current decoded video sequence in the entire bitstream.

[0215] In an embodiment, if there is no output layer set consisting of all layers of the multiple layers in the current decoded video sequence of the video data stream, there is no non-scalable nested supplementary enhancement information message with the payload type equal to 0 indicating the buffering period payload, or equal to 1 indicating the picture timing payload, or equal to 130 indicating the decoding unit payload, or equal to 203 indicating the sub-picture level information payload.

[0216] According to an embodiment, for the non-scalable nested supplemental enhancement information message with payload type equal to 0 or equal to 1 or equal to 130 or equal to 203, when present, the non-scalable nested supplemental enhancement information message applies to all the output layer sets consisting of all layers in the current decoded video sequence in the entire bitstream.

[0217] In an embodiment, if a video data stream may, for example, include at least one non-scalable nested supplemental enhancement information message for a hypothetical reference decoder, the video encoder may, for example, be configured to generate the video data stream such that at least one of the multiple output layer sets includes all layers of the multiple layers present in the video data stream.

[0218] According to an embodiment, the video encoder can be configured, for example, to generate the video data stream so that all supplemental enhancement information messages for the hypothetical reference decoder are extensible nested supplemental enhancement information messages, and all supplemental enhancement information messages are independent of the output layer sets of the multiple output layer sets, and the output layer sets of the multiple output layer sets can, for example, include all layers of the multiple layers present in the video data stream.

[0219] In an embodiment, the video encoder may, for example, be configured to generate the video data stream such that all non-scalable nested supplemental enhancement information messages for the hypothetical reference decoder are associated with output layer sets of the multiple output layer sets, which may, for example, include all layers of the multiple layers present in the video data stream.

[0220] According to an embodiment, the apparatus may, for example, be configured to decode the input bitstream to decode the video.

[0221] The current VVC draft specification provides a means to carry HRD timing information (BP SEI message, PT SEI message, DUI SEI message) in a nested form in the bitstream that applies only to sub-bitstreams (i.e., OLSs) using scalable nested SEI messages indicating the respective OLs. Other HRD SEI messages carried directly in the bitstream (as non-scalable nested SEI messages) need to apply to the 0th OLS, i.e., the single-layer sub-bitstream indicated in the VPS via the syntax element vps_layer_id[0]. However, as is typical for layered coding, removing this or any other layer from the bitstream does not require rewriting the vps, so the layer with nuh_layer_id equal to vps_layer_id[0] may not even exist.

[0222] The problem with this requirement is that when extracting OLSs other than the 0th OLS, no non-scalable nested SEI messages are allowed, even if the extracted OLS does include a layer with nuh_layer_id equal to vps_layer_id[0]. Such non-scalable nested messages can only be retained in the bitstream after rewriting the corresponding values ​​in the VPS (vps_layer_id[0]) and the NAL unit headers of each layer.

[0223] Two aspects of the present invention that address this problem are described below, as follows.

[0224] 1) As a first solution, the complexity described above is not touched, but the extraction process is adjusted in such a way that when the extracted OLS does not include a layer with nuh_layer_id equal to vps_layer_id[0], the extraction process only removes the non-scalable nesting SEI message that applies to the layer with nuh_layer_id equal to vps_layer_id[0]. Therefore, the non-scalable nesting SEI message does not need to be deleted unnecessarily (or in other words, it does not need to be copied unnecessarily for each OLS that includes a layer with nuh_layer_id equal to vps_layer_id[0]).

[0225] 2) As a second solution, when no OLS corresponding to the complete bitstream (i.e., all layers therein) is defined in the VPS, non-scalable nested SEI messages (BP, PT, DUI) are not allowed in the bitstream. Otherwise, when OLS corresponding to all layers present in the bitstream are present, individual HRD SEI messages are allowed in the bitstream in a non-scalable nested manner. All other HRD SEI messages (applicable to other OLSs) must be within the scalable nested SEI message. Vice versa, all non-scalable nested SEI messages correspond to the current bitstream as a whole.

[0226] For the non-scalable nested supplementary enhancement information message, when payloadType is equal to 0 (indicating buffering period, BP / content), 1 (indicating picture timing, PT, content), 130 (indicating decoding unit information, DUI, content) or 203 (indicating sub-picture level information, SLI, content), the non-scalable nested supplementary enhancement information message applies to all output layer sets, which, when present, consist of all layers in the current coded video sequence in the entire bitstream.

[0227] When there is no output layer set consisting of all layers in the current coded video sequence of the entire bitstream, there shall be no non-scalable nested supplemental enhancement messages with payloadType equal to 0 (BP), 1 (PT), 130 (DUI) or 203 (SLI).

[0228] Hereinafter, the rewriting sublayer PTL in the DPS is described.

[0229] According to an embodiment, an apparatus is provided for receiving a video data stream as an input bitstream, wherein the video data stream has video encoded therein. The video data stream includes initial profile level information. Furthermore, the video data stream includes a first set of two or more layers. The apparatus is configured to process the input bitstream in dependence on an output layer set of a plurality of output layer sets to generate an output bitstream by removing at least one layer from the first set of two or more layers to obtain a second set of one or more layers, such that the output bitstream includes the second set of one or more layers without including the at least one layer that has been removed from the first set. Furthermore, the apparatus is configured to remove from the initial profile level information those portions of the initial profile level information that are not relevant to at least one layer of the second set of one or more layers.

[0230] According to an embodiment, the video data stream may, for example, include initial decoding parameter set profile level information as the initial profile level information. The apparatus may, for example, be configured to remove from the initial decoding parameter set profile level information those parts of the initial decoding parameter set profile level information that are not relevant to the at least one layer of the second set of one or more layers.

[0231] Furthermore, according to an embodiment, an apparatus for receiving a video data stream as an input bitstream is provided, wherein the video data stream has video encoded therein.

[0232] The video data stream includes scalable nested supplemental enhancement information, the scalable nested supplemental enhancement information including decoding parameter set information for each output layer set of a plurality of output layer sets. The apparatus is configured to process the input bitstream in dependence on the output layer set of the plurality of output layer sets to obtain an output bitstream, such that the output bitstream includes the decoding parameter set information for the output layer set and such that the output bitstream does not include the decoding parameter set information for any other output layer set of the plurality of output layer sets.

[0233] In an embodiment, the video data stream may, for example, comprise a raw byte sequence payload.The raw byte sequence payload may, for example, comprise the decoding parameter set information for each output layer set of the plurality of output layer sets.

[0234] According to an embodiment, the video data stream may, for example, comprise a scalable nested supplemental enhancement message, which may, for example, comprise said decoding parameter set information for each of said plurality of output layer sets.

[0235] In an embodiment, the video data stream may include, for example, a nesting type flag (nesting_type), wherein the nesting type flag (nesting_type) includes one or more bits.

[0236] According to an embodiment, if the nesting type flag (nesting_type) exhibits a first value, this indicates that the extensible nested supplementary enhancement message may, for example, include one or more decoding parameter sets, the one or more decoding parameter sets including the decoding parameter set information, and indicates that no further supplementary enhancement message is nested within the extensible nested supplementary enhancement message.

[0237] In an embodiment, if the nesting type flag (nesting_type) exhibits a second value, this indicates that the one or more supplemental enhancement messages nested within the extensible nesting supplemental enhancement message apply to an output layer set.

[0238] According to an embodiment, if the nesting type flag (nesting_type) exhibits a third value, this indicates that the one or more supplementary enhancement messages nested within the extensible nesting supplementary enhancement message apply to a subset of an output layer set.

[0239] In an embodiment, if the nesting type flag (nesting_type) exhibits a fourth value, this indicates that the extensible set of supplementary enhancement messages may, for example, include one or more decoding parameter sets, and the one or more decoding parameter sets include the decoding parameter set information.

[0240] Furthermore, according to an embodiment, a video data stream having video encoded therein is provided. The video data stream may, for example, include scalable nested supplemental enhancement information including decoding parameter set information for each of a plurality of output layer sets.

[0241] In an embodiment, the video data stream may, for example, comprise a raw byte sequence payload.The raw byte sequence payload may, for example, comprise the decoding parameter set information for each output layer set of the plurality of output layer sets.

[0242] According to an embodiment, the video data stream may, for example, include an extensible nested supplementary enhancement message, and the extensible nested supplementary enhancement message may, for example, include the extensible nested supplementary enhancement information, and the extensible nested supplementary enhancement information may, for example, include the decoding parameter set information for each output layer set in the multiple output layer sets.

[0243] In an embodiment, the video data stream may include, for example, a nesting type flag (nesting_type), wherein the nesting type flag (nesting_type) includes one or more bits.

[0244] According to an embodiment, if the nesting type flag (nesting_type) exhibits a first value, this indicates that the extensible nested supplementary enhancement message may, for example, include one or more decoding parameter sets, the one or more decoding parameter sets including the decoding parameter set information, and indicates that no further supplementary enhancement message is nested within the extensible nested supplementary enhancement message.

[0245] In an embodiment, if the nesting type flag (nesting_type) exhibits a second value, this indicates that the one or more supplemental enhancement messages nested within the extensible nesting supplemental enhancement message apply to an output layer set.

[0246] According to an embodiment, if the nesting type flag (nesting_type) exhibits a third value, this indicates that one or more supplementary enhancement messages nested in the scalable nested supplementary enhancement message are applied to a subset of the output layer set, for example, wherein one or more sub-pictures are allocated to the subset.

[0247] In an embodiment, if the nesting type flag (nesting_type) exhibits a fourth value, this indicates that the extensible set of supplementary enhancement messages may, for example, include one or more decoding parameter sets, and the one or more decoding parameter sets include the decoding parameter set information.

[0248] Furthermore, according to an embodiment, a video encoder for encoding a video into a video data stream is provided, such that the video data stream has the video encoded therein. The video encoder is configured to generate the video data stream such that the video data stream includes scalable nested supplemental enhancement information, the scalable nested supplemental enhancement information including decoding parameter set information for each output layer set in a plurality of output layer sets.

[0249] In an embodiment, the video encoder may be configured to generate the video data stream such that the video data stream may include the raw byte sequence payload. The video encoder may be configured to generate the video data stream such that the raw byte sequence payload may include the decoding parameter set information for each output layer set of the plurality of output layer sets.

[0250] According to an embodiment, the video encoder may be configured to generate the video data stream, for example, so that the video data stream may include an extensible nested supplementary enhancement message, the extensible nested supplementary enhancement message may include the extensible nested supplementary enhancement information, and the extensible nested supplementary enhancement information may include the decoding parameter set information for each output layer set in the multiple output layer sets.

[0251] In an embodiment, the video encoder may, for example, be configured to generate a video data stream such that the video data stream comprises a nesting type flag (nesting_type), which may, for example, comprise one or more bits.

[0252] According to an embodiment, the video encoder can be configured, for example, to generate the video data stream so that if the nesting type flag (nesting_type) exhibits a first value, this indicates that the extensible nested supplementary enhancement message includes one or more decoding parameter sets, which can, for example, include the decoding parameter set information and indicate that no further supplementary enhancement message is nested within the extensible nested supplementary enhancement message.

[0253] In an embodiment, the video encoder may be configured to generate the video data stream such that if the nesting type flag (nesting_type) exhibits a second value, this indicates that one or more supplementary enhancement messages nested within the scalable nested supplementary enhancement message apply to an output layer set.

[0254] According to an embodiment, the video encoder can, for example, be configured to generate the video data stream such that if the nesting type flag (nesting_type) exhibits a third value, this indicates one or more supplementary enhancement messages nested within the extensible nested supplementary enhancement message to a subset of the output layer set, for example, wherein one or more sub-pictures are assigned to the subset.

[0255] In an embodiment, the video encoder may, for example, generate the video data stream such that if the nesting type flag (nesting_type) exhibits a fourth value, this indicates that the scalable nesting supplemental enhancement message may, for example, include one or more decoding parameter sets, and the one or more decoding parameter sets include the decoding parameter set information.

[0256] Furthermore, according to an embodiment, a video data stream having video encoded therein is provided. The video data stream includes an indication of a plurality of operation points. Furthermore, the video data stream includes a first mapping for each of the plurality of operation points, the first mapping assigning one or more of a plurality of profile levels to the operation point. Furthermore, the video data stream includes a plurality of second mappings, wherein each of the plurality of mappings assigns one of the plurality of operation points to one of a plurality of output level sets.

[0257] In an embodiment, the video data stream may, for example, include a decoding parameter set. The decoding parameter set may, for example, include an indication of a plurality of operating points.

[0258] According to an embodiment, the video data stream may, for example, comprise a number of layer indications indicating the number of layers of the video data stream; and / or the video data stream may, for example, comprise a number of sub-layer indications indicating the number of sub-layers of the video data stream.

[0259] In an embodiment, the number of layers of the video data stream may, for example, be constant; and / or the number of sub-layers of the video data stream may, for example, be constant.

[0260] According to an embodiment, the video data stream may, for example, comprise scalable nested supplementary enhancement information comprising decoding parameter set information for each of a plurality of output layer sets.

[0261] Furthermore, according to an embodiment, a video encoder for encoding a video into a video data stream is provided, such that the video data stream has the video encoded therein. The video encoder is configured to generate the video data stream such that the video data stream includes an indication of a plurality of operation points. Furthermore, the video encoder is configured to generate the video data stream such that the video data stream includes a first mapping for each of the plurality of operation points, the first mapping assigning one or more of a plurality of profile levels to the operation point. Furthermore, the video encoder is configured to generate the video data stream such that the video data stream includes a plurality of second mappings, wherein each of the plurality of mappings assigns one of the plurality of operation points to one of a plurality of output layer sets.

[0262] In an embodiment, the video encoder may be configured to generate the video data stream such that the video data stream may include a decoding parameter set. The video encoder may be configured to generate the video data stream such that the decoding parameter set may include the indication of the plurality of operation points.

[0263] According to an embodiment, the video encoder may, for example, be configured to generate the video data stream such that the video data stream may, for example, include a number of layer indications indicating the number of layers of the video data stream; and / or the video encoder may, for example, be configured to generate the video data stream such that the video data stream may, for example, include a number of sub-layer indications indicating the number of sub-layers of the video data stream.

[0264] In an embodiment, the video encoder may, for example, be configured to generate the video data stream such that the number of layers of the video data stream may, for example, be constant; and / or the video encoder may, for example, be configured to generate the video data stream such that the number of sub-layers of the video data stream may, for example, be constant.

[0265] According to an embodiment, the video encoder may, for example, be configured to generate the video data stream such that the video data stream may, for example, include scalable nested supplementary enhancement information including decoding parameter set information for each output layer set of a plurality of output layer sets.

[0266] Furthermore, according to an embodiment, an apparatus is provided for receiving a video data stream as an input bitstream, wherein the video data stream has video encoded therein. The video data stream includes an indication of a plurality of operation points. Furthermore, the video data stream includes a first mapping for each of the plurality of operation points, the first mapping mapping one or more profile levels of a plurality of profile levels to the operation point. Furthermore, the video data stream includes a plurality of second mappings, wherein each of the plurality of mappings maps one of the plurality of operation points to an output level in a plurality of output level sets. The apparatus is configured to process the input bitstream in dependence on an output layer set in a plurality of output layer sets and in dependence on at least one of the second mappings mapping one of the plurality of operation points in the output level set to produce an output bitstream.

[0267] In an embodiment, the video data stream may include, for example, a decoding parameter set. The decoding parameter set may include, for example, an indication of a plurality of operating points. The apparatus is configured to process the decoding parameter set.

[0268] According to an embodiment, the video data stream may, for example, comprise a number of layer indications indicating the number of layers of the video data stream; and / or the video data stream may, for example, comprise a number of sub-layer indications indicating the number of sub-layers of the video data stream.

[0269] In an embodiment, the number of layers of the video data stream may, for example, be constant; and / or the number of sub-layers of the video data stream may, for example, be constant.

[0270] According to an embodiment, the video data stream may, for example, include scalable nested supplemental enhancement information, the scalable nested supplemental enhancement information including decoding parameter set information for each output layer set in a plurality of output layer sets. The apparatus may, for example, be configured to process the scalable nested supplemental enhancement information.

[0271] According to an embodiment, the apparatus may, for example, be configured to decode the input bitstream to decode the video.

[0272] Furthermore, according to an embodiment, a system for encoding a video into a video data stream and decoding the video is provided. The system includes a video encoder as described above and a decoding device as described above. The video encoder is configured to encode the video into the video data stream so that the video data stream has the video encoded therein. The device is configured to receive the video data stream as an input bitstream. Furthermore, the device is configured to decode the input bitstream to decode the video.

[0273] The VVC draft specification includes a decoding parameter set (DPS) that describes the constraints (PTL) of the bitstream. It can be used for capability exchange and negotiation, such as SDP in RTSP / SIP, or for stream selection in adaptive streaming scenarios. Its purpose is to indicate the maximum capabilities required for a given bitstream (a concatenation of CVSs), and its scope is therefore greater than all parameter sets defined in the corresponding predecessor codec generations (HEVC and AVC), which only have the scope of coded video sequences (CVSs). The VVC draft specification now describes the constraints of the concatenation of CVSSs, i.e., the bitstream. The DPS primarily carries multiple profile level information (PTLs) including all PTLs used in the bitstream. That is, all PTLs in this list need to be supported in order to successfully decode the corresponding bitstream.

[0274]

[0275] When performing bitstream extraction or pruning, the problem arises that some of these layers are removed. After such processing, the resulting bitstream may no longer carry the entire set of layers, for example, extracting 3 layers containing versions (corresponding to OLS) from a 6-layer bitstream. Therefore, in the case where the DPS cannot accurately describe the PTL of the content of the resulting pruned bitstream, the DPS will no longer be able to serve its purpose of capability negotiation.

[0276] Therefore, as part of this invention, three options are described how to alleviate this problem and allow capability negotiation using DPS.

[0277] In the following, DPS rewriting during the extraction process is described.

[0278] The PTL information in the DPS is adjusted during extraction, in that only the PTL information corresponding to the remaining layers after extraction remains in the DPS, while the other PTL information is removed.

[0279] In the following, sub-bitstream specific DPS nesting is described.

[0280] A scalable nested SEI message is defined that carries the OLS-specific DPS RBSP and extracts the corresponding DPS RBSP via the bitstream extraction process and writes it into a new DPS in the output bitstream.

[0281] In an embodiment, a new nested SEI message carrying parameter sets is used.

[0282] In another embodiment, a single scalable nested SEI message is used to carry parameter sets or other SEI messages. In this example, syntax elements are added to the scalable nested SEI message to indicate the content of the internal execution.

[0283]

[0284] In the above syntax, it is assumed that when a parameter set is included (nesting_type==1), there is no other SEI. In another embodiment, parameter sets and other SEIs can be included in the nesting SEI message at the same time.

[0285] In another embodiment, nesting_type can be used for several purposes. Indicates:

[0286] SEI messages nested within nested SEI are applied to OLS

[0287] SEI messages nested within nested SEIs apply to a subset of the OLS, e.g., a sub-picture within the OLS

[0288] Nesting existing parameter sets within nested SEIs

[0289] In the following, the operating points in DPS are introduced.

[0290] DPS rewriting during ingestion and sub-bitstream-specific DPS nesting may work in certain situations (e.g., constrained bitstreams). For example, DPS rewriting during ingestion may only work if all VPSs of the bitstream are known in advance and the number of layers / sub-layers does not change from CVS to CVS. A similar situation occurs with sub-bitstream-specific DPS nesting, since, in principle, OLSs (e.g., their number and IDs) may change from CVS to CVS.

[0291] In another embodiment, operating points are defined. They may be defined in the DPS, with a list provided for each PTL. In addition to the OLS index, an operating point index is added as an input to the decoding process to describe and select the appropriate operating point. Some association of the operating point with the OLS is then signaled in the VPS.

[0292] The syntax elements of the DPS will be changed as follows by removing the number of sub-layers, as this may change from CVS to CVS and including the operating points.

[0293]

[0294] In another embodiment, num_sub_layer and the number of layers remain constant in the bitstream.This indication is added to the bitstream and indicates the mapping between operation points and sub-layers or multiple layers.

[0295] In another embodiment, the scalable nested SEI message defined in 4.2 is applied to the operation point, and the mapping to the OLS is also provided in the bitstream.

[0296] Hereinafter, MaxSubLayer in the Param set is described.

[0297] According to an embodiment, a video data stream having video encoded therein is provided. The video data stream includes a sequence parameter set and a decoding parameter set. The sequence parameter set includes a first indicator value that indicates a maximum number of sublayers or a constant value less than the maximum number of sublayers based on information stored in the sequence parameter set. The decoding parameter set includes a second indicator value that indicates the maximum number of sublayers or the constant value less than the maximum number of sublayers based on information stored in the decoding parameter set. The first indicator value is less than or equal to the second indicator value.

[0298] In an embodiment, a video data stream may, for example, include a sequence parameter set and a decoding parameter set. The sequence parameter set may, for example, include a first indicator value, the first indicator value indicating the maximum number of sublayers or the maximum number of sublayers minus a constant value based on information stored in the sequence parameter set. The decoding parameter set may, for example, include a second indicator value, the second indicator value indicating the maximum number of sublayers or the maximum number of sublayers minus the constant value based on information stored in the decoding parameter set. The second indicator value of the decoding parameter set is an upper bound that takes precedence over the first indicator value of the sequence parameter set.

[0299] According to an embodiment, the constant value may be 1, for example.

[0300] Furthermore, according to an embodiment, a video encoder for encoding video into a video data stream is provided, such that the video data stream has the video encoded therein. The video encoder is configured to generate the video data stream such that the video data stream includes a sequence parameter set and a decoding parameter set. The video encoder is configured to generate the video data stream such that the sequence parameter set includes a first indicator value, the first indicator value indicating a maximum number of sublayers or the maximum number of sublayers minus a constant value based on information stored in the sequence parameter set. Furthermore, the video encoder generates the video data stream such that the decoding parameter set includes a second indicator value indicating the maximum number of sublayers or the maximum number of sublayers minus the constant value based on information stored in the decoding parameter set. Furthermore, the video encoder is configured to generate the video data stream such that the first indicator value is less than or equal to the second indicator value.

[0301] Furthermore, according to an embodiment, a video encoder is provided for encoding video into a video data stream, such that the video data stream has the video encoded therein. The video encoder is configured to generate the video data stream such that the video data stream includes a sequence parameter set and a decoding parameter set. Furthermore, the video encoder is configured to generate the video data stream such that the sequence parameter set includes a first indicator value, the first indicator value indicating a maximum number of sublayers or the maximum number of sublayers minus a constant value based on information stored in the sequence parameter set. Furthermore, the video encoder generates the video data stream such that the decoding parameter set includes a second indicator value indicating the maximum number of sublayers or the maximum number of sublayers minus the constant value based on information stored in the decoding parameter set. The second indicator value of the decoding parameter set is an upper bound that takes precedence over the first indicator value of the sequence parameter set.

[0302] In an embodiment, the constant value may be, for example, 1. The video encoder may be configured to generate the video data stream such that the sequence parameter set may include, for example, the first indication value, which indicates the maximum number of sub-layers or the maximum number of sub-layers minus 1 according to information stored in the sequence parameter set. The video encoder may be configured to generate the video data stream such that the decoding parameter set may include, for example, the second indication value, which indicates the maximum number of sub-layers or the maximum number of sub-layers minus 1 according to information stored in the decoding parameter set.

[0303] Furthermore, according to an embodiment, an apparatus is provided for receiving a video data stream as an input bitstream, wherein the video data stream has video encoded therein. The video data stream includes a sequence parameter set and a decoding parameter set. The sequence parameter set includes a first indicator value, the first indicator value indicating a maximum number of sublayers or the maximum number of sublayers minus a constant value based on information stored in the sequence parameter set. The decoding parameter set includes a second indicator value, the second indicator value indicating the maximum number of sublayers or the maximum number of sublayers minus the constant value based on information stored in the decoding parameter set. If the first indicator value is greater than the second indicator value, the apparatus is configured to process the input bitstream to generate an output bitstream, such that the output bitstream includes the second indicator value as an indication of the maximum number of sublayers or the maximum number of sublayers minus the constant value.

[0304] In an embodiment, the constant value may be, for example, 1. The sequence parameter set may, for example, include the first indicator value, which indicates the maximum number of sublayers or the maximum number of sublayers minus 1 according to information stored in the sequence parameter set, and the decoding parameter set may, for example, include the second indicator value, which indicates the maximum number of sublayers or the maximum number of sublayers minus 1 according to information stored in the decoding parameter set. If the first indicator value may, for example, be greater than the second indicator value, the apparatus may, for example, be configured to process the input bitstream to generate an output bitstream such that the output bitstream may, for example, include the second indicator value as the maximum number of sublayers minus 1.

[0305] When doing bitstream pruning (also known as sub-layer extraction) and discarding NAL units belonging to a specific temporal ID from the bitstream, the parameter sets are generally not modified, although some high-level parameter sets such as DPS may be overwritten.

[0306] This leads to inconsistencies within the bitstream, making the decoding process unclear or more complicated. For example, if the DPS is rewritten as described in the previous section, and the actual number of sublayers is indicated as dps_max_sublayers_minus1 (excluding the dropped temporal ID), but the value at the SPS is not modified (which is not intended), sps_max_sublayers_minus1 will have a higher value than the value on the DPS. The decoder will not really know what value to trust.

[0307] In one embodiment, the bitstream constraint is that sps_max_sublayers_minus1 must be <= dps_max_sublayers_minus1. Therefore, if such a "problem" is encountered, some type of error resilience can be built in. With such a constraint, it is clear that when parsing a "bitstream" that violates this condition, the SPS with a higher value will simply correspond to the new bitstream (for example, due to a missing end-of-bitstream (EOB) NAL unit). This will mean that when performing bitstream pruning, dps_max_sublayers_minus1 cannot be changed unless sps_max_sublayers_minus1 is also changed.

[0308] In another embodiment, this constraint is not required and DPS signaling is used to set the operating point or output layer set of the bitstream sent to the decoder. That is, when dps_max_sublayers_minus1 is less than sps_max_sublayers_minus1, the value signaled in the DPS takes precedence over the value in the SPS indicating that the bitstream sent to the decoder has at most dps_max_sublayers_minus1.

[0309] Since different layers may have different numbers of sublayers when more than one layer is present, dps_max_sublayers_minus1 applies to all output layers, while all non-output layers have dps_max_sublayers_minus1 or its corresponding sps_max_sublayers_minus1 (if the latter is smaller than dps_max_sublayers_minus1).

[0310] Although some aspects have been described in the context of an apparatus, it is clear that these aspects also represent descriptions of corresponding methods, where blocks or devices correspond to method steps or features of method steps. Similarly, aspects described in the context of method steps also represent descriptions of corresponding blocks, items, or features of the corresponding apparatus. Some or all of the method steps can be performed by (or using) hardware devices such as microprocessors, programmable computers, or electronic circuits. In some embodiments, one or more of the most important method steps can be performed by such devices.

[0311] Depending on certain implementation requirements, embodiments of the present invention can be implemented in hardware or software, or at least partially in hardware or at least partially in software. Implementations can be performed using a digital storage medium (e.g., a floppy disk, DVD, Blu-ray, CD, ROM, PROM, EPROM, EEPROM, or flash memory) having electronically readable control signals stored thereon that cooperate (or are capable of cooperating) with a programmable computer system to cause the corresponding method to be performed. Thus, the digital storage medium can be computer-readable.

[0312] Some embodiments according to the invention comprise a data carrier having electronically readable control signals, which are capable of cooperating with a programmable computer system, such that one of the methods described herein is performed.

[0313] Generally speaking, embodiments of the present invention can be implemented as a computer program product with a program code, the program code being operative to perform one of the methods when the computer program product runs on a computer. The program code may, for example, be stored on a machine-readable carrier.

[0314] Other embodiments comprise the computer program for performing one of the methods described herein, stored on a machine readable carrier.

[0315] In other words, an embodiment of the inventive method is, therefore, a computer program having a program code for performing one of the methods described herein, when the computer program runs on a computer.

[0316] A further embodiment of the inventive method is, therefore, a data carrier (or a digital storage medium or a computer-readable medium) comprising, recorded thereon, the computer program for carrying out one of the methods described herein. The data carrier, the digital storage medium or the recorded medium are typically tangible and / or non-transitory.

[0317] A further embodiment of the inventive method is, therefore, a data stream or a sequence of signals representing the computer program for carrying out one of the methods described herein. The data stream or the sequence of signals can, for example, be configured to be transferred via a data communication connection (for example, via the Internet).

[0318] Further embodiments comprise a processing means, for example a computer or a programmable logic device, configured to or adapted to carry out one of the methods described herein.

[0319] A further embodiment comprises a computer having installed thereon the computer program for performing one of the methods described herein.

[0320] Another embodiment according to the present invention comprises an apparatus or system configured to transfer (e.g., electronically or optically) a computer program for carrying out one of the methods described herein to a receiver. The receiver may be, for example, a computer, a mobile device, a memory device, etc. The apparatus or system may, for example, comprise a file server for transferring the computer program to the receiver.

[0321] In some embodiments, a programmable logic device (e.g., a field programmable gate array) can be used to implement some or all of the functionality of the methods described herein. In some embodiments, the field programmable gate array can cooperate with a microprocessor to implement one of the methods described herein. In general, the methods are preferably implemented by any hardware device.

[0322] The devices described herein may be implemented using a hardware device or using a computer or using a combination of a hardware device and a computer.

[0323] The methods described herein may be implemented using a hardware device or using a computer or using a combination of a hardware device and a computer.

[0324] The above embodiments are merely illustrative of the principles of the present invention. It is understood that modifications and variations of the arrangements and details described herein will be readily apparent to those skilled in the art. It is therefore intended that the present invention be limited only by the scope of the appended claims and not by the specific details presented by way of description and explanation of the embodiments herein.

[0325] References

[0326] [1] ISO / IEC, ITU-T. High-Efficiency Video Coding. ITU-T Recommendation H.265 | ISO / IEC 2300810 (HEVC), Version 1, 2013; Version 2, 2014.

Claims

1. A video encoder configured to: A plurality of access units is encoded into a video data stream, the plurality of access units comprising: the first access unit, and a second access unit; and encoding a delay_for_concatenation_ensured_flag corresponding to the first access unit into the video data stream, In response to a value of delay_for_concatenation_ensured_flag being 1, determining that the first access unit is located at a position in the video data stream such that splicing of the first access unit with the second access unit having a buffering period supplemental enhancement information (BP) SEI message having the following characteristics will be seamless: concatenation_flag is equal to 1, and The selected InitCpbRemovalDelay is less than or equal to the value of max_initial_removal_delay_for_concatination.

2. The video encoder according to claim 1, wherein Seamless splicing indication: the coded picture buffer CPB removal time distance between the first access unit and the second access unit associated with the delay_for_concatenation_ensured_flag with a value of 1 is the same as the removal time distance between consecutive access units in the first access unit.

3. The video encoder according to claim 1 , further configured to: Encode the BP SEI message of the specified temporal sub-layer into the video data stream, wherein, The delay_for_concatenation_ensured_flag is encoded only for the access unit associated with the temporal sub-layer.

4. The video encoder according to claim 1, further configured to: A BP SEI message specifying the first access unit is encoded into the video data stream, wherein: Encoding the delay_for_concatenation_ensured_flag is performed only for the first access unit.

5. The video encoder according to claim 4, wherein The BP SEI message specifying the first access unit includes a signaled divisor, and the first access unit includes all of the following access units: has a cpb_removal_delay_minus1 that leaves a remainder of zero when divided by the signaled divisor, or Having a dpb_output_delay that leaves a remainder of zero when divided by the signaled divisor. The video encoder according to claim 4 , wherein: The BP SEI message specifying the first access unit includes a signaled divisor and a signaled remainder, and the first access unit includes all of the following access units: having a cpb_removal_delay_minus1 that, when divided by the signaled divisor, leaves a remainder equal to the signaled remainder, or Having a dpb_output_delay, when the dpb_output_delay is divided by the signaled divisor, a remainder equal to the signaled remainder.

7. A video decoder, configured to: Decoding a plurality of access units from a video data stream, the plurality of access units comprising: the first access unit, and Second access unit; as well as Decoding a delay_for_concatenation_ensured_flag corresponding to the first access unit from the video data stream, In response to a value of delay_for_concatenation_ensured_flag being 1, determining that the first access unit is located at a position in the video data stream such that splicing of the first access unit with the second access unit having a buffering period supplemental enhancement information (BP) SEI message having the following characteristics will be seamless: concatenation_flag is equal to 1, and The selected InitCpbRemovalDelay is less than or equal to the value of max_initial_removal_delay_for_concatination.

8. The video decoder according to claim 7, wherein: Seamless splicing indication: the coded picture buffer CPB removal time distance between the first access unit and the second access unit associated with the delay_for_concatenation_ensured_flag with a value of 1 is the same as the removal time distance between consecutive access units in the first access unit.

9. The video decoder according to claim 7, further configured to: A BP SEI message specifying a temporal sub-layer is decoded from the video data stream, wherein: Decoding the delay_for_concatenation_ensured_flag is performed only for the access unit associated with the temporal sub-layer.

10. The video decoder according to claim 7, further configured to: A BP SEI message specifying the first access unit is decoded from the video data stream, wherein: Decoding the plurality of delay_for_concatenation_ensured_flags is performed only on access units in the first access unit. The video decoder according to claim 10 , wherein: The BP SEI message specifying the first access unit includes a signaled divisor, and the first access unit includes all of the following access units: has a cpb_removal_delay_minus1 that leaves a remainder of zero when divided by the signaled divisor, or Having a dpb_output_delay that leaves a remainder of zero when divided by the signaled divisor.

12. The video decoder according to claim 10, wherein: The BP SEI message specifying the first access unit includes a signaled divisor and a signaled remainder, and the first access unit includes all of the following access units: having a cpb_removal_delay_minus1 that, when divided by the signaled divisor, leaves a remainder equal to the signaled remainder, or Having a dpb_output_delay, when the dpb_output_delay is divided by the signaled divisor, a remainder equal to the signaled remainder.

13. A video decoding method, comprising: Decoding a plurality of access units from a video data stream, the plurality of access units comprising: the first access unit, and a second access unit; and Decoding a delay_for_concatenation_ensured_flag corresponding to the first access unit from the video data stream, In response to the value of the delay_for_concatenation_ensured_flag being 1, determining that the first access unit is located at a position in the video data stream, whereby splicing of the first access unit with the second access unit having a buffering period supplemental enhancement information (BP) SEI message having the following characteristics is seamless: concatenation_flag is equal to 1, and The selected InitCpbRemovalDelay is less than or equal to the value of max_initial_removal_delay_for_concatination.

14. The video decoding method according to claim 13, wherein: Seamless splicing indication: the coded picture buffer CPB removal time distance between the first access unit and the second access unit associated with the delay_for_concatenation_ensured_flag with a value of 1 is the same as the removal time distance between consecutive access units in the first access unit.

15. The video decoding method according to claim 13, further comprising: A BP SEI message specifying a temporal sub-layer is decoded from the video data stream, wherein decoding the delay_for_concatenation_ensured_flag is performed only for access units associated with the temporal sub-layer.

16. The video decoding method according to claim 13, further comprising: A BP SEI message specifying the first access unit is decoded from the video data stream, wherein decoding of the delay_for_concatenation_ensured_flag is performed only for access units in the first access units.

17. The video decoding method according to claim 16, wherein: The BP SEI message specifying the first access unit includes a signaled divisor, and the first access unit includes all of the following access units: has a cpb_removal_delay_minus1 that leaves a remainder of zero when divided by the signaled divisor, or Having a dpb_output_delay that leaves a remainder of zero when divided by the signaled divisor.

18. The video decoding method of claim 16 , wherein the BP SEI message specifying the first access unit includes a signaled divisor and a signaled remainder, and the first access unit includes all of the following access units: having a cpb_removal_delay_minus1 that, when divided by the signaled divisor, leaves a remainder equal to the signaled remainder, or Having a dpb_output_delay, when the dpb_output_delay is divided by the signaled divisor, a remainder equal to the signaled remainder.

19. A method for encoding a video, the method comprising: A plurality of access units is encoded into a video data stream, the plurality of access units comprising: the first access unit, and a second access unit; and encoding a delay_for_concatenation_ensured_flag corresponding to the first access unit into the video data stream, In response to a value of delay_for_concatenation_ensured_flag being 1, determining that the first access unit is located at a position in the video data stream such that splicing of the first access unit with the second access unit having a buffering period supplemental enhancement information (BP) SEI message having the following characteristics will be seamless: concatenation_flag is equal to 1, and The selected InitCpbRemovalDelay is less than or equal to the value of max_initial_removal_delay_for_concatination.

Citation Information

Patent Citations

  • Low-delay buffering model in video coding

    CN104854870A

  • Image encoding device and method and image processing device and method

    CN106063275A