Video data stream, encoder, method of encoding video content and decoder

By introducing CPB operations at the sub-image level and optimizing the buffer period SEI message in the HEVC standard, the inefficiency caused by the frequent occurrence of the buffer period SEI message is solved, achieving efficient decoding and transmission efficiency of low-latency video coding, and supporting compatibility with multiple transmission protocols.

CN115442631BActive Publication Date: 2026-04-10DOLBY VIDEO COMPRESSION LLC
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
DOLBY VIDEO COMPRESSION LLC
Filing Date
2013-07-01
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

In the existing HEVC standard for low-latency video coding, the frequent occurrence of buffer period SEI messages may lead to low decoding efficiency, and the transmission protocol of high-order syntax signals is inconsistent between different application layer codecs, affecting the effectiveness and efficiency of video data streams.

Method used

By introducing CPB operations and buffer period SEI messages at the sub-image level, the frequency of use of buffer period SEI messages is optimized, and high-order syntax signaling is performed at the decoding unit level to ensure more efficient timing and buffer management of the decoding unit and support compatibility with multiple transmission protocols.

Benefits of technology

It achieves efficient decoding of low-latency video encoding, improves the transmission efficiency of video data streams and the compatibility of decoders, and meets the needs of different application layers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115442631B_ABST
    Figure CN115442631B_ABST
Patent Text Reader

Abstract

The invention relates to a video data stream, an encoder, a method of encoding video content and a decoder. The video data stream has video content encoded therein in units of sub-portions of pictures of the video content, each sub-portion is encoded into one or more payload packets of a sequence of packets of the video data stream, the sequence of packets is divided into a sequence of access units, thus each access unit collects payload packets relating to respective pictures of the video content, wherein the sequence of packets has timing control packets interspersed therein, thus the timing control packets subdivide the access units into decoding units, thus at least some of the access units are subdivided into two or more decoding units, wherein each timing control packet signals a decoder buffer retrieval time for a decoding unit, the payload packets of the decoding unit follow respective timing control packets in the sequence of packets.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of the application for invention entered into the national phase on December 29, 2014 for international application number PCT / EP2013 / 063853 with an international filing date of July 1, 2013, application number 201380034944.4, with the title "Video data stream concept technology", the entire content of which is incorporated herein by reference. TECHNICAL FIELD

[0002] The present application relates to video data stream concept technology, in particular to the concept technology which is advantageous in low-delay applications. BACKGROUND

[0003] HEVC [2] allows different means for high level syntax signaling to the application layer. These means are the NAL unit header, the parameter sets and the supplemental enhancement information (SEI) messages. SEI messages are not used in the decoding process. Other means for high level syntax signaling originate from individual transport protocol specifications, such as the MPEG2 transport protocol [3] or the real-time transport protocol [4], and their payload specific specifications, e.g. for H.264 / AVC [5], scalable video coding (SVC) [6] or HEVC [7]. These transport protocols can introduce high level signaling, which uses structures and mechanisms similar to the high level signaling of individual application layer codec specifications, e.g. HEVC [2]. An example for such signaling is the payload content scalability information (PACSI) NAL unit as described in [6], which provides supplemental information for the transport layer.

[0004] HEVC [2] allows different means for high level syntax signaling to the application layer. These means are the NAL unit header, the parameter sets and the supplemental enhancement information (SEI) messages. SEI messages are not used in the decoding process. Other means for high level syntax signaling originate from individual transport protocol specifications, such as the MPEG2 transport protocol [3] or the real-time transport protocol [4], and their payload specific specifications, e.g. for H.264 / AVC [5], scalable video coding (SVC) [6] or HEVC [7]. These transport protocols can introduce high level signaling, which uses structures and mechanisms similar to the high level signaling of individual application layer codec specifications, e.g. HEVC [2]. An example for such signaling is the payload content scalability information (PACSI) NAL unit as described in [6], which provides supplemental information for the transport layer.

[0005] In terms of parameter sets, HEVC includes a video parameter set (VPS) that compiles the most important stream information to be used by the application layer in a single and central location. In earlier approaches, this information needs to be gathered from multiple parameter sets and NAL unit headers.

[0006] Prior to the present application, the status of the standard for coded picture buffer (CPB) operation with respect to the hypothetical reference decoder (HRD) and all related syntax provided in the sequence parameter set (SPS) / video usability information (VUI), picture timing SEI, buffering period SEI, and the definition of a decoding unit (which describes sub-pictures and dependent slices as presented in the slice header and the picture parameter set (PPS)) is as follows.

[0007] To allow low-delay CPB operation on the sub-picture level, sub-picture CPB operation has been proposed and incorporated into the HEVC standard draft 7 JCTVC-I1003 [2]. In particular, in [2], a decoding unit has been defined in section 3 as follows:

[0008] Decoding unit: an access unit or a subset of an access unit. If SubPicCpbFlag is equal to 0, the decoding unit is the access unit. Otherwise, the decoding unit consists of one or more VCL NAL units and associated non-VCL NAL units from an access unit. For the first VCL NAL unit in an access unit, the associated non-VCL NAL units are the filler data NAL units (if any) immediately following the first VCL NAL unit and all non-VCL NAL units in that access unit that precede the first VCL NAL unit. For a VCL NAL unit that is not the first VCL NAL unit in an access unit, the associated non-VCL NAL units are the filler data NAL units (if any) immediately following the VCL NAL unit.

[0009] In the standard defined up to that point, "decoding unit removal timing and decoding unit decoding" has been described. To signal sub-picture timing, the buffering period SEI message and the picture timing SEI message and the HRD parameters in the VUI have been extended to support decoding units, such as sub-picture units.

[0010] The buffering period SEI message syntax of [2] is shown in Figure 1 .

[0011] When NalHrdBpPresentFlag or VclHrdBpPresentFlag is equal to 1, the buffering period SEI message can be related to any access unit in the bitstream, and the buffering period SEI message will be related to each RAP access unit and each access unit associated with a recovery point SEI message.

[0012] For some applications, frequent occurrence of buffering period SEI messages can be desirable.

[0013] The buffering period specifies a set of access units in decoding order between two occurrences of buffering period SEI messages.

[0014] The semantics are as follows:

[0015] seq_parameter_set_id specifies the sequence parameter set containing the sequence HRD properties. The value of seq_parameter_set_id shall be equal to the value of seq_parameter_set_id in the picture parameter set referred to by the primary coded picture associated with the buffering period SEI message. The value of seq_parameter_set_id shall be in the range of 0 to 31, inclusive.

[0016] rap_cpb_params_present_flag equal to 1 specifies the presence of the initial_alt_cpb_removal_delay[SchedSelIdx] and initial_alt_cpb_removal_delay_offset[SchedSelIdx] syntax elements. When not present, the value of rap_cpb_params_present_flag is inferred to be equal to 0. The value of rap_cpb_params_present_flag shall be equal to 0 when the associated picture is neither a CRA picture nor a BLA picture.

[0017] initial_cpb_removal_delay[SchedSelIdx] and initial_alt_cpb_removal_delay[SchedSelIdx] specify the initial CPB removal delay for the SchedSelIdx-th CPB. These syntax elements have a length, in bits, given by initial_cpb_removal_delay_length_minus1 + 1 and are in units of 90 kHz clock. The values of these syntax elements shall not be equal to 0 and shall not exceed 90000 * (CpbSize[SchedSelIdx] ÷ BitRate[SchedSelIdx]), which is the time equivalent of the CPB size in units of 90 kHz clock.

[0018] initial_cpb_removal_delay_offset[SchedSelIdx] and initial_alt_cpb_removal_delay_offset[SchedSelIdx] are used for the SchedSelIdx-th CPB to specify the initial delivery time of coded data units to that CPB. These syntax elements have a length in bits given by initial_cpb_removal_delay_length_minus1 + 1 and are in units of 90 kHz clock ticks. These syntax elements are not used by decoders and are only needed by the delivery scheduler (HSS).

[0019] The sum of initial_cpb_removal_delay[SchedSelIdx] and initial_cpb_removal_delay_offset[SchedSelIdx] shall be constant for each value of SchedSelIdx over the entire coded video sequence, and the sum of initial_alt_cpb_removal_delay[SchedSelIdx] and initial_alt_cpb_removal_delay_offset[SchedSelIdx] shall be constant for each value of SchedSelIdx over the entire coded video sequence.

[0020] The syntax of the picture timing SEI message of [2] is shown in Figure 2

[0021] The syntax of the picture timing SEI message depends on the content of the sequence parameter set that is active for the coded picture to which the picture timing SEI message relates. However, unless the picture timing SEI message of an IDR or BLA access unit is preceded by a buffering period SEI message located in the same access unit, the activation of the relevant sequence parameter set (and, for an IDR or BLA picture that is not the first picture in the bitstream, the determination that the coded picture is an IDR picture or a BLA picture) does not occur prior to the decoding of the first coded slice NAL unit of the coded picture. Since the coded slice NAL units of a coded picture follow the picture timing SEI message in NAL unit order, there can be a case where the decoder needs to store the RBSP containing the picture timing SEI message and then perform the parsing of the picture timing SEI message prior to determining the parameters of the sequence parameter set that will be active for the coded picture. This is not a problem for the case of a picture timing SEI message that is not associated with an IDR or BLA picture, since the picture timing SEI message is not required to be present in the bitstream. However, for the case of a picture timing SEI message that is associated with an IDR or BLA picture, the picture timing SEI message is required to be present in the bitstream, and the decoder needs to be able to determine the parameters of the sequence parameter set that will be active for the coded picture prior to decoding the first coded slice NAL unit of the coded picture.

[0022] The presence of a picture timing SEI message in a bitstream specifies the following.

[0023] ​- If CpbDpbDelaysPresentFlag is equal to 1, there shall be one picture timing SEI message present in each access unit of the coded video sequence.

[0024] - Otherwise (CpbDpbDelaysPresentFlag is equal to 0), there shall not be a picture timing SEI message present in any access unit of the coded video sequence.

[0025] The semantics are defined as follows:

[0026] cpb_removal_delay specifies how many clock ticks to wait after the removal of the cpb of the access unit related to the last buffering period SEI message in the previous access unit before removing the access unit data related to the picture timing SEI message from the buffer. This value is also used to calculate the earliest possible time of arrival of the access unit data in the CPB for HSS. This syntax element is a fixed length code whose length in bits is given by cpb_removal_delay_length_minus1 + 1. cpb_removal_delay is the remainder modulo 2 (cpb_removal_delay_length_minus1+1) counter.

[0027] The value of cpb_removal_delay_length_minus1, which determines the length (in bits) of the decision syntax element cpb_removal_delay, is the value of cpb_removal_delay_length_minus1 coded in the sequence parameter set that is active for the primary coded picture related to the picture timing SEI message, but cpb_removal_delay specifies the number of clock ticks related to the removal time of the previous access unit containing the buffering period SEI message.

[0028] dpb_output_delay is used to calculate the dpb output time of a picture. It specifies how many clock ticks to wait after the removal of the last decoding unit in the access unit from the CPB before outputting the decoded picture from the DPB.

[0029] A picture is removed from the DPB at its output time when it is still marked as "used for short-term reference" or "used for long-term reference".

[0030] For a decoded picture, only one dpb_output_delay is specified.

[0031] The length of the syntax element dpb output delay is given in bits by dpb output delay length minusl plus 1. When sps max dec pic buffering [max temporal layers minusl] is equal to 0, dpb output delay shall be equal to 0.

[0032] The output time derived from the dpb output delay of any picture output by the decoder that conforms to the output time shall precede in decoding order the output time derived from the dpb output delay of all pictures in any subsequent coded video sequence.

[0033] The picture output order determined by the value of this syntax element shall be the same as the order determined by the value of PicOrderCntVal.

[0034] For pictures that are not output by the "bumping" process because they precede in decoding order an IDR or BLA picture for which no_output_of_prior_pics_flag is equal to 1 or inferred to be equal to 1, the output time derived from dpb output delay shall increase with increasing PicOrderCntVal values with respect to all pictures within the same coded video sequence.

[0035] num_decoding_units_minusl plus 1 specifies the number of decoding units in the access unit to which the picture timing SEI message is associated. The value of num_decoding_units_minusl shall be in the range of 0 to PicWidthInCtbs * PicHeightInCtbs - 1, inclusive.

[0036] num_nalus_in_du_minusl [ i ] plus 1 specifies the number of NAL units in the i-th decoding unit of the access unit to which the picture timing SEI message is associated. The value of num_nalus_in_du_minusl [ i ] shall be in the range of 0 to PicWidthInCtbs * PicHeightInCtbs - 1, inclusive.

[0037] The first decoding unit of an access unit consists of the first num_nalus_in_du_minus[0] + 1 consecutive NAL units in decoding order in the access unit. The i-th (where i is greater than 0) decoding unit of an access unit consists of num_nalus_in_du_minus[i] + 1 consecutive NAL units immediately following the last NAL unit in the previous decoding unit of the access unit in decoding order. There will be at least one VCL NAL unit in each decoding unit. All non-VCL NAL units related to VCL NAL units shall be included in the same decoding unit.

[0038] du_cpb_removal_delay[i] specifies how many subpicture clock ticks to wait after the removal of the cpb of the first decoding unit of the access unit related to the most recent buffering period SEI message in the previous access unit before the removal of the i-th decoding unit in the access unit related to the picture timing SEI message from the CPB. This value is also used to calculate the earliest possible time of decoding unit data arrival at the CPB for the HSS. This syntax element is a fixed length code whose length in bits is given by cpb_removal_delay_length_minus1 + 1. du_cpb_removal_delay[i] is modulo 2 (cpb _removal_delay_length_minus1+1) of the counter.

[0039] The value of cpb_removal_delay_length_minus1 that determines the length (in bits) of the syntax element du_cpb_removal_delay[i] is the value of cpb_removal_delay_length_minus1 coded in the sequence parameter set that is active for the coded picture related to the picture timing SEI message, but du_cpb_removal_delay[i] specifies the number of subpicture clock ticks related to the removal time of the first decoding unit in the previous access unit that contains a buffering period SEI message.

[0040] The VUI syntax of [2] contains certain information. The VUI parameter syntax of [2] is shown in Figure 3A and Figure 3B The HRD parameter syntax of [2] is shown in Figure 4 The semantics are defined as follows:

[0041] sub_pic_cpb_params_present_flag equal to 1 specifies that sub-picture level CPB removal delay parameters are present and the CPB can operate on the access unit level or the sub-picture level. sub_pic_cpb_params_present_flag equal to 0 specifies that sub-picture level CPB removal delay parameters are not present and the CPB operates on the access unit level. When sub_pic_cpb_params_present_flag is not present, its value is inferred to be equal to 0.

[0042] num_units_in_sub_tick is the number of time units of a clock operating at the frequency time_scale Hz that corresponds to one increment of the subpicture clock tick counter, referred to as a subpicture clock tick. num_units_in_sub_tick shall be greater than 0. The subpicture clock tick is the smallest time interval that can be represented in the coded data when sub_pic_cpb_params_present_flag is equal to 1.

[0043] tiles_fixed_structure_flag equal to 1 specifies that each active picture parameter set in the coded video sequence has the same values for the following syntax elements: num_tile_columns_minus1, num_tile_rows_minus1, uniform_spacing_flag, column_width[ i ], row_height[ i ], and loop_filter_across_tiles_enabled_flag (when present). tiles_fixed_structure_flag equal to 0 specifies that the tile syntax elements in different picture parameter sets can or can not have the same values. When the tiles_fixed_structure_flag syntax element is not present, it is inferred to be equal to 0.

[0044] Signaling tiles_fixed_structure_flag equal to 1 is a guarantee to the decoder that each picture in the coded video sequence has the same number of tiles distributed in the same way, which can be useful for workload partitioning in multi-threaded decoding scenarios.

[0045] Using Figure 5 The filler data RBSP syntax shown in [2] signals the filler data of [2].

[0046] The hypothetical reference decoder of [2] that is used to check bitstream and decoder conformance is defined as follows:

[0047] Two types of bitstreams are subject to HRD conformance checking for this recommendation | International Standard. The first type of bitstream (called Type I bitstream) is a NAL unit stream containing only VCL NAL units and filler data NAL units for all access units in the bitstream. The second type of bitstream (called Type II bitstream) contains at least one of the following in addition to VCL NAL units and filler data NAL units for all access units in the bitstream:

[0048] - additional non-VCL NAL units other than filler data NAL units,

[0049] - all leading_zero_8bits, zero_byte, start_code_prefix_one_3bytes and trailing_zero_8bits syntax elements, which form the bitstream from the NAL unit stream.

[0050] Figure 6 Types of bitstream conformance points that are shown by the HRD check of [2].

[0051] Two types of HRD parameter sets (NAL HRD parameters and VCL HRD parameters) are used. The HRD parameter sets are signaled via the video usability information, which is part of the sequence parameter set syntax structure.

[0052] All sequence parameter sets and picture parameter sets referred to in VCL NAL units, as well as the corresponding buffering period and picture timing SEI messages, are conveyed to the HRD in a timely manner, either in the bitstream or by other means.

[0053] The requirement for "presence" of non-VCL NAL units is also met when they (or only some of them) are conveyed to the decoder (or to the HRD) by other means than their presence in the bitstream. Only the appropriate bit counts that are actually present in the bitstream are counted for the purposes of bit counts.

[0054] As an example, synchronization of non-VCL NAL units conveyed by means other than their presence in the bitstream with NAL units that are present in the bitstream can be achieved by indicating two points in the bitstream between which the non-VCL NAL unit would have been present in the bitstream (if the encoder decided to convey the unit in the bitstream).

[0055] When the content of a non-VCL NAL unit is conveyed for application by some means other than its presence in the bitstream, it is not required that the representation of the content of the non-VCL NAL unit uses the same syntax.

[0056] NOTE - When HRD information is contained within the bitstream, conformance to the requirements of this subclause for the bitstream can be verified based only on the information contained within the bitstream. When HRD information is not present in the bitstream (as is the case for all "standalone" type I bitstreams), conformance can only be verified when HRD data is supplied by some other means not specified in this Recommendation | International Standard.

[0057] The HRD contains coded picture buffers (CPB), an instantaneous decoding process, decoded picture buffers (DPB), and output pruning, as shown in Figure 7

[0058] The CPB size (in bits) is CpbSize[ SchedSelldx ]. For each X in the range of 0 to sps_max_temporal_layers_minus1, inclusive, the DPB size (number of picture storage buffers) for temporal layer X is sps_max_dec_pic_buffering[ X ].

[0059] The variable SubPicCpbPreferredFlag is specified by external means or, when not specified by external means, is set equal to 0.

[0060] The variable SubPicCpbFlag is derived as follows:

[0061] SubPicCpbFlag = SubPicCpbPreferredFlag && sub_pic_cpb_params_present_flag

[0062] If SubPicCpbFlag is equal to 0, the CPB operates at the access unit level and each decoding unit is an access unit. Otherwise, the CPB operates at the sub-picture level and each decoding unit is a subset of an access unit.

[0063] The HRD operates as follows. Data related to decoding units that are to arrive in the CPB according to a specified arrival schedule is delivered by the HSS. Data related to each decoding unit is removed and decoded instantaneously by the instantaneous decoding process at the CPB removal time. Each decoded picture is placed in the DPB. Decoded pictures are removed from the DPB at a later DPB output time or at a time when the inter prediction reference is no longer needed.

[0064] The HRD is initialized as specified in the buffering period SEI. The removal timing of decoding units removed from the CPB and the output timing of decoded pictures output from the DPB are specified in the picture timing SEI message. All timing information related to a particular decoding unit will arrive prior to the cpb removal time of the decoding unit.​

[0065] The HRD is used to check the bitstream and decoder conformance.

[0066] While conformance is guaranteed under the assumption that all frame rates and clocks used to generate the bitstream exactly match the values signaled in the bitstream, in a real system each of these frame rates and clocks can be different from the signaled or specified values.

[0067] All arithmetic is done with the real values, so rounding errors are not propagated. For example, the number of bits in the CPB immediately before or after removal of the decoding units is not necessarily an integer.

[0068] The variable t c is derived as follows and is referred to as the clock tick:

[0069] t c = num units in tick ÷ time scale

[0070] The variable t c_sub is derived as follows and is referred to as the subpicture clock tick:

[0071] t c_sub = num units in sub tick ÷ time scale

[0072] The following is specified to express the constraints:

[0073] - Access unit n is the nth access unit in decoding order, where the first access unit is access unit 0.

[0074] - Picture n is the coded picture or decoded picture of access unit n.

[0075] - Decoding unit m is the mth access unit in decoding order, where the first decoding unit is decoding unit 0.

[0076] In [2], the segment header syntax allows for so-called dependent segments.

[0077] Figure 8 The segment header syntax of [2] is shown.

[0078] The segment header semantics are defined as follows:

[0079] dependent_slice_flag equal to 1 specifies that the value of each non-existing slice header syntax element is inferred to be equal to the value of the corresponding slice header syntax element in the previous slice pair containing the coding tree block for which the coding tree block address is SliceCtbAddrRS - 1. When not present, the value of dependent_slice_flag is inferred to be equal to 0. When SliceCtbAddrRS is equal to 0, the value of dependent_slice_flag shall be equal to 0.

[0080] slice_address specifies the address in the slice granularity parsing at which the slice starts. The length of the slice_address syntax element is (Ceil( Log2( PicWidthInCtbs * PicHeightInCtbs ) + SliceGranularity ) ) bits.

[0081] The variable SliceCtbAddrRS specifying the coding tree block in which the slice starts in the raster scan order of the coding tree blocks is derived as follows.

[0082] SliceCtbAddrRS = ( slice_address » SliceGranularity )

[0083] The variable SliceCbAddrZS specifying the address of the first coding block in the slice in the z-scan order at the smallest coding block granularity is derived as follows.

[0084] SliceCbAddrZS = slice_address

[0085] << ( log2_diff_max_min_coding_block_size - SliceGranularity ) « 1 )

[0086] Slice decoding starts at the possible largest coding unit at the slice start coordinate.

[0087] first_slice_in_pic_flag indicates whether the slice is the first slice of the picture. If first_slice_in_pic_flag is equal to 1, the variables SliceCbAddrZS and SliceCtbAddrRS are both set to 0 and the decoding starts at the first coding tree block in the picture.

[0088] pic_parameter_set_id specifies the picture parameter set in use. The value of pic_parameter_set_id shall be in the range of 0 to 255, inclusive.

[0089] num_entry_point_offsets specifies the number of entry_point_offset[ i ] syntax elements in the slice header. When tiles_or_entropy_coding_sync_idc is equal to 1, the value of num_entry_point_offsets shall be in the range of 0 to ( num_tile_columns_minusl + 1 ) * ( num_tile_rows_minusl + 1 ) - 1, inclusive. When tiles_or_entropy_coding_sync_idc is equal to 2, the value of num_entry_point_offsets shall be in the range of 0 to PicHeightInCtbs - 1, inclusive. When not present, the value of num_entry_point_offsets is inferred to be equal to 0.

[0090] offset_len_minusl plus 1 specifies the length in bits of the entry_point_offset[ i ] syntax element.

[0091] entry_point_offset[ i ] specifies the i-th entry point offset in bytes and shall be represented by offset len minusl plus 1 bits. The coded slice data after the slice header consists of num entry point offsets + 1 subsets, with subset index values in the range of 0 to num entry point offsets, inclusive. Subset 0 consists of bytes 0 to entry_point_offset[ 0 ] - 1, inclusive, of the coded slice data, subset k, where k is in the range of 1 to num entry point offsets - 1, inclusive, consists of bytes entry_point_offset[ k - 1 ] to entry_point_offset[ k ] + entry_point_offset[ k - 1 ] - 1, inclusive, of the coded slice data, and the last subset, where the subset index is equal to num entry point offsets, consists of the remaining bytes of the coded slice data.

[0092] When tiles_or_entropy_coding_sync_idc is equal to 1 and num_entry_point_offsets is greater than 0, each subset shall contain exactly all coded bits of one tile, and the number of subsets (i.e., the value of num_entry_point_offsets + 1) shall be equal to or less than the number of tiles in the slice.

[0093] When tiles_or_entropy_coding_sync_idc is equal to 1, each slice must include either a subset of one tile (in which case signaling of entry points is not necessary) or an integer number of complete tiles.

[0094] When tiles_or_entropy_coding_sync_idc is equal to 2 and num_entry_point_offsets is greater than 0, each subset k, where k is in the range of 0 to num_entry_point_offsets - 1, inclusive, shall contain exactly one column of Coded Tree Blocks of all coded bits, and the last subset, where the subset index is equal to num_entry_point_offsets, shall contain all coded bits of the remaining Coded Tree Blocks included in the slice, where the remaining Coded Tree Blocks consist of exactly one column of Coded Tree Blocks or a subset of a column of Coded Tree Blocks, and the number of subsets, i.e., the value of num_entry_point_offsets + 1, shall be equal to the number of columns of Coded Tree Blocks in the slice, where subsets of a column of Coded Tree Blocks are also counted.

[0095] When tiles_or_entropy_coding_sync_idc is equal to 2, a slice can include a number of columns of Coded Tree Blocks and a subset of a column of Coded Tree Blocks. For example, if a slice includes 2.5 columns of Coded Tree Blocks, and the number of subsets, i.e., the value of num_entry_point_offsets + 1, shall be equal to 3.

[0096] Figure 9 The picture parameter set RBSP syntax of [2] is shown in [2], the picture parameter set RBSP semantics is defined as follows:

[0097] dependent_slice_enabled_flag equal to 1 specifies the presence of the syntax element dependent_slice_flag in the slice header of coded pictures referring to this picture parameter set. dependent_slice_enabled_flag equal to 0 specifies the absence of the syntax element dependent_slice_flag in the slice header of coded pictures referring to this picture parameter set. When tiles_or_entropy_coding_sync_idc is equal to 3, the value of dependent_slice_enabled_flag shall be equal to 1.

[0098] tiles_or_entropy_coding_sync_idc equal to 0 specifies that there will be only one tile in each picture referring to the picture parameter set, that the specific synchronization process for context variables will not be invoked before decoding the first coding tree block of a column of coding tree blocks in each picture referring to the picture parameter set, and that the values of cabac_independent_flag and dependent_slice_flag for coded pictures referring to the picture parameter set will not both equal 1.

[0099] When cabac_independent_flag and dependent_slice_flag both equal 1 for a slice, the slice is an entropy slice.

[0100] tiles_or_entropy_coding_sync_idc equal to 1 specifies that there will be more than one tile in each picture referring to the picture parameter set, that the specific synchronization process for context variables will not be invoked before decoding the first coding tree block of a column of coding tree blocks in each picture referring to the picture parameter set, and that the values of cabac_independent_flag and dependent_slice_flag for coded pictures referring to the picture parameter set will not both equal 1.

[0101] tiles_or_entropy_coding_sync_idc equal to 2 specifies that there will be only one tile in each picture referring to the picture parameter set, that the specific synchronization process for context variables will be invoked before decoding the first coding tree block of a column of coding tree blocks in each picture referring to the picture parameter set, that the specific memory process for context variables will be invoked after decoding two coding tree blocks of a column of coding tree blocks in each picture referring to the picture parameter set, and that the values of cabac_independent_flag and dependent_slice_flag for coded pictures referring to the picture parameter set will not both equal 1.

[0102] tiles_or_entropy_coding_sync_idc equal to 3 specifies that there will be only one tile in each picture referring to the picture parameter set, that the specific synchronization process for context variables will not be invoked before decoding the first coding tree block of a column of coding tree blocks in each picture referring to the picture parameter set, and that the values of cabac_independent_flag and dependent_slice_flag for coded pictures referring to the picture parameter set can both equal 1.

[0103] When dependent_slice_enabled_flag shall not be equal to 0, tiles_or_entropy_coding_sync_idc shall not be equal to 3.

[0104] The requirement for bitstream conformance is that the value of tiles_or_entropy_coding_sync_idc shall be the same for all picture parameter sets activated within a coded video sequence.

[0105] For each slice referring to the picture parameter set, when tiles_or_entropy_coding_sync_idc is equal to 2 and the first coded block in the slice is not the first coded block in the first coding tree block of a column of coding tree blocks, the last coded block in the slice shall belong to the same column of coding tree blocks as the first coded block in the slice.

[0106] num_tile_columns_minus1 plus 1 specifies the number of tile rows that partition the picture.

[0107] num_tile_rows_minus1 plus 1 specifies the number of tile columns that partition the picture. When num_tile_columns_minus1 is equal to 0, num_tile_rows_minus1 shall not be equal to 0.

[0108] uniform_spacing_flag equal to 1 specifies that row boundaries and column boundaries are uniformly distributed across the picture. uniform_spacing_flag equal to 0 specifies that row boundaries and column boundaries are not uniformly distributed across the picture, but are explicitly signaled using the syntax elements column_width[ i ] and row_height[ i ].

[0109] column_width[ i ] specifies the width of the i-th tile column in units of coding tree blocks.

[0110] row_height[ i ] specifies the height of the i-th tile row in units of coding tree blocks.

[0111] The vector colWidth[ i ] specifies the width of the i-th tile column in units of CTBs, where i is in the range of 0 to num_tile_columns_minus1, inclusive.

[0112] The vector CtbAddrRStoTS[ ctbAddrRS ] specifies the conversion from a CTB address in the raster scan order to a CTB address in the tile scan order, where the index ctbAddrRS is in the range of 0 to ( picHeightInCtbs * picWidthInCtbs ) - 1, inclusive.

[0113] The vector CtbAddrTStoRS[ ctbAddrTS ] specifies the conversion from a CTB address in the tile scan order to a CTB address in the raster scan order, where the index ctbAddrTS is in the range of 0 to ( picHeightInCtbs * picWidthInCtbs ) - 1, inclusive.

[0114] The vector TileId[ ctbAddrTS ] specifies the conversion from a CTB address in the tile scan order to a tile id, where ctbAddrTS is in the range of 0 to ( picHeightInCtbs * picWidthInCtbs ) - 1, inclusive.

[0115] The values of colWidth, CtbAddrRStoTS, CtbAddrTStoRS, and TileId are derived by invoking the CTB raster and tile scan conversion process as specified in subclause 6.5.1 with PicHeightInCtbs and PicWidthInCtbs as inputs and assigning the outputs to colWidth, CtbAddrRStoTS, CtbAddrTStoRS, and TileId.

[0116] The value of ColumnWidthInLumaSamples[ i ] (which specifies the width of the i-th tile column in luma samples) is set equal to colWidth[ i ] « Log2CtbSize.

[0117] The array MinCbAddrZS[ x ][ y ], which specifies the conversion from a location in minCBs (x, y) to the address of the smallest CB in z-scan order, where x is in the range of 0 to picWidthInMinCbs - 1, inclusive, and y is in the range of 0 to picHeightInMinCbs - 1, inclusive, is derived by invoking the Z-scan order array initialization process as specified in subclause 6.5.2 with Log2MinCbSize, Log2CtbSize, PicHeightInCtbs, PicWidthInCtbs, and vector CtbAddrRStoTS as inputs and assigning the output to MinCbAddrZS.

[0118] loop_filter_across_tiles_enabled_flag equal to 1 specifies that the in-loop filtering operations are performed across tiles. loop_filter_across_tiles_enabled_flag equal to 0 specifies that the in-loop filtering operations are not performed across tiles. The in-loop filtering operations include the deblocking filter, sample adaptive offset, and adaptive loop filter operations. When not present, the value of loop_filter_across_tiles_enabled_flag is inferred to be equal to 1.

[0119] cabac_independent_flag equal to 1 specifies that the CABAC decoding of the coded blocks in the slice is independent of the state of any previously decoded slice. cabac_independent_flag equal to 0 specifies that the CABAC decoding of the coded blocks in the slice depends on the state of the previously decoded slice. When not present, the value of cabac_independent_flag is inferred to be equal to 0.

[0120] The derivation process of the availability of coded blocks with the smallest coded block address is described as follows:

[0121] The inputs to this process are

[0122] - the smallest coded block address in z-scan order, minCbAddrZS

[0123] - the current smallest coded block address in z-scan order, currMinCBAddrZS

[0124] The output of this process is the availability of coded blocks with the smallest coded block address in z-scan order, cbAvailable.

[0125] Note that the meaning of determining availability is when this process is invoked.

[0126] Note that any coding block, regardless of its size, is associated with a minimum coding block address, which is the address of the coding block with the smallest coding block size in z-scan order.

[0127] - If one or more of the following conditions are true, cbAvailable is set to false.

[0128] - minCbAddrZS is less than 0

[0129] - minCbAddrZS is greater than currMinCBAddrZS

[0130] - the coding block with the minimum coding block address minCbAddrZS belongs to a different slice than the coding block with the current minimum coding block address currMinCBAddrZS and dependent_slice_flag of the slice containing the coding block with the current minimum coding block address currMinCBAddrZS is equal to 0

[0131] - the coding block with the minimum coding block address minCbAddrZS is included in a different tile than the coding block with the current minimum coding block address currMinCBAddrZS.

[0132] - Otherwise, cbAvailable is set to true.

[0133] The CABAC parsing process for slice data of [2] is as follows:

[0134] This process is invoked when parsing the syntax element with descriptor ae(v).

[0135] The input to this process is a request for the value of a syntax element and the values of previously parsed syntax elements.

[0136] The output of this process is the value of the syntax element.

[0137] The initialization process of the CABAC parsing process is invoked when starting the parsing of the slice data of a slice.

[0138] The minimum coding block address of the coding tree block containing the spatially neighboring blocks T( Figure 10A ) is derived as follows using the position (x0, y0) of the top-left luma sample of the current coding tree block.

[0139] x = x0 + 2 « Log2CtbSize - 1

[0140] y = y0 - 1

[0141] ctbMinCbAddrT = MinCbAddrZS[x » Log2MinCbSize][y » Log2MinCbSize]

[0142] The variable availableFlagT is obtained by invoking the coding block availability derivation process with ctbMinCbAddrT as input.

[0143] When starting the parsing of a coding tree, the following ordered steps apply.

[0144] 1. The arithmetic decoding engine is initialized as follows.

[0145] - If CtbAddrRS is equal to Slice_address, dependent_slice_flag is equal to 1 and entropy_coding_reset_flag is equal to 0, the following applies.

[0146] - The synchronization process of the CABAC parsing process is invoked with TableStateIdxDS and TableMPSValDS as inputs.

[0147] - The decoding process for bin decisions before termination is invoked, followed by the initialization process for the arithmetic decoding.

[0148] - Otherwise, if tiles_or_entropy_coding_sync_idc is equal to 2 and CtbAddrRS % PicWidthInCtbs is equal to 0, the following applies.

[0149] - When availableFlagT is equal to 1, the synchronization process of the CABAC parsing process is invoked with TableStateIdxWPP and TableMPSValWPP as inputs.

[0150] - The decoding process for bin decisions before termination is invoked, followed by the initialization process for the arithmetic decoding engine.

[0151] 2. When cabac_independent_flag is equal to 0 and dependent_slice_flag is equal to 1, or when tiles_or_entropy_coding_sync_idc is equal to 2, the following memory process applies.

[0152] - When tiles_or_entropy_coding_sync_idc is equal to 2 and CtbAddrRS % PicWidthlnCtbs is equal to 2, the memory process of the CABAC parsing process is invoked with TableStateIdxWPP and TableMPSValWPP as inputs.

[0153] - When cabac_independent_flag is equal to 0, dependent_slice_flag is equal to 1, and end_of_slice_flag is equal to 1, the memory process of the CABAC parsing process is invoked with TableStateIdxDS and TableMPSValDS as outputs.

[0154] The parsing of the syntax elements is done as follows:

[0155] For each requested value of a syntax element, a binarization is derived.

[0156] The binarization of the syntax element and the decoded process flow of the parsed bins.

[0157] For each bin of the binarization of the syntax element (indexed by the variable binldx), a context index ctxldx is derived.

[0158] The arithmetic decoding process is invoked for ctxldx.

[0159] After decoding each bin, the resulting sequence of parsed bins (b 0.. b binIdx ) is compared with the set of bin strings given by the binarization process. When the sequence matches a bin string in the given set, the corresponding value is assigned to the syntax element.

[0160] When the request for the value of the syntax element is processed for the syntax element pcm-flag and the decoded value of pcm_flag is equal to 1, the decoding engine is initialized after decoding any pcm_alignment_zero_bit, num_subsequent_pcm, all pcm_sample_luma, and pcm_sample_chroma data.

[0161] In the design framework described so far, the following problems arise.

[0162] Before encoding and sending data in a low delay situation, the timing of the decoding units needs to be known, where the NAL units would have been issued by the encoder while the encoder is still encoding part of the picture, i.e. other sub-picture decoding units. That is, because the NAL unit order in an access unit only allows SEI messages before the VCL (video coding NAL units) in the access unit, but in this low delay situation the non-VCL NAL units need to be already on-line, i.e. issued, if the encoder starts to encode the decoding unit. Figure 10B The structure of an access unit as defined in [2] is illustrated. [2] does not specify the end of a sequence or stream, so its presence in an access unit is assumed.

[0163] Furthermore, in a low delay situation the number of NAL units related to a sub-picture also needs to be known in advance, because the picture timing SEI message contains this information and has to be composed and issued before the encoder starts to encode the actual picture. Application designers who do not want to insert filler data NAL units (possibly without filler data to fit the NAL unit number, as signaled in the picture timing SEI for each decoding unit) need a means to signal this information on the sub-picture level. The same applies to the sub-picture timing, which is currently fixed at the beginning of the access unit by parameters given in the timing SEI message.

[0164] Further drawbacks of the draft specification [2] include the large amount of signaling on the sub-picture level, which is required for certain applications, such as ROI signaling or segment size signaling.

[0165] The problems outlined above are not specific to the HEVC standard. Rather, the same problems arise with other video codecs as well. Figure 11 More generally, a video transmission scenario is shown in which a pair of an encoder 10 and a decoder 12 are connected via a network 14 for transmitting video 16 from the encoder 10 to the decoder 12 with a short end-to-end delay. The figures outlined above have been summarized above. The encoder 10 encodes a sequence of frames 18 of the video 16 according to a certain decoding order which generally, but not necessarily, follows a reproduction order 20 of the frames 18, and proceeds through the frame regions of the frames 18 in some defined manner within each frame 18, such as in a raster scan manner with or without tile partitioning of the frames 18. The decoding order controls the availability of information for the encoding techniques used by the encoder 10, such as prediction and / or entropy encoding, i.e. the availability of information about spatial and / or temporal neighboring parts of the video 16 which can be used as a basis for prediction or context selection. Even though the encoder 10 can be able to use parallel processing for encoding the frames 18 of the video 16, the encoder 10 necessarily needs some time to encode a certain frame 18, such as the current frame. For example, Figure 11An instant is illustrated where the encoder 10 has finished encoding a portion 18a of the current frame 18, while another portion 18b of the current frame 18 has not yet been encoded. Since the encoder 10 has not yet encoded the portion 18b, the encoder 10 can not be able to predict how the available bit rate for encoding the current frame 18 should be spatially distributed over the current frame 18 to achieve an optimization, e.g. in terms of rate / distortion. Thus, the encoder 10 has only two options: the encoder 10 pre-estimates a near-optimal distribution of the available bit rate of the current frame 18 over the segments into which the current frame 18 is spatially subdivided, accepting that the estimate can be wrong accordingly; or the encoder 10 finishes encoding the current frame 18 before transmitting the packets containing the segments from the encoder 10 to the decoder 12. In any case, in order to be able to exploit any transmission of segment packets of the current encoded frame 18 before finishing the encoding of the current encoded frame 18, the segments 14 should be informed about the bit rate associated with each such segment packet in the form of an encoded picture buffer retrieval time. However, as indicated above, although the encoder 10, according to the current version of HEVC, is able to vary the bit rate distributed over the frame 18 by using decoder buffer retrieval times defined separately for sub-picture regions, the encoder 10 needs to transmit or signal this information via the network 14 at the beginning of each access unit that collects all data related to the current frame 18, thereby forcing the encoder 10 to choose between the two alternatives just outlined, one leading to lower delay but worse rate / distortion, the other leading to optimal rate / distortion but increased end-to-end delay.

[0166] Thus, no video codec has allowed to achieve this low delay so far, so that the encoder would be allowed to start transmitting packets related to a portion 18a of the current frame before finishing encoding the remaining portion 18b of the current frame, the decoder being able to exploit this intermediate transmission of packets related to the preliminary portion 18a via the network 16, which complies with the decoding buffer retrieval timing conveyed within the video data stream sent from the encoder 12 to the decoder 14. Applications that would exploit this low delay exemplarily include industrial applications, such as workpiece or manufacturing monitoring for automation or inspection purposes or the like. No satisfactory solution has been available so far to inform the decoding side about the association of packets to tiles constituting the current frame and to interesting regions (regions of interest) of the current frame, thus allowing intermediate network entities within the network 16 to collect this information from the data stream without having to deeply inspect the interior of the packets, i.e. the segment syntax.

[0167] It is thus an object of the present application to provide a video data stream encoding concept technique that more efficiently allows for low end-to-end delay and / or makes the identification of portions of the data stream to regions of interest or to specific tiles easier.

[0168] The objects described by the attached independent claims are achieved. SUMMARY

[0169] The present application relates to a video data stream having video content encoded therein in units of sub-portions of pictures of the video content, each sub-portion being encoded into one or more payload packets of a packet sequence of the video data stream, the packet sequence being divided into a sequence of access units, thus each access unit collecting payload packets pertaining to respective pictures of the video content, wherein the packet sequence has timing control packets interspersed therein, thus the timing control packets subdividing the access units into decoding units, thus at least some access units being subdivided into two or more decoding units, wherein each timing control packet signals a decoder buffer retrieval time for a decoding unit, the payload packets of the decoding unit following the respective timing control packet in the packet sequence.

[0170] One idea underlying the present application is that decoder retrieval timing information, ROI information and tile identification information should be conveyed within a video data stream on a level that allows easy access by network entities such as MANEs or decoders; and in order to achieve this level, such types of information should be conveyed within the video data stream by means of packets interspersed among the packets of the access units of the video data stream. According to one embodiment, the interspersed packets are of a removable packet type, i.e. removal of such interspersed packets maintains the ability of a decoder to fully reconstruct the video content conveyed via the video data stream.

[0171] According to one aspect of the present application, the achievement of low end-to-end delay is made more efficient by using interspersed packets to convey information about the decoder buffer retrieval time for a decoding unit within the current access unit (the decoding unit being formed by payload packets following a respective timing control packet in the video data stream). By this measure, the encoder is allowed to determine the decoder buffer retrieval time on the fly while encoding the current frame, the encoder in turn being able to, while encoding the current frame, continue to determine the bit rate actually used for the part of the current frame that has already been encoded into payload packets and transmitted or emitted (which is prefixed by a timing control packet), and to adapt the distribution of the remaining bit rate available for the current frame over the remaining part of the current frame that has not yet been encoded accordingly. By this measure, the available bit rate is efficiently utilized, and still the delay is kept short, because the encoder does not need to wait for the encoding of the current frame to be completed.

[0172] According to another aspect of the present application, the information about the region of interest is conveyed using a packet that is interspersed in the payload packets of the access unit, which allows easy access to this information by network entities as outlined above, since these network entities do not have to inspect the intermediate payload packets. Moreover, the encoder can still decide in real time, during the encoding of the current frame, which packets belong to the ROI, without having to decide in advance how the current frame is subdivided into subparts and individual payload packets. Furthermore, according to an embodiment (according to which the interspersed packet is of a removable packet type), a video stream receiver that is not interested in or capable of processing the ROI information can ignore the ROI information.

[0173] According to another aspect, a similar idea is used in the present application according to which the interspersed packet conveys information about which tile a particular packet within an access unit belongs to. BRIEF DESCRIPTION OF DRAWINGS

[0174] Advantageous implementations of the present application are subject matter of the dependent claims. Preferred embodiments of the present application are described in more detail below with respect to the figures, in which:

[0175] Figures 1 to 10B A current state of HEVC is shown, wherein Figure 1 A buffer period SEI message syntax is shown, Figure 2 A picture timing SEI message syntax is shown, Figure 3A and Figure 3B A VUI parameter syntax is shown, Figure 4 An HRD parameter syntax is shown, Figure 5 A filler data RBSP syntax is shown, Figure 6 The structure of the bitstream and the NAL unit stream for HRD conformance checking is shown, Figure 7 An HRD buffer model is shown, Figure 8 A slice header syntax is shown, Figure 9 A picture parameter set RBSP syntax is shown, Figure 10A A diagram illustrating a spatially neighboring coding tree block T that can be used to invoke a coding tree block availability derivation process related to the current coding tree block is shown, and Figure 10B A definition of the structure of an access unit is shown;

[0176] Figure 11 A pair of encoder and decoder connected via a network is schematically shown in order to illustrate the problem occurring in the transmission of a video stream;

[0177] Figure 12 A schematic block diagram of an encoder according to an embodiment using timing control packets is shown;

[0178] Figure 13A flowchart showing the operational mode of an encoder according to an embodiment Figure 12 A flowchart showing the operational mode of an encoder according to an embodiment

[0179] Figure 14 A block diagram showing one embodiment of a decoder to explain the functionality of the decoder with respect to a video data stream generated by an encoder according to Figure 12 A block diagram showing one embodiment of a decoder to explain the functionality of the decoder with respect to a video data stream generated by an encoder according to

[0180] Figure 15 A schematic block diagram showing an encoder, a network entity and a video data stream according to another embodiment using ROI;

[0181] Figure 16 A schematic block diagram showing an encoder, a network entity and a video data stream according to another embodiment using tile identification packets;

[0182] Figure 17 A structure of an access unit according to an embodiment is shown. The dashed line reflects the case of a non-mandatory slice prefix NAL unit;

[0183] Figure 18 The use of tiles in attention signaling is shown;

[0184] Figure 19 A first simple syntax / version 1 is shown;

[0185] Figure 20 An extended syntax / version 2 is shown, which includes tile_id signaling, decoding unit start identifier, slice prefix ID and slice header data in addition to the SEI message concept;

[0186] Figure 21 NAL unit type codes and NAL unit type categories are shown;

[0187] Figure 22 A possible syntax of a slice header is shown, wherein specific syntax elements present in the slice header according to the current version are transformed into lower level syntax elements called slice_header_data();

[0188] Figure 23A , Figure 23B and Figure 23C A table of all syntax elements removed from the slice header signaled via the syntax element slice header data is shown;

[0189] Figure 24 A supplemental enhancement information message syntax is shown;

[0190] Figure 25A and Figure 25BAn adapted SEI payload syntax is shown in order to introduce new tile or subpicture SEI message types;

[0191] Figure 26 An example of a subpicture buffering SEI message is shown;

[0192] Figure 27 An example of a subpicture timing SEI message is shown;

[0193] Figure 28 An example of how a subpicture segment information SEI message can look like is shown;

[0194] Figure 29 An example of a subpicture tile information SEI message is shown;

[0195] Figure 30 An example of the syntax of a subpicture segment size information SEI message is shown;

[0196] Figure 31 A first variant of the syntax example of a region of interest SEI message is shown, where each ROI is signaled in an individual SEI message;

[0197] Figure 32 A second variant of the syntax example of a region of interest SEI message is shown, where all ROIs are signaled in a single SEI message;

[0198] Figure 33 A possible syntax of a timing control packet according to another embodiment is shown;

[0199] Figure 34 A possible syntax of a tile identification packet according to an embodiment is shown;

[0200] Figures 35 to 38 A possible subdivision of an image according to different subdivision settings according to an embodiment is shown; and

[0201] Figure 39 An example of a portion of a video data stream according to an embodiment using timing control packets is shown, interspersed between the payload packets of access units. DETAILED DESCRIPTION

[0202] With regard to Figure 12 An encoder 10 and its operational modes according to one embodiment of the present application are described. The encoder 10 is configured to encode video content 16 into a video data stream 22. The encoder is configured to do this in units of sub-portions of frames / images 18 of the video content 16, which can be, for example, segments 24 into which the images 18 are partitioned, or some other spatial section, such as tiles 26 or WPP sub-streams 28, all of which are described in more detail in Figure 12It is noted that this is only for illustrative purposes and does not imply that the encoder 10 needs to be able to support e.g. tile or WPP parallel processing or that the sub- portions need to be segments.

[0203] When encoding the video content 16 in units of sub-portions 24, the encoder 10 can adhere to a decoding order (or encoding order) defined between the sub-portions 24, which e.g. traverses the pictures 18 of the video 16 according to a picture decoding order (which e.g. does not necessarily comply with the reproduction order 20 defined between the pictures 18) and according to a raster scan order within each picture 18, where the sub-portions 24 represent consecutive runs of such blocks along the decoding order. In detail, the encoder 10 can be configured to adhere to this decoding order when deciding on the availability of spatially and / or temporally neighboring portions of the portion currently to be encoded for use in e.g. predictive encoding and / or entropy encoding, describing properties of such neighboring portions. Only previously visited (encoded / decoded) portions of the video are available. Otherwise, the properties just mentioned are set to default values, or some other alternative measures are taken.

[0204] On the other hand, the encoder 10 does not need to encode the sub-portions 24 along the decoding order serially. Rather, the encoder 10 can use parallel processing to speed up the encoding process, or be able to perform more complex encoding in real time. Likewise, the encoder 10 can or can not be configured to transmit or issue data encoding the sub-portions along the decoding order. For example, the encoder 10 can output / transmit the encoded data in some other order, such as according to the order in which the encoder 10 finishes encoding the sub-portions (which can deviate from the decoding order just mentioned due to e.g. parallel processing).

[0205] To make the encoded versions of the sub-portions 24 suitable for transmission over a network, the encoder 10 encodes each sub-portion 24 into one or more payloads packets of a sequence of packets of the video data stream 22. In case the sub-portions 24 are segments, the encoder 10 can e.g. be configured to put each segment data (i.e. each encoded segment) into one payload packet (such as a NAL unit). This packetization can be used to make the video data stream 22 suitable for transmission over a network. Thus, the packets can represent the smallest units at which the video data stream 22 can occur, i.e. can each be the smallest unit that the encoder 10 issues for transmission over a network to a recipient.

[0206] In addition to the payload packets and the timing control packets interspersed therebetween and discussed below, other packets (i.e. other types of packets) can likewise be present, such as filler data packets, picture or sequence parameter set packets for conveying syntax elements that do not change often, or EOF (end of file) or AUE (access unit end) packets or the like.

[0207] The encoder performs encoding into the payload packets such that the sequence of packets is divided into a sequence of access units 30, and each access unit collects the payload packets 32 that are related to one picture 18 of the video content 16. That is, the sequence of packets 34 forming the video data stream 22 is subdivided into non-overlapping portions called access units 30, each of which is related to an individual one of the pictures 18. The sequence of access units 30 can follow the decoding order of the pictures 18 to which the access units 30 relate. Figure 12 For example, the access unit 30 shown in the middle of the illustrated portion of the data stream 22 includes one payload packet 32 for each subpart 24 into which the picture 18 is subdivided. That is, each payload packet 32 carries a corresponding subpart 24. The encoder 10 is configured to intersperse timing control packets 36 in the sequence of packets 34, so that the timing control packets subdivide the access units 30 into decoding units 38, so that at least some of the access units 30, such as the middle access unit shown in Figure 12 The middle access unit shown in the middle of the illustrated portion of the data stream 22 includes one payload packet 32 for each subpart 24 into which the picture 18 is subdivided. That is, each payload packet 32 carries a corresponding subpart 24. The encoder 10 is configured to intersperse timing control packets 36 in the sequence of packets 34, so that the timing control packets subdivide the access units 30 into decoding units 38, so that at least some of the access units 30, such as the middle access unit shown in Figure 12 For example, the case is illustrated in which every other packet 32 represents a first payload packet of a decoding unit 38 of an access unit 30. As illustrated in Figure 12 As illustrated in the middle of the illustrated portion of the data stream 22, the amount of data or bit rate for each decoding unit 38 varies, and the decoder buffer retrieval time can be related to this variation in bit rate between decoding units 38, because the decoder buffer retrieval time for a decoding unit 38 can follow the decoder buffer retrieval time signaled by the timing control packet 36 of the immediately preceding decoding unit 38 plus a time interval that corresponds to the bit rate for this immediately preceding decoding unit 38.

[0208] That is, the encoder 10 can be configured to signal the decoder buffer retrieval time for each decoding unit 38 in the sequence of access units 30, so that the decoder buffer retrieval time for each decoding unit 38 is signaled in the sequence of packets 34. Figure 13The encoder 10 can operate as illustrated in the flowchart of Fig. 4. In detail, as mentioned above, the encoder 10 can subject a current sub-section 24 of the current picture 18 to encoding in step 40. As already mentioned, the encoder 10 can sequentially loop through the sub-sections 24 in the above-mentioned decoding order, as exemplified by the arrow 42, or the encoder 10 can use some kind of parallel processing, such as WPP and / or tile processing, in order to encode several "current sub-sections" 24 at the same time. Whether or not parallel processing is used, the encoder 10 forms a decoding unit from one or several of the sub-sections that have just been encoded in step 40, and proceeds with step 44 in which the encoder 10 sets the decoder buffer retrieval time of this decoding unit and transmits this decoding unit prefixed with a time control packet that signals the just set decoder buffer retrieval time of this decoding unit. For example, the encoder 10 can determine the decoder buffer retrieval time in step 44 based on the bit rate used for encoding the sub-sections that have been encoded into the payload packet forming the current decoding unit, which includes, for example, all other intermediate packets (if any) within this decoding unit, i.e. the "prefixing packets".

[0209] Then, in step 46, the encoder 10 can adapt the available bit rate based on the bit rate that has been used for the decoding unit that has just been transmitted in step 44. For example, if the image content within the decoding unit that has just been transmitted in step 44 is very complex in terms of compression rate, the encoder 10 can reduce the available bit rate for the next decoding unit in order to comply with some externally set target bit rate that has been determined based on, for example, the current bandwidth situation faced by the network with respect to the transmitted video data stream 22. Steps 40 to 46 are then repeated. By this measure, the picture 18 is encoded and transmitted (i.e. signaled) in units of decoding units, each prefixed with a corresponding timing control packet.

[0210] In other words, the encoder 10 encodes 40 a current sub-section 24 of the current picture 18 under encoding of the video content 16 into a current payload packet 32 of a current decoding unit 38, transmits 44 the current decoding unit 38 prefixed with a current timing control packet 36 by setting the decoder buffer retrieval time signaled by the current timing control packet (36) at a first time instant within the data stream, and encodes 44 another sub-section 24 of the current picture 18 by looping back from step 46 to 40 at a second time instant (second visit of step 40) later than the first time instant (first visit of step 44).

[0211] Because the encoder can issue such decoding units before the remaining portion of the current picture to which the decoding unit belongs, the encoder 10 can reduce the end-to-end delay. On the other hand, the encoder 10 need not waste available bit rate because the encoder 10 can react to the specific nature of the content of the current picture and to the spatial distribution of its complexity.

[0212] On the other hand, the intermediate network entity responsible for further transmitting the video data stream 22 from the encoder to the decoder can use the timing control packets 36 to ensure that any decoder receiving the video data stream 22 receives the decoding units in time to be able to exploit the encoding and transmission of the decoding units by the encoder 10 on a decoding unit by decoding unit basis. See, for example, Figure 14 which shows an example of a decoder for decoding the video data stream 22. The decoder 12 receives the video data stream 22 at a coded picture buffer CPB 48 by way of a network via which the encoder 10 transmits the video data stream 22 to the decoder 12. In particular, because the network 14 is assumed to be able to support low delay applications, the network 10 inspects the decoder buffer retrieval times in order to feed the packet sequence 34 of the video data stream 22 to the coded picture buffer 48 of the decoder 12 so that each decoding unit is present in the coded picture buffer 48 before the decoder buffer retrieval time signaled by the timing control packet prefixed to the decoding unit. By this measure, the decoder can use the decoder buffer retrieval times in the timing control packets to empty the coded picture buffer 48 of the decoder in units of decoding units rather than complete access units without stalling, i.e., without using up the available payload packets in the coded picture buffer 48. See, for example, Figure 14 The processing unit 50 is shown connected to the output of the coded picture buffer 48 for which the input receives the video data stream 22 for illustrative purposes. Like the encoder 10, the decoder 12 can be able to perform parallel processing, such as using tile parallel processing / decoding and / or WPP parallel processing / decoding.

[0213] As will be outlined in more detail below, the decoder buffer retrieval times mentioned so far do not necessarily relate to the retrieval times of the coded picture buffer 48 of the decoder 12. Rather, the timing control packets can additionally or otherwise manipulate the retrieval of decoded picture data of a corresponding decoded picture buffer of the decoder 12. See, for example, Figure 14The decoder 12 is shown to comprise a decoder picture buffer in which the decoded version of the video content, i.e. the decoded version of the video data stream 22 as obtained by the processing unit 50, is buffered, i.e. stored and output, in units of decoded units. The decoded picture buffer 22 of the decoder can thus be connected between the output of the decoder 12 and the output of the processing unit 50. By having the ability to set the retrieval time at which the decoded version of a decoded unit is output from the decoded picture buffer 52, the encoder 10 has the opportunity to control the rendering of the video content at the decoding side or the end-to-end delay of the rendering in real time, i.e. during the encoding of the current picture, even to a degree smaller than the image rate or frame rate. Obviously, over-segmenting each picture 18 into a large number of sub-portions 24 at the encoding side will adversely affect the bit rate for transmitting the video data stream 22, but on the other hand, the end-to-end delay can be minimized as the time required for encoding and transmitting and decoding and outputting this decoded unit will be minimized. On the other hand, increasing the size of the sub-portions 24 will increase the end-to-end delay. Hence, a trade-off has to be found. Using the decoder buffer retrieval time just mentioned to manipulate the output timing of the decoded version of the sub-portions 24 in units of decoded units allows the encoder 10 or some other unit at the encoding side to spatially adapt this trade-off over the content of the current picture. By this measure, it will be possible to control the end-to-end delay in a way that the end-to-end delay spatially varies across the content of the current picture.

[0214] In implementing the embodiments outlined above, it is possible to use packets of the removable packet type as timing control packets. Packets of the removable packet type are not necessary for the reconstruction of the video content at the decoding side. In the following, these packets are referred to as SEI packets. Other packets of the removable packet type can also exist, i.e. another type of removable packet, such as a redundant packet, if transmitted within the stream. As another alternative, the timing control packets can be packets of a specific removable packet type, however, which additionally carry a specific SEI packet type field. For example, the timing control packets can be SEI packets, wherein each SEI packet carries one or several SEI messages, and only those SEI packets comprising a specific type of SEI message form the timing control packets described above.

[0215] Hence, according to another embodiment, the embodiments described so far with respect to Figures 12 to 14 are applied to the HEVC standard, thereby forming a possible conceptual technique for making HEVC more efficient in achieving lower end-to-end delays. In this way, the packets mentioned above are formed by NAL units, and the payload packets described above are VCL NAL units of the NAL units, wherein the slices form the sub-portions mentioned above.

[0216] However, before describing the more detailed embodiments, other embodiments are described that conform to the embodiments outlined above in that they use interleaved packets to efficiently convey information describing the video data stream, but the classification of information differs from the embodiments described above, where timing control packets convey timing information captured by the decoder buffer. In the embodiments further described below, the types of information conveyed via interleaved packets interspersed within payload packets belonging to access units relate to Region of Interest (ROI) information and / or tile identification information. The embodiments further described below may or may not relate to the embodiments described above. Figures 12 to 14 The described embodiments are combined.

[0217] Figure 15 Encoder 10 is shown, in addition to the interleaving of timing control packets and the above-mentioned... Figure 13 The described functionality (its relation to) Figure 15 Apart from the encoder 10 series (which can be selected), this encoder 10 is similar to the one described above. Figure 12 The encoder is operated as explained. However, Figure 15 The encoder 10 is configured to encode the video content 16 into the video data stream 22 in units of sub-parts 24 of the image 18 of the video content 16, as described above. Figure 11 As explained. When encoding video content 16, encoder 10 is interested in transmitting information about the Region of Interest (ROI) 60 along with the video data stream 22 to the decoding side. ROI 60 is a spatial sub-region of the current image 18 that the decoder should pay particular attention to, for example. The spatial location of ROI 60 can be input externally, such as by user input (as illustrated by dashed line 62), or can be automatically determined in real time by encoder 10 or by some other entity during the encoding of the current image 18. In either case, encoder 10 faces the following problem: indicating the location of ROI 60 is not a problem for encoder 10 in principle. For this purpose, encoder 10 can easily indicate the location of ROI 60 within data stream 22. However, in order to make this information easily accessible, Figure 15 The encoder 10 uses ROI packets interleaved among the payload packets of the access unit, thus the encoder 10 can select, on an online basis, the number of sub-parts 24 and / or payload packets to be segmented into the current image 18, with the sub-parts 24 spatially outside and within the ROI 60 being packetized into these payload packets. Using interleaved ROI packets, any network entity can easily identify payload packets belonging to an ROI. On the other hand, if a removable packet type is used for such ROI packets, they may be easily ignored by any network entity.

[0218] Figure 15An example is shown in which ROI packets 64 are interspersed among payload packets 32 of an access unit 30. ROI packets 64 indicate where in the sequence 34 of video data stream 22 there is coded data related to (i.e., encoding) an ROI 60. How ROI packets 64 indicate the location of an ROI 60 can be implemented in a variety of ways. For example, the mere presence / appearance of an ROI packet 64 can indicate that coded data related to an ROI 60 is incorporated in one or more of the subsequent payload packets 32 that follow in the sequential order of sequence 34, i.e., that belong to the prefix payload packets. Alternatively, a syntax element within an ROI packet 64 can indicate whether one or more subsequent payload packets 32 are related to an ROI 60, i.e., at least partially encode an ROI 60. A large variety of variations also arise from possible variations in the "scope" of an individual ROI packet 64, i.e., the number of prefix payload packets that are prefixed by one ROI packet 64. For example, an indication of whether any coded data related to an ROI 60 is incorporated within one ROI packet can be related to all payload packets 32 that follow in the sequential order of sequence 34 up to the occurrence of the next ROI packet 64, or can be related only to the immediately following payload packets 32, i.e., the payload packets 32 that immediately follow an individual ROI packet 64 in the sequential order of sequence 34. In Figure 15 In the following, FIG. 66 exemplarily illustrates a case in which an ROI packet 64 indicates ROI relevance (i.e., incorporation of any coded data related to an ROI 60) or ROI irrelevance (i.e., absence of any coded data related to an ROI 60) for all payload packets 32 that occur downstream of an individual ROI packet 64 up to the occurrence of the next ROI packet 64 or the end of the current access unit 30, whichever occurs earlier along the sequence 34. In particular, Figure 15 In the following, FIG. 66 exemplarily illustrates a case in which an ROI packet 64 indicates ROI relevance (i.e., incorporation of any coded data related to an ROI 60) or ROI irrelevance (i.e., absence of any coded data related to an ROI 60) for all payload packets 32 that occur downstream of an individual ROI packet 64 up to the occurrence of the next ROI packet 64 or the end of the current access unit 30, whichever occurs earlier along the sequence 34. In particular,

[0219] Any network entity 68 receiving the video data stream 22 can utilize the indication of ROI relevance, as implemented by use of the ROI packets 64, to process ROI- relevant portions of, for example, the packet sequence 34 with higher priority than other portions of, for example, the packet sequence 34. Alternatively, the network entity 68 can use the ROI relevance information to perform other tasks related to the transmission of, for example, the video data stream 22. The network entity 68 can be, for example, a MANE or a decoder that is used to decode and play back the video content 60 conveyed via the video data stream 22. In other words, the network entity 68 can use the identification of the ROI packets to make decisions regarding transmission tasks related to the video data stream. Transmission tasks can include retransmission requests regarding defective packets. The network entity 68 can be configured to process the region of interest 70 with increased priority and assign higher priority to the ROI packets 72 that signal that they cover the region of interest and their associated payload packets (i.e., the payload packets that are prefixed by the ROI packets 72) than to the ROI packets that do not cover the region of interest and their associated payload packets. The network entity 68 can first request retransmission of the payload packets that are assigned higher priority before requesting any retransmission of the payload packets that are assigned lower priority.

[0220] Figure 15 Embodiments of the present application can readily be combined with the embodiments previously described with respect to Figures 12 to 14 For example, the ROI packets 64 mentioned above can also be SEI packets that contain a specific type of SEI message, i.e., ROI SEI packets. That is, the SEI packets can be, for example, timing control packets and at the same time ROI packets, i.e., in the case where individual SEI packets contain both timing control information as well as ROI indication information. Alternatively, the SEI packets can be one but not the other of timing control packets and ROI packets, or can be neither ROI packets nor timing control packets.

[0221] According to the embodiments shown in Figure 16 In the embodiments shown in Figure 16In the example, it is shown that the current picture 18 is subdivided into four tiles 70, which here are exemplarily formed by four quarters of the current picture 18. The subdivision of the current picture 18 into tiles 70 can be signaled, e.g., in a unit containing a sequence of pictures within a video data stream, such as in a VPS or SPS packet also interspersed in the sequence of packets 34. As will be described in more detail below, the tile subdivision of the current picture 18 can be a regular subdivision of the picture 18 in rows and columns of tiles. The number of rows and columns of tiles as well as the row width and column height can vary. In particular, the width and height of the rows / columns of tiles can be different for different columns and different rows, respectively. Figure 16 It is further shown that the sub-sections 24 are examples of slices of the picture 18. The slices 24 subdivide the picture 18. As will be outlined in more detail below, the subdivision of the picture 18 into slices 24 can be subject to constraints, under which each slice 24 can be completely contained within one single tile 70, or completely cover two or more tiles 70. Figure 16 It is exemplarily shown that the picture 18 is subdivided into five slices 24. The first four of these slices 24 cover the first two tiles 70 in the aforementioned decoding order, while the fifth slice completely covers the third and the third tile 70. Further, Figure 16 It is exemplarily shown that each slice 24 is encoded into a respective payload packet 32. Each tile identification packet 72, in turn, indicates for its immediately following payload packet 32 which of the sub-sections 24 encoded into this payload packet 32 covers which of the tiles 70. Thus, while the first two tile identification packets 72 indicate the first tile of the picture 18 in the access unit 30 related to the current picture 18, the third and fourth tile identification packets 72 indicate the second tile 70 of the picture 18, and the fifth tile identification packet 72 indicates the third and fourth tile 70. With respect to Figure 16 Embodiments of the present application, as described above, e.g., with respect to Figure 15 The same variations are possible as described above, e.g., with respect to

[0222] With respect to the tiles, the encoder 10 can be configured to encode each tile 70 such that no spatial prediction or context selection occurs across tile boundaries. The encoder 10 can, e.g., encode the tiles 70 in parallel. Likewise, any decoder, such as the network entity 68, can decode the tiles 70 in parallel.

[0223] The network entity 68 can be a MANE or a decoder or some other device between the encoder 10 and the decoder and can be configured to use the information conveyed by the tile identification packets 72 to make decisions for a specific transmission task. For example, the network entity 68 can treat a specific tile of a current picture 18 of the video 16 with higher priority, i.e. the individual payload packets indicated to be related to this tile can be sent earlier or with safer FEC protection or the like. In other words, the network entity 68 can use the identification results to make decisions for a transmission task related to the video data stream. The transmission task can include a retransmission request for packets received in defective state, i.e. with any FEC protection for the video data stream being overdone, if present. The network entity can treat different tiles 70, for example, with different priorities. To this end, the network entity can assign a higher priority to the tile identification packets 72 and their payload packets related to a higher priority tile (i.e. the payload packets prefixed by the tile identification packets 72) compared to the tile identification packets 72 and their payload packets related to a lower priority tile. The network entity 68 can request a retransmission of the payload packets assigned with a higher priority first, for example, before requesting any retransmission of the payload packets assigned with a lower priority.

[0224] The embodiments described so far can be constructed into the HEVC framework as described in the introductory part of the specification of this application as described below.

[0225] In particular, in the sub-picture CPB / HRD case, SEI messages can be assigned to the slices of a decoding unit. That is, buffering period and timing SEI messages can be assigned to the NAL units containing the slices of a decoding unit. This can be achieved by a new NAL unit type that is allowed as a non-VCL NAL unit immediately preceding one or more slice / VCL NAL units of a decoding unit. This new NAL unit can be referred to as a slice prefix NAL unit. Figure 17 An access unit structure is illustrated that omits any hypothetical NAL units for the end of a sequence and a stream.

[0226] According to Figure 17The access unit 30 is understood as follows: In the sequential order of the packets of the packet sequence 34, the access unit can start with the occurrence of a special type of packet, namely the access unit delimiter 80. Then, within the access unit 30, one or more SEI packets 82 of one of the SEI packet types relevant for the whole access unit can follow. Both packet types 80 and 82 are optional. That is, no packets of this type can occur within the access unit 30. Then, the sequence of decoding units 38 follows. Each decoding unit 38 optionally starts with a slice prefix NAL unit 84, which contains, for example, timing control information, or, depending on the Figure 15 Embodiments of the slice prefix NAL unit 84 according to the present application comprise the ROI information or tile information, or even more generally the individual sub-picture SEI messages 86 according to the embodiments of the present application of 0 or 16. Then, the actual slice data 88 in individual payload packets or VCL NAL units follows, as indicated in 88. Thus, each decoding unit 38 comprises a sequence of slice prefix NAL units 84, followed by individual slice data NAL units 88. Figure 17 The skip arrow 90 skipping the slice prefix NAL unit in the middle would indicate that, in case of no decoding unit subdivision of the current access unit 30, there can be no slice prefix NAL unit 84.

[0227] As already pointed out above, all information signaled in the slice prefix and relevant for the sub-picture SEI messages can be valid for all VCL NAL units in the access unit or up to the occurrence of the second prefix NAL unit, or valid for the following VCL-NAL units in decoding order, depending on the flags given in the slice prefix NAL unit.

[0228] The slice VCL NAL units for which the information signaled in the slice prefix is valid are referred to in the following as prefix slices. The prefix slices related to a single slice prefix do not necessarily constitute a complete decoding unit, but can be part of a decoding unit. However, a single slice prefix cannot be valid for multiple decoding units (sub-pictures) and the start of a decoding unit is signaled in the slice prefix. If the means for signaling is not given via the slice prefix syntax (as in the "simple syntax" / version 1 indicated in the following), the occurrence of a slice prefix NAL unit signals the start of a decoding unit. Only certain SEI messages (identified via payloadType in the following syntax description) can be sent uniquely within a slice prefix NAL unit on the sub-picture level, while some SEI messages can be sent in a slice prefix NAL unit on the sub-picture level or as regular SEI messages on the access unit level.

[0229] As already pointed out above, the slice prefix NAL unit 84 according to the present application comprises the ROI information or tile information, or even more generally the individual sub-picture SEI messages 86 according to the embodiments of the present application of 0 or 16. Figure 16As discussed, in addition or otherwise, tile ID SEI messages / tile ID signaling can be implemented in high level syntax. In earlier designs of HEVC, slice header / slice data contains an identifier for the tiles contained in the individual slice. For example, slice data semantics indicate:

[0230] tile_idx_minus_1 specifies the TileID in the order of raster scan. The first tile in the picture will have a TileID of 0. The value of tile_idx_minus_1 will be in the range of 0 to (num_tile_columns_minus1 + 1) * (num_tile_rows_minus1 + 1) - 1.

[0231] However, this parameter is not considered useful, as it can easily be derived from the slice address and slice size (as signaled in the picture parameter set) if tiles_or_entropy_coding_sync_idc is equal to 1.

[0232] Although the tile ID can be derived implicitly in the decoding process, it is also important to know this parameter on the application level for different use cases, such as in a video conferencing situation, where different tiles can have different playback priority (those tiles typically forming the region of interest, which contains the speaker in a conversational use case, can have a higher priority than other tiles. In case network packets are lost in the transmission of multiple tiles, those network packets containing tiles representing the region of interest can be retransmitted with higher priority in order to maintain the quality of experience at the receiver terminal higher than in case of retransmission without any priority order. Another use case can be the assignment of tiles (if their size and their position are known) to different screens, e.g. in a video conferencing situation.

[0233] In order to allow this application level to handle tiles with specific priority in a transmission situation, the tile_id can be provided as subpicture or slice specific SEI message, or in a special NAL unit in front of one or more NAL units of a tile or in a special header section of NAL units belonging to a tile.

[0234] As described above with respect to Figure 15 In addition or otherwise, a region of interest SEI message can also be provided. This SEI message can allow the signaling of a region of interest (ROI), in particular, the ROI to which a specific tile_id / tile belongs. The message can allow giving a region of interest ID plus a region of interest priority.

[0235] Figure 18 The use of tiles in region of interest signaling is exemplified.

[0236] In addition to what has been described above, fragment header signals can also be implemented. The fragment prefix NAL unit can also contain fragment headers for subsequent dependent fragments (i.e., fragments prefixed with individual fragment prefixes). If the fragment header is provided only in the fragment prefix NAL unit, the actual fragment type needs to be derived from the NAL unit type containing the individual dependent fragments or from a flag in the fragment prefix, which signals whether subsequent fragment data belongs to the fragment type that serves as a random access point.

[0237] Furthermore, the fragment prefix NAL unit can carry fragment or sub-image specific SEI messages to convey non-mandatory information, such as sub-image timing or tile identifiers. While the HEVC specification described in the introductory section of this application does not support the transmission of non-mandatory sub-image specific messages, they are crucial for certain applications.

[0238] The following describes techniques for implementing the tile prefixing concept outlined above. Specifically, it describes which changes are sufficient at the tile level when using the HEVC state as outlined in the introductory section of this application as a basis.

[0239] In detail, the following presents two versions of the possible fragment prefix syntax: one version has only SEI message passing functionality, and the other version has extended functionality by signaling a portion of the fragment header of the subsequent fragment. Figure 19 The first simple syntax / version 1 is shown in the text.

[0240] As a preliminary note Figure 19 Therefore, this demonstrates the method used to implement the above regarding Figures 11 to 16 Possible implementations of any of the described embodiments. The interleaved packets shown therein can be as follows: Figure 19 The above situation will be understood through the examples shown in the document, and will be described in more detail below using specific implementation examples.

[0241] exist Figure 20 The table provides the extended syntax / version 2, which includes tile_id signaling, decoder start identifier, fragment prefix ID and fragment header data, in addition to the SEI message concept technology.

[0242] The semantics can be defined as follows:

[0243] A rap_flag value of 1 indicates that an access cell containing a fragment prefix is ​​a RAP image. A rap_flag value of 0 indicates that an access cell containing a fragment prefix is ​​not a RAP image.

[0244] The decoding_unit_start_flag indicates the start of a decoding unit within an access unit, thus indicating that the subsequent slices until the end of the access unit or the start of another decoding unit belong to the same decoding unit.

[0245] The single_slice_flag with value 0 indicates that the information provided in the slice prefix NAL unit and the associated subpicture SEI message are valid for all subsequent VCL-NAL units until the start of the next access unit, the occurrence of another slice prefix, or another complete slice header. The single_slice_flag with value 1 indicates that the information provided in the slice prefix NAL unit and the associated subpicture SEI message are valid only for the next VCL-NAL unit in decoding order.

[0246] The tile_idc indicates the number of tiles that will be present in the subsequent slice. A tile_idc equal to 0 indicates that no tiles are used in the subsequent slice. A tile_idc equal to 1 indicates that a single tile is used in the subsequent slice and its tile identifier is signaled accordingly. A tile_idc with value 2 indicates that multiple tiles are used in the subsequent slice and the number of tiles and the first tile identifier are signaled accordingly.

[0247] The prefix_slice_header_data_present_flag indicates that data corresponding to the slice header of the slice following in decoding order is signaled in the given slice prefix.

[0248] slice_header_data() is defined later in this document. It contains the relevant slice header information that is not covered by the slice header if dependent_slice_flag is set equal to 1.

[0249] Note that decoupling the slice header from the actual slice data allows for more flexible transmission schemes of the header and the slice data.

[0250] The num_tiles_in_prefixed_slices_minus1 indicates the number of tiles used in the subsequent decoding unit minus 1.

[0251] The first_tile_id_in_prefixed_slices indicates the tile identifier of the first tile in the subsequent decoding unit.

[0252] For the simple syntax / Version 1 of the slice prefix, the following syntax elements can be set to the following default values (if not present):

[0253] decoding_unit_start is equal to 1, i.e., the slice prefix always indicates the start of a decoding unit.

[0254] The single_slice_flag is equal to 0, meaning that the fragment prefix is ​​valid for all tiles in the decoding unit.

[0255] The proposed fragment prefix NAL unit has NAL unit type 24 and the NAL unit type overview table will be based on Figure 21 To expand.

[0256] That is, to summarize briefly. Figures 19 to 21 The syntactic details shown indicate that a particular packet type can be considered to belong to the interleaved packet identified above, exemplified here as NAL unit type 24. Furthermore, especially... Figure 20 The syntactic examples demonstrate that the aforementioned alternatives to the "category" of interleaved packets, controlled by a switching mechanism governed by individual syntactic elements within these interleaved packets themselves (exemplarily single_slice_flag here), can be used to control this category, i.e., to switch between different alternatives defined for this category. Furthermore, it has been shown that extensibility... Figures 1 to 16 In the above embodiments, since the interleaved packets also include common segment header data of segments 24 contained in packets belonging to the "scope" of individual interleaved packets, there may be a mechanism controlled by individual flags within these interleaved packets that indicates whether individual interleaved packets contain common segment header data.

[0257] Of course, the conceptual technique just presented (which transforms a portion of the fragment header data into a fragment header prefix) requires changes to the fragment header as specified in the current version of HEVC. Figure 22 The table in the table shows the possible syntax for this fragment header, where specific syntax elements present in the fragment header according to the current version are transformed into lower-level syntax elements called slice_header_data(). This syntax and fragment header data for the fragment header are only applicable to the following options, according to which the extended fragment header prefix NAL unit concept technique is used.

[0258] Figure 22 In the code, slice_header_data_present_flag indicates that the value of the signal emitted in the last slice prefix NAL unit (i.e., the most recently appearing slice prefix NAL unit) in the self-access unit is used to predict the slice header data of the current slice.

[0259] via such Figure 23A , Figure 23B and Figure 23C The syntax element fragment header data given in the table is emitted by a signal from all syntax elements removed from the fragment header.

[0260] i.e. the concept of Figure 22 and Figure 23A , Figure 23B and Figure 23C is transferred to embodiments of Figures 12 to 16 wherein the described interleaved packets can be extended by the concept of incorporating a part of the slice header syntax of the segment (subpart) 24 encoded into the payload packets (i.e. VCL NAL units) into such interleaved packets. This incorporation can be optional. I.e. an individual syntax element in the interleaved packet can indicate whether the individual interleaved packet contains such slice header syntax or not. If incorporated, the individual slice header data incorporated into the individual interleaved packet can apply to all segments contained in the packets belonging to the "scope" of the individual interleaved packet. Whether the segments encoded into any of the payload packets belonging to the scope of the interleaved packet employ the slice header data contained in the interleaved packet can be signaled by an individual flag such as Figure 22 slice_header_data_present_flag of Figures 12 to 16 . By this measure, the slice header of the segments encoded into the packets belonging to the "scope" of the individual interleaved packet can be reduced accordingly using the just mentioned flag in the slice segment header, and any decoder receiving the video data stream such as the network entity shown above in

[0261] Further continuing the syntax example for implementing embodiments of Figures 12 to 16 , the SEI message syntax can be as shown in Figure 24 . For the introduction of the segment or subpicture SEI message type, the SEI payload syntax can be adapted as presented in the tables of Figure 25A and Figure 25B . Only SEI messages with payloadType in the range of 180 to 184 can be sent uniquely on the segment level in segment prefix NAL units. In addition, the region of interest SEI message with payloadType equal to 140 can be sent on the segment level in segment prefix NAL units or as regular SEI messages on the access unit level.

[0262] i.e. transferring the details shown in Figure 25A and Figure 25B and Figure 24 to the above regarding Figures 12 to 16The described embodiments can be implemented by using segment prefix NAL units with a specific NAL unit type (e.g. 24) Figures 12 to 16 The interleaved packets shown in these embodiments include a specific type of SEI message signaled by payloadType at the beginning of each SEI message, e.g. within a segment prefix NAL unit. In the specific syntax embodiments now described, payloadType = 180 and payloadType = 181 lead to timing control packets according to the embodiment of Figures 11 to 14 payloadType = 140 leads to ROI packets according to the embodiment of Figure 15 and payloadType = 182 leads to tile identification packets according to the embodiment of Figure 16 The specific syntax examples described below herein can only include one or a subset of the just mentioned payloadType options. Furthermore, Figure 25A and Figure 25B indicate that: Figures 11 to 16 Any of the above embodiments can be combined with each other. Further, Figure 25A and Figure 25B indicate that: Figures 12 to 16 Any of the above embodiments or any combination thereof can be extended by another interleaved packet (explained later by payloadType = 184). As already described above, the extension described below with respect to payloadType = 183 finally leads to the possibility that any interleaved packet can have incorporated in it common segment header data of segments of any payload packets belonging to its category.

[0263] The tables in the following figures define SEI messages that can be used on segment or subpicture level. Also presented are region of interest SEI messages that can be used on subpicture and access unit level.

[0264] For example, Figure 26 An example of a subpicture buffering SEI message is shown that occurs whenever a segment prefix NAL unit of NAL unit type 24 has SEI message type 180 included in it (thus forming a timing control packet).

[0265] The semantics can be defined as follows:

[0266] seq_parameter_set_id specifies the sequence parameter set containing the sequence HRD parameters. The value of seq_parameter_set_id shall be equal to the value of seq_parameter_set_id in the picture parameter set referred to by the primary coded picture associated with the buffering period SEI message. The value of seq_parameter_set_id shall be in the range of 0 to 31, inclusive.

[0267] initial_cpb_removal_delay[SchedSelIdx] and initial_alt_cpb_removal_delay[SchedSelIdx] specify the initial CPB removal delay for the SchedSelIdx-th CPB of the decoding unit (sub-picture). These syntax elements have a length in bits given by initial_cpb_removal_delay_length_minus1 + 1 and are in units of 90 kHz clock. The value of these syntax elements shall not be equal to 0 and shall not exceed 90000 * (CpbSize[SchedSelIdx] ÷ BitRate[SchedSelIdx]), which is the time equivalent of the CPB size in units of 90 kHz clock.

[0268] The sum of initial_cpb_removal_delay[SchedSelIdx] and initial_cpb_removal_delay_offset[SchedSelIdx] and the sum of initial_alt_cpb_removal_delay[SchedSelIdx] and initial_alt_cpb_removal_delay_offset[SchedSelIdx] shall be constant over the entire coded video sequence for each value of SchedSelIdx.

[0269] Figure 27 An example of a sub-picture timing SEI message is also shown, where the semantics can be described as follows:

[0270] du_cpb_removal_delay specifies how many clock ticks to wait after the removal of the cpb of the decoding unit (subpicture) related to the last subpicture buffering period SEI message (if present) in the same decoding unit (subpicture), or else the last buffering period SEI message in the previous access unit, before removing the decoding unit (subpicture) data related to the subpicture timing SEI message from the buffer. This value is also used to calculate the earliest possible time of arrival of the decoding unit (subpicture) data in the CPB for the HSS (Hypothetical Stream Scheduler [2]0). This syntax element is a fixed length code whose length in bits is given by cpb_removal_delay_length_minus1 + 1. cpb_removal_delay is modulo 2 (cpb_removal_delay_length_minus1+1) The remainder of the counter.

[0271] du_dpb_output_delay is used to calculate the dpb output time of a decoding unit (subpicture). It specifies how many clock ticks to wait after the removal of the decoded decoding unit (subpicture) from the CPB before outputting the decoding unit (subpicture) from the DPB.

[0272] Note that this allows subpicture updates. In this case, the non-updated decoding unit can keep the last decoded picture unchanged, i.e. it remains visible.

[0273] SUMMARY Figure 26 and Figure 27 and transferring specific details therein to the Figures 12 to 14 embodiments, it can be said that the decoder buffer retrieval times of decoding units can be signaled in a differentially encoded manner (i.e., in delta fashion) relative to another decoder buffer retrieval time in the relevant timing control packet. That is, to obtain the decoder buffer retrieval time of a particular decoding unit, a decoder receiving the video data stream adds the decoder retrieval time obtained from the timing control packet for the particular decoding unit, prefixed by a plus sign, to the decoder retrieval time of the immediately preceding decoding unit (i.e., the decoding unit preceding the particular decoding unit), and continues in this manner for subsequent decoding units. At the beginning of an encoded video sequence of several images (each or portions thereof), the timing control packet can additionally or otherwise contain decoder buffer retrieval time values that are independently encoded rather than differentially encoded relative to the decoder buffer retrieval time of any preceding decoding unit.

[0274] Figure 28 An example of how a subpicture segment information SEI message can look like is shown. The semantics can be defined as follows:

[0275] slice_header_data_flag with value equal to 1 specifies that slice header data is present in the SEI message. The slice header data provided in the SEI is valid for all slices following in decoding order until the end of the access unit, another SEI message, slice NAL unit, or slice prefix NAL unit in which slice data occurs.

[0276] Figure 29 An example of a sub-picture tile information SEI message is shown, where the semantics can be defined as follows:

[0277] tile_priority specifies the priority of all tiles in the prefixed slice following in decoding order. The value of tile_priority shall be in the range of 0 to 7, inclusive, where 7 specifies the highest priority.

[0278] multiple_tiles_in_prefixed_slices_flag with value equal to 1 specifies that there is more than one tile in the prefixed slice following in decoding order. multiple_tiles_in_prefixed_slices_flag with value equal to 0 specifies that the following prefixed slice contains only one tile.

[0279] num_tiles_in_prefixed_slices_minus1 specifies the number of tiles in the prefixed slice following in decoding order.

[0280] first_tile_id_in_prefixed_slices specifies the tile id of the first tile in the prefixed slice following in decoding order.

[0281] That is, the tile identification package mentioned in Figure 29 can be implemented using the syntax of Figure 16 Embodiments of Figure 16 can be implemented. As Figure 29As shown in the middle, a specific flag (here: multiple_tiles_in_prefixed_slices_flag) can be used to signal within a punctured tile identification package: whether any sub-portion of the current picture 18 encoded into any of the payload packages belonging to the scope of the individual punctured tile identification package covers only one tile or more than one tile. If the flag signals coverage to more than one tile, another syntax element, here exemplarily num_tiles_in_prefixed_slices_minus1, is contained in the individual punctured package, which indicates the number of tiles covered by any sub-portion of any payload package belonging to the scope of the individual punctured tile identification package. Finally, another syntax element, here exemplarily first_tile_id_in_prefixed_slices, indicates the ID of the tile being the first one according to the decoding order among the tiles indicated by the current punctured tile identification package. The Figure 29 syntax of the embodiment of Figure 16 the prefix tile identification package 72 prefixed to the fifth payload package 32 can for example have all three syntax elements just discussed, with multiple_tiles_in_prefixed_slices_flag being set to 1, num_tiles_in_prefixed_slices_minus1 being set to 1, indicating thereby that two tiles belong to the current scope, and first_tile_id_in_prefixed_slices being set to 3, indicating that the run of tiles according to the decoding order belonging to the scope of the current tile identification package 72 starts with the third tile (having tile_id = 2).

[0282] Figure 29 It is also shown that the tile identification package 72 can also possibly indicate tile_priority, i.e. the priority of the tile belonging to its scope. Similar to the ROI pattern, the network entity 68 can use this priority information to control transmission tasks, such as requests for retransmission of specific payload packages.

[0283] Figure 30 An example of a syntax of a subpicture tile size information SEI message is shown, where the semantics can be defined as follows:

[0284] A multiple_tiles_in_prefixed_slices_flag value of 1 indicates that there is more than one tile in the following prefixed slice in decoding order. A multiple_tiles_in_prefixed_slices_flag value of 0 indicates that the following prefixed slice contains only one tile.

[0285] num_tiles_in_prefixed_slices_minus1 indicates the number of tiles in the prefix segments that are appended in the order of decoding.

[0286] tile_horz_start[i] indicates the start of the i-th tile in the horizontal direction within the pixels of the image.

[0287] tile_width[i] indicates the width of the i-th tile in the image.

[0288] tile_vert_start[i] indicates the start of the i-th tile in the horizontal direction within the pixels of the image.

[0289] tile_height[i] indicates the height of the i-th tile in the image.

[0290] Please note that the size SEI information can be used in display operations, such as in multi-screen display situations, to assign tiles to screens.

[0291] Figure 30 Therefore, it indicates that the information about can be changed. Figure 16 The image recognition packet Figure 29 The implementation syntax example is as follows: tiles belonging to the scope of individual tile identification packets are indicated by their position within the current image 18 rather than their tile ID. That is, instead of signaling the tile ID of the first tile in the decoding order covered by individual sub-parts (which are encoded into any of the payload packets belonging to the scope of individual interleaved tile identification packets), for each tile belonging to the current tile identification packet, the position of the tile can be signaled by signaling, for example, the following: the top-left corner position of each tile i (exemplarily here by tile_horz_start and tile_vert_start), and the width and height of tile i (exemplarily here by tile_width and tile_height).

[0292] Syntax examples of SEI messages in the attention zone are shown in Figure 31 In order to be even more precise, Figure 32 This demonstrates the first variant. Specifically, one or more regions of interest (SEIs) can be signaled using SEI messages, for example, at the access unit level or at the sub-image level. According to... Figure 32 The first variant signals each ROI SEI message once for an individual ROI, instead of signaling all ROIs within the scope of an individual ROI packet in a single ROI SEI message (if multiple ROIs are within the current scope).

[0293] According to Figure 31 , the region of interest SEI message signals each ROI, respectively. The semantics can be defined as follows:

[0294] roi id indicates an identifier of the region of interest.

[0295] roi priority indicates the priority of all tiles belonging to the region of interest in the prefix slice or all slices following in decoding order, depending on whether the SEI message is sent on subpicture level or access unit level. The value of roi priority will be in the range of 0 to 7, inclusive, where 7 indicates the highest priority. If both roi priority in the given roi information SEI message and tile priority in the subpicture tile information SEI message are given, the highest value of the two is valid for the priority of the individual tile.

[0296] num tiles in roi minusl indicates the number of tiles belonging to the region of interest in the prefix slice following in decoding order.

[0297] roi tile id[i] indicates the tile id of the i-th tile belonging to the region of interest in the prefix slice following in decoding order.

[0298] That is, Figure 31 It is shown that the ROI packet as shown in Figure 15 may signal the ID of the region of interest in which the individual ROI packet and the payload packets belonging to its scope are about. Optionally, an ROI priority index can be signaled together with the ROI ID. However, both syntax elements are optional. Then, the syntax element num tiles in roi minusl can indicate the number of tiles belonging to the individual ROI 60 within the scope of the individual ROI packet. Then, roi tile id indicates the tile-ID of the i-th tile belonging to the ROI 60. For example, imagine that the picture 18 is subdivided into tiles 70 in the manner shown in Figure 16 and that the ROI 60 of Figure 15 corresponds to the left half of the picture 18, which would be formed by the first and third tiles in decoding order. Then, the ROI packet would be in Figure 16The first ROI packet can be placed in front of the first payload packet 32 of the access unit 30, followed by another ROI packet in between the fourth and fifth payload packets 32 of this access unit 30. Then, the first ROI packet will have num tile in roi minusl set to 0 and roi tile id[0] to 0 (further referring to the first tile in decoding order), whereas the second ROI packet in front of the fifth payload packet 32 will have num tile in roi minusl set to 0 and roi tile id[0] set to 2 (further referring to the third tile in decoding order, which is located in the lower left quarter of the picture 18).

[0299] According to a second variant, the syntax of the region of interest SEI message can be as shown in Figure 32 Here, all ROIs in a single SEI message are signaled. In detail, the same syntax as discussed above with respect to Figure 31 will be used, but the syntax elements for each of the several ROIs (for which the individual ROI SEI messages or ROI packets are about) will be multiplied by one number signaled by a syntax element (here exemplarily num rois minusl). Optionally, another syntax element (here exemplarily roi presentation on separate screen) can be signaled for each ROI whether the individual ROI is suitable for being presented on a separate screen or not.

[0300] The semantics can be as follows:

[0301] num rois minusl indicates the number of ROIs in the following prefix slice or regular slice in decoding order.

[0302] roi id[i] indicates the identifier of the i-th region of interest.

[0303] roi priority[i] indicates the priority of all tiles belonging to the i-th region of interest in the following prefix slice or all slices in decoding order, depending on whether the SEI message is sent on subpicture level or access unit level. The value of roi priority will be in the range of 0 to 7, inclusive, where 7 indicates the highest priority. If both ROI priority in the roi info SEI message and tile priority in the subpicture tile information SEI message are given, the highest value of the two is valid for the priority of the individual tile.

[0304] num_tiles_in_roi_minus1[ i ] indicates the number of tiles belonging to the i-th region of interest in the prefix slice that follows in decoding order.

[0305] roi_tile_id[ i ][ n ] indicates the tile id of the n-th tile belonging to the i-th region of interest in the prefix slice that follows in decoding order.

[0306] roi_presentation_on_seperate_screen[ i ] indicates that the region of interest associated with the i-th roi_id is suitable for presentation on a separate screen.

[0307] Thus, to briefly summarize the various embodiments described so far, an additional high-level syntax signaling strategy has been proposed that allows for the application of SEI messages as well as additional high-level syntax items beyond those included in the NAL unit header at the slice level. Thus, a slice prefix NAL unit has been described. Syntax and semantics for slice prefix and slice level / sub-picture SEI messages have been described, along with use cases for low-delay / sub-picture CPB operation, tile signaling, and ROI signaling. Extended syntax has been presented to additionally signal portions of the slice header of the following slice in the slice prefix.

[0308] For completeness, Figure 33 Another example of syntax is shown that can be used according to Figures 12 to 14 Timing control packet for embodiments of The semantics can be:

[0309] du_spt_cpb_removal_delay_increment specifies the duration, in clock ticks, of the nominal CPB time of the last decoding unit in the current access unit in decoding order and the nominal CPB time of the decoding unit associated with the decoding unit information SEI message. This value is also used to calculate the earliest possible time of arrival of the decoding unit data in the CPB for the HSS. This syntax element is represented by a fixed length code, the length of which in bits is given by du_cpb_removal_delay_increment_length_minus1 + 1. When the decoding unit associated with the decoding unit information SEI message is the last decoding unit in the current access unit, the value of du_spt_cpb_removal_delay_increment shall be equal to 0.

[0310] dpb_output_du_delay_present_flag equal to 1 specifies that the pic_spt_dpb_output_du_delay syntax element is present in the decoding unit information SEI message. dpb_output_du_delay_present_flag equal to 0 specifies that the pic_spt_dpb_output_du_delay syntax element is not present in the decoding unit information SEI message.

[0311] pic_spt_dpb_output_du_delay is used to calculate the dpb output time of a picture when SubPicHrdFlag is equal to 1. It specifies how many sub-clock ticks to wait after removing the last decoding unit of an access unit from the CPB before outputting the decoded picture from the DPB. When not present, the value of pic_spt_dpb_output_du_delay is inferred to be equal to pic_dpb_output_du_delay. The length of the syntax element pic_spt_dpb_output_du_delay is given in bits by dpb_output_delay_du_length_minus1 + 1.

[0312] It is a requirement of bitstream conformance that all decoding unit information SEI messages related to the same access unit, applicable to the same operation point and having dpb_output_du_delay_present_flag equal to 1 shall have the same value of pic_spt_dpb_output_du_delay. The output time derived from pic_spt_dpb_output_du_delay of any picture output by a decoder conforming to the output timing shall precede the output time derived from pic_spt_dpb_output_du_delay of all pictures in any subsequent CVS in decoding order.

[0313] The picture output order determined by the value of this syntax element shall be the same as the order determined by the value of PicOrderCntVal.

[0314] For pictures that are not output by the "bumping" process (because they precede an IRAP picture with no_output_of_prior_pics_flag equal to 1 or inferred to be equal to 1 and NoRaslOutputFlag equal to 1 in decoding order), the output times derived by pic_spt_dpb_output_du_delay shall increase with increasing PicOrderCntVal values for all pictures within the same CVS. The difference between the output times of any two pictures in a CVS shall be the same when SubPicHrdFlag is equal to 1 as when SubPicHrdFlag is equal to 0.

[0315] Furthermore, Figure 34 Another example of signaling the ROI region using a ROI packet is shown. According to Figure 34 The syntax of the ROI packet contains only one flag that indicates whether all sub- portions of the picture 18 coded into any payload packet 32 belonging to its scope belong to the ROI. The "scope" extends until the ROI packet or the region_refresh_info SEI message occurs. If the flag is 1, it indicates that the region is coded into individual subsequent payload packets and if it is 0, the opposite is the case, i.e. individual sub-portions of the picture 18 do not belong to the ROI 60.

[0316] Before discussing some of the above embodiments again (in other words, explaining some of the words used above, such as tiles, segments and WPP sub-stream subdivision) or in other words, explaining some of the words used above, such as tiles, segments and WPP sub-stream subdivision), it should be noted that the high-level signaling of the above embodiments can alternatively be defined in a transport specification such as [3-7]. In other words, the packets mentioned above and forming the sequence 34 can be transport packets, some of which have application layer sub-portions (such as segments) incorporated into them (such as fully packetized or split), some of which are the above-mentioned and discussed interlaced between the latter in the manner and for the purpose discussed above. In other words, the above-mentioned interlaced packets are not limited to being SEI messages defined as other types of NAL units defined in the video codec of the application layer, but can alternatively be additional transport packets defined in the transport protocol.

[0317] In other words, according to one aspect of this specification, the above embodiments disclose a video data stream having video content encoded in units of sub-parts of images of video content (see coding tree blocks or segments), each sub-part being encoded in one or more payload packets (see VCL NAL units) of a sequence of packets (NAL units) of the video data stream, the packet sequence being divided into a sequence of access units, such that each access unit collects payload packets related to individual images of the video content, wherein the packet sequence has timing control packets (segment prefixes) interspersed therein, such timing control packets subdivide access units into decoding units, thus at least some access units are subdivided into two or more decoding units, wherein each timing control packet signals the decoder buffer fetch time of the decoding unit, and its payload packet follows the individual timing control packets in the packet sequence.

[0318] As described above, encoding video content into the data stream in units of sub-parts of an image can encompass syntax elements related to predictive coding, such as coding patterns (e.g., intra-frame patterns, inter-frame patterns, segmentation information, and the like), prediction parameters (e.g., motion vectors, extrapolation directions, or the like), and / or residual data (e.g., transformation coefficient levels), wherein these syntax elements are respectively related to local parts of the image (e.g., coding tree blocks, prediction blocks, and residual (e.g., transformation) blocks).

[0319] As described above, the payload packet may each contain one or more segments (each complete). These segments may be independently decodeable or may exhibit interrelationships that would prevent their independent decoding. For example, an entropy segment may be independently entropy-decoded but may prevent prediction beyond the segment boundaries. Dependency segments may allow WPP processing, i.e., encoding / decoding using entropy and prediction coding beyond the segment boundaries, with the ability to encode / decode dependency segments in parallel in a temporally overlapping manner, but interleaving the encoding / decoding procedures of individual dependency segments and the segments / segments referenced by the dependency segments.

[0320] The decoder can know in advance the order in which the payload packets of the access units are arranged within individual access units. For example, the encoding / decoding order can be defined between sub-parts of an image, such as the scanning order between blocks in the coding tree in the example above.

[0321] For example, see the diagram below. The currently encoded / decoded image 100 can be divided into tiles, such tiles, for example in... Figure 35 and Figure 36 The four quarters of image 110 are exemplary and indicated by reference numerals 112a-112d. That is, the entire image 110 can form a single block, as shown in... Figure 37In case of a picture 110 that is not divisible by the tile size, or in case of a picture 110 that is not divisible by the tile size, the picture 110 can be segmented into more than one tile. The tile segmentation can be limited to a regular segmentation, where tiles are arranged only by rows and columns. Different examples are presented below.

[0322] As can be seen, the picture 110 is further subdivided into coding (tree) blocks (small squares in the figure, and referred to as CTBs above) 114, between which a coding order 116 is defined (here, a raster scan order, but it can be different). The subdivision of the picture into tiles 112a-d can be restricted such that a tile is a disjoint set of blocks 114. Also, both blocks 114 and tiles 112a-d can be limited to a regular arrangement by rows and columns.

[0323] If tiles (i.e., more than one) exist, the coding (decoding) order 116 first performs a raster scan over the first complete tile, and then transitions to the next tile in tile order also in raster scan tile order.

[0324] Since tiles can be coded / decoded independently from each other due to the tile boundaries being disjoint, this coding / decoding is by spatial prediction and context selection inferred from spatial neighbors, so the encoder 10 and the decoder 12 can be coded / decoded in parallel independently from each other, except for, e.g., in-loop or post-filtering (which can be allowed to intersect with tile boundaries).

[0325] The picture 110 can be further subdivided into segments 118a-d, 180 (previously indicated by reference sign 24). A segment can contain only a part of a tile, one complete tile, or more than one complete tile. Thus, segmentation into segments can also be a subdivision of tiles, as in Figure 35 the case. Each segment contains at least one complete coding block 114 and consists of consecutive coding blocks 114 in coding order 116, thus defining an order between segments 118a-d, where the index in the figure is assigned following this order. The segment partitioning in Figures 35 to 37 is chosen for illustration purposes only. Tile boundaries can be signaled in the data stream. The picture 110 can form a single tile, as depicted in Figure 37 .

[0326] The encoder 10 and the decoder 12 can be configured to respect tile boundaries, since spatial prediction is not applied across tile boundaries. Context adaptation (i.e., probability adaptation) for various entropy (arithmetic) contexts can continue over the whole tile. However, as soon as a segment intersects with a tile boundary (if present inside the segment), such as in Figure 36With respect to slices 118a, 118b, the slices are further subdivided into sub-slices (sub-streams or tiles), where the slice contains an index (i.e., entry_point_offset) to the beginning of each sub-slice. In a decoder loop, filters can be allowed to cross tile boundaries. Such filters can involve one or more of a de-blocking filter, a sample adaptive offset (SAO) filter, and an adaptive loop filter (ALF). The latter can be applied on tile / slice boundaries if enabled.

[0327] Each optional second and subsequent sub-slice can have its beginning positioned within the slice in byte-aligned fashion, with the index indicating an offset from the beginning of one sub-slice to the beginning of the next sub-slice. The sub-slices are arranged in the slice in scan order 116. Figure 38 It is shown that Figure 37 slice 180c is subdivided into sub-slices 119 i of the example.

[0328] With respect to the figures, note that the tiles forming the sub- portions of a slice need not end in a column in tile 112a. For example, see slice 118a in Figure 37 and Figure 38 .

[0329] The following figure shows an exemplary portion of a data stream with respect to an access unit related to picture 110 of the above Figure 38 . Here, each payload packet 122a-d (previously indicated by reference number 32) is exemplarily applied to only one slice 118a. Two timing control packets 124a, 124b (previously indicated by reference number 36) are shown to be interleaved in access unit 120 for illustration purposes: 124a precedes packet 122a in packet order 126 (corresponding to a decoding / encoding time axis), and 124b precedes packet 122c. Thus, access unit 120 is divided into two decoding units 128a, 128b (previously indicated by reference number 38), where the first decoding unit contains packets 122a, 122b (and optionally a filler data packet (after the first packet 122a and the second packet 122b, respectively) and an optional access unit leading SEI packet (before the first packet 122a)) and the second decoding unit contains packets 118c, 118d (and optionally filler data (after packets 122c, 122d, respectively)).

[0330] As mentioned above, each packet of a packet sequence can be assigned only one of a plurality of packet types (nal_unit_type). Payload packets and timing control packets (and optional filler data and SEI packets) are, for example, different packet types. The instantiation of packets of a particular packet type in a packet sequence can be subject to certain restrictions. Such restrictions can define an order among the packet types (see Figure 17 ), the packets within each access unit will adhere to this order so that the access unit boundaries 130a, 130b can be detected and remain at the same position within the packet sequence even if any packets of a removable packet type are removed from the video data stream. Payload packets, for example, are not removable packet types. However, timing control packets, filler data packets and SEI packets, as discussed above, can be removable packet types, i.e. these packets can be non-VCL NAL units.

[0331] In the above example, timing control packets have been explicitly illustrated above by the syntax of slice_prefix_rbsp ().

[0332] Using this interleaving of timing control packets, the encoder is enabled to adjust the buffer scheduling at the decoder side during the encoding of individual pictures of the video content. For example, the encoder is enabled to optimize the buffer scheduling to minimize the end-to-end delay. In this regard, the encoder is enabled to take into account the individual distribution of encoding complexity across the picture regions of the video content for individual pictures of the video content. In detail, the encoder can continuously output a sequence of packets 122, 122a-d, 122a-d on a packet-by-packet basis (i.e. as soon as a current packet is finished, it is output). By using timing control packets, the encoder is enabled to adjust the buffer scheduling at the decoder side at a time when some of the sub-portions of a current picture have already been encoded into individual payload packets, but the remaining sub-portions have not yet been encoded. 1-3

[0333] Thus, an encoder for encoding video content into a video data stream in units of sub-portions (see coding tree blocks, tiles or slices) of pictures of the video content, wherein each sub-portion is encoded into one or more payload packets (see VCL NAL units) of a sequence of packets (NAL units) of the video data stream, respectively, so that the sequence of packets is divided into a sequence of access units and each access unit collects the payload packets related to an individual picture of the video content, can be configured to interleave timing control packets (slice prefixes) in the sequence of packets, so that the timing control packets subdivide the access units into decoding units, so that at least some of the access units are subdivided into two or more decoding units, wherein each timing control packet signals a decoder buffer retrieval time of a decoding unit, the payload packets of which follow the individual timing control packet in the sequence of packets. ​

[0334] Any decoder receiving the video data stream just outlined can make use of or not make use of the scheduling information contained in the timing control packets. However, while the decoder can make use of this information, a decoder conforming to the codec level must be able to decode the data following the indicated timing. If use is made, the decoder feeds its decoder buffer and empties its decoder buffer in decoding unit units. As outlined above, a "decoder buffer" can relate to a decoded picture buffer and / or a coded picture buffer.

[0335] Thus, a decoder for decoding a video data stream having video content coded into it in units of sub-portions of pictures of the video content (see coding tree blocks, tiles or slices), wherein each sub-portion is coded into one or more payload packets (see VCL NAL units) of a sequence of packets (NAL units) of the video data stream, thus the sequence of packets being divided into a sequence of access units and each access unit collecting the payload packets related to an individual picture of the video content, can be configured to look for timing control packets interspersed in the sequence of packets, at which timing control packets the access units are subdivided into decoding units, thus at least some access units being subdivided into two or more decoding units, derive from the decoding unit's decoder buffer retrieval time of each timing control packet, whose payload packets follow in the sequence of packets behind the individual timing control packet, and retrieve the decoding units from the decoder's buffer at the time scheduled by the decoding unit's decoder buffer retrieval time.

[0336] Looking for timing control packets can involve the decoder checking the NAL unit header and its contained syntax element, namely nal_unit_type. If the value of the latter flag equals a certain value, namely 124 according to the above example, the currently checked packet is a timing control packet. That is, a timing control packet can contain or convey the information explained above with respect to the pseudo code subpic_buffering and subpic_timing. That is, a timing control packet can convey or specify the initial CPB removal delay of the decoder or specify how many clock ticks to wait after removing an individual decoding unit from the CPB.

[0337] In order to allow repeated transmission of timing control packets without inadvertently further dividing access units into other decoding units, a flag within the timing control packet can explicitly signal whether the current timing control packet participates in subdividing an access unit into decoding units (compare decoding_unit_start_flag = 1 indicating the start of a decoding unit and decoding_unit_start_flag = 0 signaling the opposite).

[0338] The use of tile identification information in connection with interlaced decoding units differs from the use of timing control packets in connection with interlaced decoding units in that the tile identification information is interlaced in the data stream. The timing control packets mentioned above can additionally be interlaced in the data stream or the decoder buffer retrieval time can be communicated together with the tile identification information explained below in the same packet. Thus, the details presented in the above sections can be used to clarify the issues in the following description.

[0339] Another aspect of the present specification, which can be derived from the above embodiments, discloses a video data stream, having video content encoded therein using prediction and entropy coding, in units of segments into which pictures of the video content are spatially subdivided, using an encoding order between the segments, wherein the prediction coding and / or the entropy coding is limited to within tiles into which pictures of the video content are spatially subdivided, wherein a sequence of segments in the encoding order is packetized into payload packets of a sequence of packets (NAL units) of the video data stream, the sequence of packets being divided into a sequence of access units, thus each access unit collecting payload packets related to a respective picture of the video content, the payload packets having segments packetized therein, wherein the sequence of packets has tile identification packets interlaced therein, the tile identification packets identifying tiles, possibly only one, covered by segments, possibly only one, which are packetized into one or more payload packets immediately following the respective tile identification packet in the sequence of packets.

[0340] For example, see the diagram immediately preceding showing the data stream. Packets 124a and 124b will now represent tile identification packets. By explicit signaling (compare single_slice_flag = 1) or by convention, a tile identification packet can identify only the tile covered by the segment packetized into the immediately following payload packet 122a. Alternatively, by explicit signaling or by convention, tile identification packet 124a can identify the tiles covered by the segments packetized into one or more payload packets immediately following the respective tile identification packet 124a in the sequence of packets up to the earlier of the end 130b of the current access unit 120 and the beginning of the respective decoding unit 128b, respectively. For example, see Figure 35 , if each segment 118a-d 1-3 is packetized into a separate packet 122a-d 1-3 , wherein the subdivision into decoding units is such that the packets are grouped into three decoding units according to {122a 1-3}, {122b 1-3}, and {122c 1-3 , 122d 1-3}, then the packets packetized into the third decoding unit {122c 1-3,122d 1-3} the fragments {118c 1-3 ,118d 1-3} would, for example, cover tiles 112c and 112d, and the corresponding fragment prefixes would indicate "c" and "d", respectively, i.e. that these tiles 112c and 112d.

[0341] Thus, the network entity referred to further below can use this explicit signaling or convention to correctly relate each tile identification packet to one or more payload packets immediately following the tile identification packet in the packet sequence. The way in which this relation can be signaled has been exemplarily described above by means of the pseudo code subpic_tile_info. This example can be modified in its essence. For example, the syntax element "tile_priority" can be omitted. Furthermore, the order between the syntax elements can be switched, and the description of possible bit lengths and descriptors for coding principles with respect to the syntax elements is merely illustrative.

[0342] The network entity receiving a video data stream, i.e. a video data stream having video content encoded therein using prediction and entropy coding, in units of segments into which pictures of the video content are spatially subdivided, using an encoding order between the segments, wherein the prediction coding and / or the entropy coding is limited to be within tiles into which pictures of the video content are spatially subdivided, wherein a sequence of segments in the encoding order is packetized into payload packets of a sequence of packets (NAL units) of the video data stream, the packet sequence being divided into a sequence of access units, thus each access unit collecting payload packets related to a respective picture of the video content, the payload packets having tiles packetized therein, wherein the packet sequence has tile identification packets interspersed therein, can be configured to identify, based on the tile identification packets, tiles covered by segments, the tiles being packetized into one or more payload packets immediately following the respective tile identification packet in the packet sequence. The network entity can use the identification result for making decisions on transmission tasks. For example, the network entity can treat different tiles with different playing priorities. For example, in case of packet loss, it can be more willing to retransmit payload packets related to tiles with higher priority than payload packets related to tiles with lower priority. That is, the network entity can first request retransmission of lost payload packets related to tiles with higher priority. Only if there is enough time left (depending on the transmission rate), the network entity continues to request retransmission of lost payload packets related to tiles with lower priority. However, the network entity can also be capable of assigning tiles or payload packets related to a particular tile to different playing units of different screens.

[0343] As to the usage of the in-penetrated ROI information, it should be noted that the ROI packets mentioned in the following can coexist with the timing control packets and / or tile identification packets mentioned in the above, which are implemented by combining the information content of these packets in a common packet as described above with respect to the segment prefix, or in the form of separate packets.

[0344] In other words, using the usage of the in-penetrated ROI information as described above discloses a video data stream, having video content encoded using prediction and entropy coding, in units of segments into which images of the video content are spatially subdivided, using an encoding order between the segments, wherein the prediction of the predictive coding and / or the entropy coding is limited to within tiles into which images of the video content are divided, wherein a sequence of segments in the encoding order is packetized into payload packets of a sequence of packets (NAL units) of the video data stream, which sequence of packets is divided into a sequence of access units, thus each access unit collects payload packets related to an individual image of the video content, the payload packets having segments packetized therein, wherein the sequence of packets has ROI packets in-penetrated therein, the ROI packets identifying tiles of the image, the tiles respectively belonging to a ROI of the image.

[0345] As to the ROI packets, similar annotations as provided previously with respect to the tile identification packets are valid: A ROI packet can identify tiles of the image only among those tiles covered by the segments contained in one or more payload packets, the tiles belonging to a ROI of the image, the individual ROI packet pertaining to the one or more payload packets (by immediately preceding the one or more payload packets), as described above with respect to the "prefix segment".

[0346] A ROI packet can allow for more than one ROI to be identified per prefix segment, wherein for each of these ROIs (i.e. num rois minus 1) the relevant tiles are identified. Then, for each ROI, a priority can be transmitted, allowing for the ROIs to be ordered according to the priority (i.e. roi priority [i]). In order to allow for "tracking" of ROIs over time during a sequence of images of the video, the ROIs can be indexed with a ROI index, thus the ROIs indicated in the ROI packet to exceed / cross image boundaries (i.e. over time) can be related to each other (i.e. roi id [i]).

[0347] A network entity receiving a video data stream, i.e. a video data stream having video content encoded using prediction and entropy coding, in units of slices into which pictures of the video content are spatially subdivided, with prediction coding and / or entropy coding being restricted to be within tiles into which pictures of the video content are divided, while probability adaptation of entropy coding continues across slices, with a sequence of slices in coding order being packetized into payload packets of a sequence of packets (NAL units) of the video data stream, the sequence of packets being divided into a sequence of access units, thus each access unit collecting payload packets related to an individual picture of the video content, the payload packets having slices packetized therein, can be configured to identify, based on tile identification packets, packets to which slices continue to be packetized, the slices covering tiles belonging to a ROI of the picture.

[0348] The network entity can make use of the information conveyed by the ROI packets in a similar way as explained above in this previous section with respect to the tile identification packets.

[0349] With respect to the current section as well as the previous sections, it should be noted that any network entity such as a MANE or decoder is able to determine the tile covered by the slice of the currently inspected payload packet only by investigating the slice order of the slices of the picture and investigating the progress (with respect to the position of the tiles in the picture) of the current picture portion covered by the slices, which can have been explicitly signaled in the data stream as explained above or can be known to encoder and decoder according to convention. Alternatively, each slice (except the first slice in scanning order of a picture) can be provided with an indication / index of the same reference (same code) of the first coding block (e.g. CTB) (measured in units of coding tree blocks), slice_address, so that the decoder can place each slice (whose reconstruction) in the picture starting from the first coding block in the direction of the slice order. Thus, the above tile information packets only need to contain the index of the first tile covered by any slice of the one or more relevant payload packets immediately following the individual slice identification packet, first tile id in prefixed slices, because the network entity upon sequentially encountering the next tile identification packet is clear that if the index conveyed by the latter tile identification packet differs from the former tile identification packet by more than one, the payload packets between the two tile identification packets cover tiles having tile indices in between. This holds true in case, as mentioned, both the tile subdivision and the coding tree subdivision are based on a column-by-row subdivision, e.g. column-wise for both tiles and coding blocks, with a raster scan order defined therebetween, i.e. the tile indices increase in this raster scan order and the slices follow each other along this raster scan order between coding blocks according to the slice order.

[0350] The packetized and interleaved fragment header signaling pattern described in the above embodiments can also be combined with any one or any combination of the above patterns. The previously explicitly described fragment prefix (e.g., according to version 2) combines all such patterns. The advantage of this pattern is that it makes it easier for network entities to access fragment header data because it is conveyed in a separate packet located outside the prefix fragment / payload packet, and allows for repeated transmission of fragment header data.

[0351] Therefore, another aspect of this specification is the encapsulated and interspersed segment header signaling aspect, and in other words, it can be regarded as disclosing a video data stream having sub-parts of images of video content (see video content encoded in units of coding tree blocks or segments), each sub-part being encoded in one or more payload packets (see VCL NAL units) of a sequence of packets (NAL units) of the video data stream, the sequence of packets being divided into a sequence of access units, so that each access unit collects payload packets related to individual images of the video content, wherein the sequence of packets has interspersed segment header packets (segment prefixes), which convey segment header data for (and any missing segments) of one or more payload packets following individual segment header packets in the sequence of packets.

[0352] A network entity receiving a video data stream (i.e., a video data stream having video content encoded in units of sub-parts (see coding tree blocks or segments) of images of video content, each sub-part being encoded in one or more payload packets (see VCL NAL units) of a sequence of packets (NAL units) of the video data stream, the sequence of packets being divided into a sequence of access units, such that each access unit collects payload packets related to individual images of the video content, wherein the sequence of packets has segment header packets interspersed therein) can be configured to: read segment headers and payload data from the packets, however, wherein segment header data is derived from the segment header packets, and skips reading segment headers for one or more payload packets following individual segment header packets in the packet sequence, instead using segment headers derived from segment header packets following one or more payload packets.

[0353] As is true for the above-mentioned aspects, it is possible that the packet (here: slice header packet) can also have the functionality of indicating to any network entity such as a MANE or decoder the beginning of a decoding unit or the beginning of the run of one or more payload packets that are prefixed by the individual packet. Thus, a network entity according to the present aspects can identify payload packets for which the slice header has to be skipped based on the above syntax elements in this packet (i.e. single_slice_flag in combination with e.g. decoding_unit_start_flag, where the latter flag allows for retransmission of a copy of a particular slice header packet within a decoding unit as discussed above). This is useful, for example, because the slice headers of slices within one decoding unit can change along the slice sequence and, thus, while the slice header packet at the beginning of a decoding unit can have the decoding_unit_start_flag set (equal to 1), slice header packets positioned in between can have this flag not set in order to prevent any network entity from wrongly interpreting the occurrence of this slice header packet as the beginning of a new decoding unit.

[0354] Although some aspects have been described in the context of an apparatus, it is clear that such aspects also represent a description of the corresponding method, where a block or device corresponds to a method step or a feature of a method step. Analogously, aspects described in the context of a method step also represent a description of a corresponding block or item or feature of a corresponding apparatus. Some or all of the method steps can be performed by (or using) hardware, e.g. microprocessors, programmable computer systems or electronic circuitry. In some embodiments, one or more of the method steps

[0355] The encoded data stream of the present application can be stored on a digital storage medium or can be transmitted on a transmission medium such as a wireless transmission medium or a wired transmission medium such as the Internet.

[0356] Depending on certain implementation requirements, embodiments of the application can be implemented in hardware or in software. The implementation can be performed using a digital storage medium, for example a floppy disk, a DVD, a Blu-Ray, a CD, a ROM, a PROM, an EPROM, an EEPROM or a FLASH memory, having electronically readable control signals stored thereon, which cooperate with a programmable computer system such that individual steps of the method are performed. Therefore, the digital storage medium can be computer readable.

[0357] Some embodiments according to the application comprise a data carrier having electronically readable control signals, which are capable of cooperating with a programmable computer system, such that one of the methods described herein is performed.

[0358] Generally, embodiments of the present application can be implemented as a computer program product having a program code, the program code being operative for performing one of the methods when the computer program product is run on a computer. The program code may, for example, be stored on a machine-readable carrier.

[0359] Other embodiments comprise the computer program for performing one of the methods described herein, stored on a machine-readable carrier.

[0360] In other words, an embodiment of the inventive method is, therefore, a computer program having a program code for performing one of the methods described herein, when the computer program is run on a computer.

[0361] Another embodiment of the inventive method is, therefore, a data carrier (or a digital storage medium, or a computer-readable medium) comprising, recorded thereon, the computer program for performing one of the methods described herein. The data carrier, the digital storage medium or the recorded medium are typically tangible and / or non-transitionary.

[0362] A further embodiment of the inventive method is, therefore, a data stream or a sequence of signals representing the computer program for performing one of the methods described herein. The data stream or the sequence of signals can for example be configured to be transferred via a data communication connection, e.g. via the Internet.

[0363] A further embodiment comprises a processing means, such as a computer, or a programmable logic device, configured to or adapted for performing one of the methods described herein.

[0364] A further embodiment comprises a computer having installed thereon the computer program for performing one of the methods described herein.

[0365] A further embodiment comprises an apparatus or a system configured to transfer (for example, electronically or optically) a computer program for performing one of the methods described herein to a receiver. The receiver may, for example, be a computer, a mobile device, a memory device or the like. The apparatus or system may, for example, comprise a file server for transferring the computer program to the receiver.

[0366] In some embodiments, a programmable logic device (for example a field programmable gate array) can be used to perform some or all of the functionalities of the methods described herein. In some embodiments, a field programmable gate array can cooperate with a microprocessor in order to perform one of the methods described herein. Generally, the methods are preferably performed by any hardware device.

[0367] In some embodiments of the present application, the present application can also be configured in the following way:

[0368] 1. A video data stream having encoded therein video content (16) in units of sub- portions (24) of pictures (18) of the video content (16), each sub-portion (24) being encoded into one or more payload packets (32) of a packet sequence (34) of the video data stream (22), the packet sequence (34) being divided into a sequence of access units (30) so that each access unit (30) collects payload packets (32) pertaining to respective pictures (18) of the video content (16), wherein the packet sequence (34) has interspersed thereinto timing control packets (36) so that the timing control packets (36) subdivide the access units (30) into decoding units (38) so that at least some access units (30) are subdivided into two or more decoding units (38), wherein each timing control packet (38) signals a decoder buffer retrieval time for a decoding unit (38) whose payload packets (32) follow in the packet sequence (34) after the respective timing control packet (38).

[0369] 2. The video data stream according to clause 1, wherein the sub-portions (24) are tiles and each payload packet (32) contains one or more tiles.

[0370] 3. The video data stream according to clause 2, wherein the tiles contain independently decodable tiles and dependent slices allowing for entropy and prediction decoding beyond tile boundaries for WPP processing.

[0371] 4. The video data stream according to any of clauses 1 to 3, wherein each packet of the packet sequence (34) can be assigned only one of a plurality of packet types, wherein payload packets (32) and timing control packets (36) are different packet types, wherein the occurrence of packets of the plurality of packet types in the packet sequence (34) is subject to specific restrictions defining an order between the packet types, packets within each access unit (30) will adhere to the order, so that access unit boundaries can be detected using these restrictions by detecting cases of violation thereof, and access unit boundaries remain at the same positions within the packet sequence even if packets of any removable packet type are removed from the video data stream, wherein payload packets (32) are a non-removable packet type and timing control packets (36) are the removable packet type.

[0372] 5. The video data stream according to any of clauses 1 to 4, wherein each packet contains a packet type indicating a portion of syntax elements.

[0373] 6. The video data stream according to item 5, wherein the packet type indicating the part of syntax elements is contained in a packet type field in the header of each of the packets, the content of each of the packets differs between payload packets and timing control packets, and for timing control packets the SEI packet type field distinguishes between the timing control packets on the one hand and between different types of SEI packets on the other hand.

[0374] 7. An encoder for encoding video content (16) into a video data stream (22) in units of sub-portions (24) of pictures (18) of the video content (16), wherein each sub-portion (24) is encoded into one or more payload packets (32) of a sequence (34) of packets of the video data stream (22), whereby the sequence (34) of packets is divided into a sequence of access units (30), and each access unit (30) collects the payload packets (32) related to each picture (18) of the video content (16), the encoder being configured to intersperse timing control packets (36) in the sequence (34) of packets, whereby the timing control packets (36) subdivide the access units (30) into decoding units (38), whereby at least some access units (30) are subdivided into two or more decoding units (38), wherein each timing control packet (36) signals a decoder buffer retrieval time for a decoding unit (38), the payload packets (32) of which follow each of the timing control packets (36) in the sequence (34) of packets.

[0375] 8. The encoder according to item 7, wherein in the course of encoding a current picture of the video content,

[0376] a current sub-portion (24) of the current picture (18) is encoded into a current payload packet (32) of a current decoding unit (38),

[0377] at a first time instant, the current decoding unit (38) prefixed by a current timing control packet (36) is transmitted within the data stream by setting the decoder buffer retrieval time signaled by the current timing control packet (36); and

[0378] at a second time instant, another sub-portion of the current picture is encoded, the second time instant being later than the first time instant.

[0379] 9. A method for encoding video content (16) into a video data stream (22) in units of sub-portions (24) of pictures (18) of the video content (16), wherein each sub-portion (24) is encoded into one or more payload packets (32) of a packet sequence (34) of the video data stream (22), whereby the packet sequence (34) is divided into a sequence of access units (30) and each access unit (30) collects the payload packets (32) related to a respective picture (18) of the video content (16), the method comprising interspersing timing control packets (36) in the packet sequence (34), whereby the timing control packets (36) subdivide the access units (30) into decoding units (38), whereby at least some access units (30) are subdivided into two or more decoding units (38), wherein each timing control packet (36) signals a decoder buffer retrieval time for a decoding unit (38) whose payload packets (32) follow in the packet sequence (34) after the respective timing control packet (36).

[0380] 10. A decoder for decoding a video data stream (22) having video content (16) encoded therein in units of sub-portions (24) of pictures (18) of the video content (16), wherein each sub-portion is encoded into one or more payload packets (32) of a packet sequence (34) of the video data stream (22), the packet sequence (34) being divided into a sequence of access units (30) whereby each access unit (30) collects the payload packets (32) related to a respective picture (18) of the video content (16), the decoder comprising a buffer for buffering the video data stream or a reconstruction of the video content obtained therefrom by decoding the video data stream, and the decoder being configured to find timing control packets (36) interspersed in the packet sequence, at which the access units (30) are subdivided into decoding units (38), whereby at least some access units are subdivided into two or more decoding units, and to empty the buffer in units of the decoding units.

[0381] 11. The decoder according to item 10, wherein the decoder is configured to, when looking for the timing control packets (36), check in each packet for a packet type indicating a syntax element portion, and to recognize the respective packet as a timing control packet (36) if a value of the packet type indicating a syntax element portion is equal to a predetermined value.

[0382] 12. A method for decoding a video data stream (22), the video data stream (22) having video content (16) encoded therein in units of sub-portions (24) of pictures (18) of the video content (16), each sub-portion being encoded into one or more payload packets (32) of a sequence (34) of packets of the video data stream (22), the sequence (34) of packets being divided into a sequence of access units (30), thus each access unit (30) collecting payload packets (32) pertaining to respective pictures (18) of the video content (16), the method using a buffer to buffer the video data stream or a reconstruction of the video content obtained therefrom by decoding the video data stream, and the method comprising finding timing control packets (36) interspersed in the sequence of packets, subdividing the access units (30) into decoding units (38) at the timing control packets (36), thus at least some access units being subdivided into two or more decoding units, and emptying the buffer in units of the decoding units.

[0383] 13. A network entity for transmitting a video data stream (22), the video data stream (22) having video content (16) encoded therein in units of sub-portions (24) of pictures (18) of the video content (16), each sub-portion being encoded into one or more payload packets (32) of a sequence (34) of packets of the video data stream (22), the sequence (34) of packets being divided into a sequence of access units (30), thus each access unit (30) collecting payload packets (32) pertaining to respective pictures (18) of the video content (16), the decoder being configured to find timing control packets (36) interspersed in the sequence of packets, to subdivide the access units into decoding units at the timing control packets (36), thus at least some access units (30) being subdivided into two or more decoding units (38), to derive a decoding buffer retrieval time for a decoding unit (38) from each timing control packet (36), the payload packets (32) of the decoding unit following respective ones of the timing control packets (36) in the sequence (34) of packets, and to perform the transmission of the video data stream in dependence on the decoder buffer retrieval time for the decoding unit (38).

[0384] 14. A method for transmitting a video data stream (22), the video data stream (22) having video content (16) encoded therein in units of sub-portions (24) of pictures (18) of the video content (16), each sub-portion being encoded into one or more payload packets (32) of a packet sequence (34) of the video data stream (22), the packet sequence (34) being divided into a sequence of access units (30), thus each access unit (30) collecting payload packets (32) relating to respective pictures (18) of the video content (16), the method comprising finding timing control packets (36) interspersed in the packet sequence, subdividing the access units into decoding units at the timing control packets (36), thus at least some access units (30) being subdivided into two or more decoding units (38), deriving a decoder buffer retrieval time for a decoding unit (38) from each timing control packet (36), the decoding unit's payload packets (32) following the respective timing control packet (36) in the packet sequence (34), and performing the transmission of the video data stream in dependence on the decoder buffer retrieval times of the decoding units (38).

[0385] 15. A video data stream having video content (16) encoded therein in units of segments (24) using prediction and entropy encoding, pictures (18) of the video content (16) being ordered between the segments (24) using the encoding order, wherein the prediction encoding and / or entropy encoding is spatially subdivided into the segments (24) by limiting the prediction to within tiles (70), the pictures of the video content being spatially subdivided into tiles (70), wherein a sequence of the segments (24) is packetized into payload packets (32) of a packet sequence of the video data stream in the encoding order, the packet sequence (34) being divided into a sequence of access units (30), thus each access unit collecting the payload packets (32) having segments (24) relating to respective pictures (18) of the video content (16) packetized therein, wherein the packet sequence (34) has tile identification packets (72) interspersed between the payload packets (32) of one access unit, the tile identification packets identifying one or more tiles (70) covered by any segments (24) packetized into one or more payload packets (32) immediately following the respective tile identification packet (72) in the packet sequence (34).

[0386] 16. The video data stream according to clause 15, wherein the tile identification packets (72) identify the one or more tiles (70) covered by any segments (24) packetized into only the immediately following payload packets (32).

[0387] 17. The video data stream according to item 15, wherein the tile identification packet (72) identifies one or more tiles (70) covered by any segment (24) in one or more payload packets (32) immediately following the respective tile identification packet (72) in the packet sequence (34) up to the earlier of the end of the current access unit (30) and the next tile identification packet (72) in the packet sequence.

[0388] 18. A network entity configured to receive a video data stream according to any of items 15 to 16 and to identify, based on the tile identification packet (72), tiles (70) covered by segments (24) packetized into one or more payload packets (72) immediately following the respective tile identification packet (72) in the packet sequence.

[0389] 19. The network entity according to item 18, wherein the network entity is configured to use the result of the identification for making a decision on a transmission task related to the video data stream.

[0390] 20. The network entity according to item 18 or 19, wherein the transmission task comprises a retransmission request with respect to defective packets.

[0391] 21. The network entity according to item 18 or 19, wherein the network entity is configured to treat different tiles (70) with different priorities, which is achieved by assigning a higher priority to tile identification packets (72) and the payload packets immediately following the respective tile identification packets (72) having segments packetized therein covering tiles of a higher priority identified by the respective tile identification packets (72) than to tile identification packets (72) and the payload packets immediately following the respective tile identification packets (72) having segments packetized therein covering tiles of a lower priority identified by the respective tile identification packets.

[0392] 22. The network entity according to item 21, wherein the network entity is configured to first request retransmission of payload packets assigned with a higher priority before requesting any retransmission of payload packets assigned with a lower priority.

[0393] 23. A method comprising receiving a video data stream according to any of clauses 15 to 16 and identifying tiles (70) covered by segments (24) based on said tile identification packets (72), said segments (24) being packetized into one or more payload packets (72) immediately following each said tile identification packet (72) in said packet sequence.

[0394] 24. A video data stream having video content (16) encoded therein in units of sub-portions (24) of pictures (18) of said video content (16), each sub-portion (24) being encoded into one or more payload packets (32) of a packet sequence (34) of said video data stream (22), said packet sequence (34) being divided into a sequence of access units (30) so that each access unit (30) collects payload packets (32) relating to a respective picture (18) of said video content (16), wherein at least some access units (30) have said packet sequence (34) with ROI packets (64) interspersed therein, so that said timing control packets (36) subdivide said access units (30) into decoding units (38), so that at least some access units (30) have ROI packets interspersed between payload packets (32) relating to said picture of said respective access unit, wherein each ROI packet relates to one or more subsequent payload packets following each said ROI packet in said packet sequence (34) and identifies whether the sub-portion (24) encoded into any of said one or more payload packets relating to each said ROI packet covers a region of interest of said video content.

[0395] 25. A video data stream according to clause 24, wherein said sub-portions are segments and said video content is encoded into said video data stream using prediction and entropy encoding, wherein said prediction and / or entropy encoding is limited to within tiles into which said picture of said video content is divided, wherein each ROI packet also identifies a tile within which the sub-portion (24) encoded into any of said one or more payload packets relating to each said ROI packet covers said region of interest.

[0396] 26. A video data stream according to clause 24 or 25, wherein each ROI packet is uniquely related to said payload packet immediately following.

[0397] 27. The video data stream according to any of the items 24 or 25, wherein each ROI packet is related to all payload packets in the packet sequence immediately following each of the ROI packets up to the end of the access unit and the earlier of a respective ROI packet, each of the ROI packets being arranged within the access unit.

[0398] 28. A network entity configured to receive a video data stream according to any of the items 24 to 27 and to identify the ROIs of the video content based on the ROI packets.

[0399] 29. The network entity according to item 27, wherein the network entity is configured to use the result of the identification for making decisions on transmission tasks related to the video data stream.

[0400] 30. The network entity according to any of the items 28 or 29, wherein the transmission tasks comprise retransmission requests on defective packets.

[0401] 31. The network entity according to any of the items 28 or 29, wherein the network entity is configured to handle the region of interest with increased priority, which is achieved by assigning a higher priority to a ROI packet (72) and the one or more payload packets immediately following each of the ROI packets (72) in relation to which each of the ROI packets signals an overlap of the region of interest with the sub portion (24) encoded into any of the one or more payload packets in relation to which each of the ROI packets, than to a ROI packet and the one or more payload packets immediately following each of the ROI packets (72) in relation to which each of the ROI packets signals no overlap.

[0402] 32. The network entity according to item 31, wherein the network entity is configured to first request retransmission of payload packets assigned with a higher priority before requesting any retransmission of payload packets assigned with a lower priority.

[0403] 33. A method comprising receiving a video data stream according to any of the items 23 to 26 and identifying the ROIs of the video content based on the ROI packets.

[0404] 34. A computer program having a program code for performing the method according to any of the items 9, 12, 14, 23 or 33 when the computer program is run on a computer.

[0405] The embodiments described above are merely illustrative of the principles of the present application. It is understood that modifications and variations of the arrangements and details described herein will be apparent to others skilled in the art. It is the intent, therefore, to be limited only by the scope of the appended claims and not by the specific details presented by way of description and explanation of embodiments herein.

[0406] References

[0407] [1] Thomas Wiegand, Gary J. Sullivan, Gisle Bjontegaard, Ajay Luthra, "Overview of the H.264 / AVC Video Coding Standard", IEEE Trans. Circuits Syst. Video Technol., vol. 13, N7, July 2003.

[0408] [2] JCT-VC, "High-Efficiency Video Coding (HEVC) text specification Working Draft 7", JCTVC-I1003, May 2012.

[0409] [3] ISO / IEC 13818-1 : MPEG-2 Systems specification.

[0410] [4] IETF RFC 3550 - Real-time Transport Protocol.

[0411] [5] Y.-K. Wang et al., "RTP Payload Format for H.264 Video", IETF RFC6184, http: / / tools.ietf.org / html / rfc6184

[0412] [6] S. Wenger et al., "RTP Payload Format for Scalable Video Coding", IETF RFC 6190, http: / / tools.ietf.org / html / rfc6190

[0413] [6] T. Schierl et al., "RTP Payload Format for High Efficiency Video Coding", IETF internet draft, http: / / datatracker.ietf.org / doc / draft-schierl-payload-rtp-h265 / 。

Claims

1. A method for transmitting video content, the method comprising: transmitting a video data stream having the video content encoded therein in units of sub-portions of pictures of the video content, each sub-portion being encoded into one or more payload packets of a sequence of packets of the video data stream, the sequence of packets being divided into a sequence of access units, thus each access unit collecting payload packets pertaining to an individual picture of the video content, wherein at least some access units have ROI packets interleaved between payload packets pertaining to the picture of the individual access unit, wherein each ROI packet pertains to one or more subsequent payload packets following the individual ROI packet in the sequence of packets and identifies whether the sub-portion encoded into any of the one or more payload packets pertaining to the individual ROI packet covers a region of interest of the video content.

2. The method of claim 1, wherein, the sub-portions are segments and the video content is encoded into the video data stream using prediction and entropy encoding, wherein the prediction and / or the entropy encoding is limited to within tiles into which the pictures of the video content are divided, wherein each ROI packet additionally identifies a tile within which the sub-portion encoded into any of the one or more payload packets pertaining to the individual ROI packet covers the region of interest.

3. The method of claim 1 or 2, wherein, each ROI packet pertains to the immediately following payload packet only.

4. The method of claim 1 or 3, wherein, each ROI packet pertains to all payload packets following the individual ROI packet in the sequence of packets up to the earlier of the end of the access unit and a ROI packet following the individual ROI packet respectively, arranged within the access unit.

5. The method of claim 1, wherein, the video data stream is a Scalable Video Coding (SVC) data stream.

6. A network device positioned between an encoder and a decoder, configured to receive a video data stream transmitted according to the method of any of claims 1 to 5 and identify the ROI of the video content based on the ROI packets.

7. The network device of claim 6, wherein, the network device is configured to use the result of the identification to make decisions on transmission tasks pertaining to the video data stream.

8. The network device of claim 6 or 7, wherein, transmission tasks include retransmission requests with respect to defective packets.

9. The network device of claim 6 or 7, wherein, the network device is configured to increase the priority of processing the region of interest by assigning higher priority to ROI packets signaling coverage of the region of interest by the sub-portion encoded into any of the one or more payload packets pertaining to the individual ROI packet following the individual ROI packet and the one or more payload packets following the individual ROI packet pertaining to the individual ROI packet than to ROI packets signaling no coverage.

10. The network device of claim 9, wherein, The network device is configured to request retransmission of payload packets assigned with higher priority first before requesting any retransmission of payload packets assigned with lower priority.

11. The network device of claim 6, wherein, The video data stream is a Scalable Video Coding (SVC) data stream.

12. A decoder configured to: decoding the video content from a video data stream in units of sub-portions of images of the video content by decoding each sub-portion from one or more payload packets of a sequence of packets of the video data stream, the sequence of packets being divided into a sequence of access units, thus each access unit collecting payload packets related to an individual image of the video content, wherein, At least some of the access units have ROI packets interleaved between the payload packets relating to the pictures of the individual access units, wherein each ROI packet relates to one or more subsequent payload packets following the individual ROI packet in the packet sequence and identifies whether the subpart encoded into any of the one or more payload packets the individual ROI packet relates to covers a region of interest of the video content, and identify a ROI of the video content based on the ROI packets.

13. The decoder of claim 12, wherein, The video data stream is a Scalable Video Coding (SVC) data stream.

14. An encoder configured to: encoding the video content into a video data stream in units of sub-portions of images of the video content, wherein, encode each subpart into one or more payload packets of a packet sequence of the video data stream, the packet sequence being divided into a sequence of access units, thus each access unit collecting payload packets relating to an individual picture of the video content, interleave ROI packets between the payload packets relating to the pictures of the individual access units in at least some of the access units, wherein each ROI packet relates to one or more subsequent payload packets following the individual ROI packet in the packet sequence and identifies whether the subpart encoded into any of the one or more payload packets the individual ROI packet relates to covers a region of interest of the video content, thus a ROI of the video content being identifiable based on the ROI packets.

15. The encoder of claim 14, wherein, The video data stream is a Scalable Video Coding (SVC) data stream.

16. A method for decoding, comprising: receiving a video data stream having a video content encoded into it in units of subparts of pictures of the video content, each subpart being encoded into one or more payload packets of a packet sequence of the video data stream, the packet sequence being divided into a sequence of access units, thus each access unit collecting payload packets relating to an individual picture of the video content, wherein at least some of the access units have ROI packets interleaved between the payload packets relating to the pictures of the individual access units, wherein each ROI packet relates to one or more subsequent payload packets following the individual ROI packet in the packet sequence and identifies whether the subpart encoded into any of the one or more payload packets the individual ROI packet relates to covers a region of interest of the video content, and identify a ROI of the video content based on the ROI packets.

17. The method of claim 16, wherein, The video data stream is a Scalable Video Coding (SVC) data stream.

18. A digital storage medium having stored thereon a computer program having a program code for performing the method according to claim 1, 16 or 17, when the program code is executed on a computer.

Citation Information

Patent Citations

  • Video data stream, encoder, method for encoding video content, and decoder

    CN110536136A