Video data stream concept

The enhanced video coding standard addresses low-latency decoding challenges by allowing early transmission of non-VCL NAL units and optimizing sub-picture level signaling, ensuring timely and efficient decoding in video transmission.

JP2025188074APending Publication Date: 2025-12-25DOLBY VIDEO COMPRESSION LLC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2025148578
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2012-06-29
Filing Date
2025-09-08
Publication Date
2025-12-25

AI Technical Summary

Technical Problem

Existing video coding standards, such as HEVC, face challenges in low-latency scenarios where the timing of decoding units and the number of NAL units associated with sub-pictures must be known in advance, and require significant sub-picture level signaling, especially for applications like ROI signaling or tile dimension signaling.

Method used

The proposed solution involves enhancing the video coding standard to allow for the transmission of non-VCL NAL units before the encoder finishes encoding a picture, and reducing the need for filler data by using advanced signaling mechanisms that provide necessary information on the sub-picture level, such as ROI and tile dimensions, ensuring timely and efficient decoding.

Benefits of technology

This approach enables timely decoding in low-latency scenarios by providing necessary information ahead of encoding, reducing the need for filler data, and optimizing signaling for efficient video transmission.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025188074000001_ABST
    Figure 2025188074000001_ABST
Patent Text Reader

Abstract

To provide a video data stream encoding concept which enables low inter-terminal delay and is efficient to facilitate identifying a portion of a data stream in a region of interest or in a specific tile.SOLUTION: Decoder retrieval timing information, ROI information and tile identification information are conveyed within a video data stream at a level which allows easy access by network entities such as MANEs or decoders. To reach such a level, information of such types are conveyed within a video data stream by way of packets interspersed in packets of access units of a video data stream. In accordance with an embodiment, the interspersed packets are of a removable packet type, i.e., the removal of these interspersed packets maintains decoder's ability to completely recover video content 16 conveyed via the video data stream.SELECTED DRAWING: Figure 11
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a video data stream concept that is particularly advantageous in connection with low latency applications. [Background technology]

[0002] HEVC [2] allows different means of High Level Syntax signaling to the application layer. These means are NAL unit headers, Parameter Sets, and Supplemental Enhancement Information (SEI) Messages. The latter are not used in the decoding process. Other means of High Level Syntax signaling originate from respective transport protocol specifications, such as the MPEG2 Transport Protocol [3] or the Realtime Transport Protocol [4], and their payload-specific specifications, such as the proposal for H.264 / AVC [5], Scalable Video Coding (SVC) [6], or HEVC [7]. Such transport protocols can introduce High Level Signaling using structures and mechanisms similar to the High Level Signaling of the respective application layer codec specifications, such as HEVC [2]. One example of such signaling is the Payload Content Scalability Information (PACSI) NAL unit, as described in [6], which provides auxiliary information to the transport layer.

[0003] For parameter sets, HEVC includes a Video Parameter Set (VPS) that contains the most important stream information used by the application layer in one central location. In earlier approaches, this information had to be collected from multiple parameter sets and NAL unit headers.

[0004] Prior to this application, the current state of the standard regarding the Coded Picture Buffer (CPB) operation of the Hypothetical Reference Decoder (HRD) and all related syntax set in the Sequence Parameter Set (SPS) / Video Usability Information (VUI), Picture Timing SEI, Buffering Period SEI, and the definition of decoding units describing subpictures, and the syntax of slice headers and dependent slices present in the Picture Parameter Set (PPS) was as follows:

[0005] To enable low-latency CPB operations at the sub-picture level, sub-picture CPB operations were proposed and incorporated into the HEVC draft standard 7JCTVC-I1003 [2], where, in particular, the decoding unit was specified in section 3 of [2] as follows: Decoding device: An access unit or a subset of an access unit. If SubPicCpbFlag is equal to 0, the decoding unit is an access unit. Otherwise, the decoding unit consists of one or more VCL NAL units in the access unit and associated non-VCL NAL units. For the first VCL NAL unit in an access unit, there is an associated non-VCL NAL unit, and filler data NAL units, if any, immediately follow the first VCL NAL unit and follow all non-VCL NAL units in the access unit that precede the first VCL NAL unit. For a VCL NAL unit that is not the first VCL NAL unit in an access unit, the associated non-VCL NAL units are the filler data NAL units, if any, immediately follow the VCL NAL unit.

[0006] In the standard defined up to that time, "Removal timing of decoding units and decoding of decoding units" is described and added to Annex C "Hypothential reference decoder". To send sub-picture timing, the buffering period SEI and picture timing SEI messages, along with the HRD parameter in the VUI, are extended to support decoding units as sub-picture units.

[0007] The Buffering period SEI message syntax of [2] is shown in Figure 1.

[0008] When NalHrdBpPresentFlag or VclHrdBpPresentFlag is equal to 1, a buffering period SEI message can be associated with any access unit of the bitstream, and a buffering period SEI message is associated with each RAP access unit and each access unit associated with a recovery point SEI message.

[0009] In some applications, the frequent presence of a buffering period SEI message is desirable.

[0010] A buffering period is defined as the set of access units between two instances of the buffering period SEI message in decoding order.

[0011] The semantics were as follows:

[0012] seq_parameter_set_id specifies the sequence parameter set that contains the sequence HRD attribute. The value of seq_parameter_set_id is equal to the value of seq_parameter_set_id in the picture parameter set referenced by the primary coded picture associated with the buffering period SEI message. The value of seq_parameter_set_id is in the range of 0 to 31.

[0013] rap_cpb_params_present_flag equal to 1 specifies the presence of the initial_alt_cpb_removal_delay[SchedSelIdx] and initial_alt_cpb_removal_delay_offset[SchedSelIdx] syntax elements. If not present, the value of rap_cpb_params_present_flag is inferred to be equal to 0. If the associated image is neither a CRA nor a BLA image, the value of rap_cpb_params_present_flag is equal to 0.

[0014] initial_cpb_removal_delay[SchedSelIdx] and initial_alt_cpb_removal_delay[SchedSelIdx] specify the initial CPB removal delay for SchedSelIdx-th-CPB. The syntax element has a bit length given by initial_cpb_removal_delay_length_minus1+1 and is in units of 90 kHz clock. The value of the syntax element shall not be equal to 0 and shall not exceed 90000 * (CpbSize[SchedSelIdx] ÷ BitRate[SchedSelIdx]), which is the time equivalent of the CPB size in 90 kHz clock units.

[0015] initial_cpb_removal_delay_offset[SchedSelIdx] and initial_alt_cpb_removal_delay_offset[SchedSelIdx] are used by the SchedSelIdx-th CPB to specify the time of the initial delivery of an encoded data unit to the CPB. The syntax elements have a length in bits given by initial_cpb_removal_delay_length_minus1+1 and are in units of a 90 kHz clock. These syntax elements are not used by the decoder and are only needed for the delivery scheduler (HSS).

[0016] Across all coded video sequences, the sum of initial_cpb_removal_delay[SchedSelIdx] and initial_cpb_removal_delay_offset[SchedSelIdx] is constant for each value of SchedSelIdx, and the sum of initial_alt_cpb_removal_delay[SchedSelIdx] and initial_alt_cpb_removal_delay_offset[SchedSelIdx] is constant for each value of SchedSelIdx.

[0017] The picture timing SEI message syntax of [2] is shown in Figure 2.

[0018] The syntax of a picture timing SEI message depends on the contents of the sequence parameter set in effect for the coded picture associated with the picture timing SEI message. However, unless a buffering period SEI message exists within the access unit prior to the picture timing SEI message for an IDR or BLA access unit, activation of the associated sequence parameter set (and, for an IDR or BLA picture that is not the first picture in the bitstream, determination that the coded picture is an IDR or BLA picture) does not occur until decoding the first coded slice NAL unit of the coded picture. Because the coded slice NAL units of a coded picture follow the picture timing SEI message in NAL unit order, it may be necessary to store the RBSP containing the picture timing SEI message until the decoder determines the parameters of the sequence parameter set in effect for the coded picture, at which time it performs parsing of the picture timing SEI message.

[0019] The presence of a picture timing SEI message in the bitstream is specified as follows: - If CpbDpbDelaysPresentFlag is equal to 1, one picture timing SEI message is present in every access unit of a coded video sequence. - Otherwise (CpbDpbDelaysPresentFlag is equal to 0), no picture timing SEI message is present in any access unit of the coded video sequence.

[0020] The semantics are defined as follows:

[0021] cpb_removal_delay specifies how many clock ticks to wait after the removal from the CPB of the access unit associated with the most recent Buffering Period SEI message in the preceding access unit before removing the access unit data associated with a picture timing SEI message from the buffer. This value is also used to calculate the earliest possible time of appearance of access unit data in the CPB for HSS. The syntax element is a fixed length code whose length in bits is given by cpb_removal_delay_length_minus1+1. cpb_removal_delay is modulo 2. (cpb_removal_delay_length_minus1+1) This is the remainder of the counter.

[0022] The value of cpb_removal_delay_length_minus1, which determines the length (in bits) of the syntax element cpb_removal_delay, is the value of cpb_removal_delay_length_minus1 coded in the sequence parameter set that serves the primary coded picture associated with the picture timing SEI message, except that cpb_removal_delay specifies the number of clock ticks relative to the removal time of the preceding access unit containing the buffering period SEI message, which may be an access unit of a different coded video sequence.

[0023] dpb_output_delay is used to calculate the DPB output time of a picture. It specifies how many clock ticks to wait after removal of the last decoding unit of an access unit from the CPB before the decoded picture is output from the DPB.

[0024] The image is not removed from the DPB at output time if it is still marked as "Use for short term reference" or "Use for long term reference".

[0025] Only one dpb_output_delay is specified for the decoded image.

[0026] The length of the syntax element dpb_output_delay is given in bits by dpb_output_delay_length_minus1 + 1. When sps_max_dec_pic_buffering[max_temporal_layers_minus1] is equal to 0, dpb_output_delay shall be equal to 0.

[0027] The output time resulting from the dpb_output_delay of some pictures output from an output timing synchronous decoder precedes the output times resulting from the dpb_output_delay of all pictures in the next coded video sequence in decoding order.

[0028] The image output order specified by the value of this syntax element is the same order as specified by the value of PicOrderCntVal.

[0029] For pictures that are not output by the "bumping" process because they precede in decoding order an IDR or BLA picture with no_output_of_prior_pics_flag equal to 1 or inferred to be equal to 1, the output time obtained from dpb_output_delay increases with increasing values ​​of PicOrderCntVal associated with all pictures in the same coded video sequence.

[0030] num_decoding_units_minus1 plus 1 specifies the number of decoding units in the access unit to which the picture timing SEI message is associated. The value of num_decoding_units_minus1 is in the range from 0 to PicWidthInCtbs*PicHeightInCtbs-1, inclusive.

[0031] num_nalus_in_du_minus1[i]plus1 specifies the number of NAL units in the i-th decoding unit of the access unit to which the picture timing SEI message is associated. The value of num_nalus_in_du_minus1[i] is in the range from 0 to PicWidthInCtbs*PicHeightInCtbs-1.

[0032] The first decoding unit of an access unit consists of the first num_nalus_in_du_minus1[0] consecutive NAL units in decoding order in the access unit. The i-th (i greater than 0) decoding unit of an access unit consists of the num_nalus_in_du_minus1[i]+1 consecutive NAL units that immediately follow, in decoding order, the last NAL unit in the previous decoding unit of the access unit. Each decoding unit has at least one VCL NAL unit. All non-VCL NAL units associated with a VCL NAL unit are included in the same decoding unit.

[0033] du_cpb_removal_delay[i] specifies how many sub-picture clock ticks to wait after the removal from the CPB of the first decoding unit in the access unit associated with the most recent buffering period SEI message in the previous access unit before removing from the CPB the ith decoding unit in the access unit associated with the picture timing SEI message. This value is also used to calculate the earliest possible time of arrival of decoding unit data in the CPB for HSS. The syntax element is a fixed length code whose length in bits is given by cpb_removal_delay_length_minus1+1. du_cpb_removal_delay[i] is a modulo 2 (cpb_removal_delay_length_minus1+1) This is the remainder of the counter.

[0034] The value of cpb_removal_delay_length_minus1 that determines the length (in bits) of the syntax element du_cpb_removal_delay[i] is the value of cpb_removal_delay_length_minus1 that is coded in the sequence parameter set that serves the coded picture associated with the picture timing SEI message, while du_cpb_removal_delay[i] specifies the number of subpicture clock ticks relative to the removal time of the first decoding unit in the preceding access unit containing the buffering period SEI message, which may be an access unit of a different coded video sequence.

[0035] Some information was included in the VUI syntax of [2]. The VUI parameter syntax of [2] is shown in Figure 3. The HRD parameter syntax of [2] is shown in Figure 4. The semantics were defined as follows:

[0036] sub_pic_cpb_params_present_flag equal to 1 indicates that there is a subpicture-level CPB removal delay parameter and that CPB can operate at the access unit level or the subpicture level. sub_pic_cpb_params_present_flag equal to 0 indicates that there is no subpicture-level CPB removal delay parameter and that CPB operates at the access unit level. When sub_pic_cpb_params_present_flag is not present, its value is inferred to be equal to 0.

[0037] num_units_in_sub_tick is the number of time units of a clock running at frequency time_scale Hz that corresponds to one increment (called a subpicture clock tick) of the subpicture clock tick counter. num_units_in_sub_tick must be greater than 0. A subpicture clock tick is the smallest interval of time that can be represented in the coded data when sub_pic_cpb_params_present_flag is equal to 1.

[0038] A tiles_fixed_structure_flag equal to 1 indicates that, if present, each picture parameter set featured in the coded video sequence has the same values ​​for the syntax elements num_tile_columns_minus1, num_tile_rows_minus1, uniform_spacing_flag, column_width[i], row_height[i], and loop_filter_across_tiles_enabled_flag. A tiles_fixed_structure_flag equal to 0 indicates whether the tiles syntax elements of different picture parameter sets can have the same values. When the tiles_fixed_structure_flag syntax element is not present, it is inferred to be equal to 0.

[0039] Signaling tiles_fixed_structure_flag equal to 1 is a guarantee to the decoder that each picture in a coded video sequence has the same number of tiles distributed in an identical manner, which may be helpful for workload distribution in case of multi-threaded decoding.

[0040] The filler data in [2] was signaled using the filter data RBSP syntax shown in FIG.

[0041] The hypothetical reference decoder in [2] used to check bitstream and decoder conformance was defined as follows:

[0042] Two types of bitstreams are subject to HRD conformance checking with this Recommendation | International Standard. The first such type of bitstream, called Type I bitstream, is a NAL unit stream that contains only VCL NAL units and filler data NAL units for all access units of the bitstream. The second type of bitstream, called Type II bitstream, contains, in addition to VCL NAL units and filler data NAL units for all access units of the bitstream, at least one of the following: - further non-VCL NAL units other than filler data NAL units, - All leading_zero_8bits, zero_byte, start_code_prefix_one_3bytes and trailing_zero_8bits syntax elements that form a byte stream from a NAL unit stream.

[0043] Figure 6 shows the types of bitstream conformance points checked by the HRD in [2].

[0044] Two types of HRD parameter sets are used: NAL HRD parameters and VCL HRD parameters. HRD parameter sets are signaled by video availability information, which is part of the sequence parameter set syntax structure.

[0045] All sequence parameter sets and picture parameter sets referenced for a VCL NAL unit, and the corresponding buffering period and picture timing SEI messages, are conveyed to the HRD in a timely manner, in the bitstream or by other means.

[0046] The specification for "presence" of non-VCL NAL units is also met when those NAL units (or just some of them) are conveyed to the decoder (or to the HRD) by other means not specified by this Recommendation. For bit counting purposes, only the appropriate bits actually present in the bitstream are counted.

[0047] As an example, synchronization of a non-VCL NAL unit that is conveyed by means other than its presence in the bitstream with a NAL unit that is present in the bitstream can be achieved by indicating two points in the bitstream between which a non-VCL NAL unit is present in the bitstream and having the encoder decide to convey it in the bitstream.

[0048] When the content of non-VCL NAL units is conveyed to an application by some means other than their presence within the bitstream, the representation of the content of the non-VCL NAL units is not required to use the same syntax specified in this appendix.

[0049] It should be noted that when HRD information is included in a bitstream, conformance of the bitstream with this subordinate requirement can be verified solely based on the information contained in the bitstream. When HRD information is not present in the bitstream, conformance can only be verified when HRD data is provided by some other means not specified in this Recommendation|International Standard, as is the case for all "standalone" Type I bitstreams.

[0050] The HRD includes a coded picture buffer (CPB), an instantaneous decoding process, a decoded picture buffer (DPB) and output cropping, as shown in FIG.

[0051] The CPB size (number of bits) is CpbSize[SchedSelIdx]. The DPB size (number of picture store buffers) for temporal layer X is sps_max_dec_pic_buffering[X], for X in the range 0 to sps_max_temporal_layers_minus1 inclusive.

[0052] The variable SubPicCpbPreferredFlag is either specified by external means or set to 0 if not specified by external means.

[0053] The variable SubPicCpbFlag is derived as follows: SubPicCpbFlag=SubPicCpbPreferredFlag&&sub_pic_cpb_params_present_flag

[0054] If SubPicCpbFlag is equal to 0, the CPB operates at the access unit level and each decoding unit is an access unit. Otherwise, the CPB operates at the subpicture level and each decoding unit is a subset of an access unit.

[0055] The HRD operates as follows: Data associated with decoding units flowing into the CPB according to a specified arrival schedule is distributed by the HSS. Data associated with each decoding unit is removed and immediately decoded by an instantaneous decoding process at the CPB removal time. Each decoded image is placed in the DPB. Decoded images are removed from the DPB when they are no longer needed for later DPB output or inter-prediction reference.

[0056] The HRD is initialized as specified by the buffer period SEI. The removal timing of a decoding unit from the CPB and the output timing of the decoded image from the DPB are specified in the image timing SEI message. All timing information for a particular decoding unit arrives before the CPB removal time of the decoding unit.

[0057] The HRD is used to check the consistency of the bitstream and the decoder.

[0058] While agreement is guaranteed under the assumption that all frame rates and clocks used to generate the bitstream exactly match the values ​​signaled in the bitstream, in a real system each of these can vary from the signaled or specified value.

[0059] All calculations are done with real values, so that rounding errors cannot be propagated. For example, the number of bits in a CPB just before or after removal of a decoding unit is not necessarily an integer.

[0060] The variable tc is called a clock tick and is derived as follows: tc=num_units_in_tick time_scale The variable tc_sub is called the subpicture clock tick and is derived as follows: tc_sub=num_units_in_sub_tick time_scale

[0061] The following are provided to express constraints: Let access unit n be the nth access unit in decoding order, with the first access unit being access unit 0. Let image n be the coded or decoded image of access unit n. Let decoding unit m be the mth decoding unit in the decoding order with the first decoding unit being decoding unit 0.

[0062] In [2], the slice header syntax allows for so-called dependent slices.

[0063] Figure 8 shows the slice header syntax for [2].

[0064] The slice header semantics are defined as follows:

[0065] A dependent_slice_flag equal to 1 indicates that the value of each unindicated slice header syntax element is inferred to be equal to the value of the corresponding slice header syntax element in the previous slice containing the coding tree block whose coding tree block address is SliceCtbAddrRS-1. When not indicated, the value of dependent_slice_flag is inferred to be equal to 0. The value of dependent_slice_flag is equal to 0 when SliceCtbAddrRS is equal to 0.

[0066] slice_address specifies the address in slice granularity where the slice begins. The slice_address syntax element is (CEil(Log2(PicWidthInCtbs*PicHeightInCtbs))+SliceGranularity) bits in length.

[0067] The variable SliceCtbAddrRS specifies the coding tree block where the slice starts in coding tree block raster scan order and is derived as follows: SliceCtbAddrRS=(slice_address>>SliceGranularity)

[0068] The variable SliceCbAddrZS specifies the address of the first coding block of a slice of minimum coding block granularity in z-scan order and is derived as follows: SliceCbAddrZS=slice_address <<((log2_diff_max_min_coding_block_size=SliceGranularity)<<1)

[0069] Slice decoding begins at the maximum possible coding unit of the slice start coordinate.

[0070] first_slice_in_pic_flag indicates whether the slice is the first slice of the picture. If first_slice_in_pic_flag is equal to 1, the variables SliceCbAddrZS and SliceCtbAddrRS are both set to 0, and decoding starts from the first coding tree block of the picture.

[0071] pic_parameter_set_id specifies the picture parameter that was introduced. The value of pic_parameter_set_id ranges from 0 to 255 inclusive.

[0072] num_entry_point_offsets specifies the number of entry_point_offset[i] syntax elements in the slice header. When tiles_or_entropy_coding_sync_idc is equal to 1, the value of num_entry_point_offsets ranges from 0 to (num_tile_columns_minus1+1)*(num_tile_rows_minus1+1)-1, inclusive. When tiles_or_entropy_coding_sync_idc is equal to 2, the value of num_entry_point_offsets ranges from 0 to PicHeightInCtbs-1, inclusive. When not specified, the value of num_entry_point_offsets is inferred to be equal to 0.

[0073] offset_len_minus1 plus1 specifies the length, in bits, of the entry_point_offset[i] syntax element.

[0074] entry_point_offset[i] specifies the offset of the ith entry point in bytes, represented by offset_len_minus1 plus 1 bit. The coded slice data after the slice header consists of num_entry_point_offsets+1 subsets, with subset index values ​​ranging from 0 to num_entry_point_offsets inclusive. Subset 0 consists of bytes 0 to entry_point_offset[0]-1 inclusive of the coded slice data, subset k consists of bytes entry_point_offset[k-1] to entry_point_offset[k]+entry_point_offset[k-1]-1 inclusive of the coded slice data, where k ranges from 1 to num_entry_point_offset-1 inclusive, and the last subset (with subset index equal to num_entry_point_offsets) consists of the remaining bytes of the coded slice data.

[0075] When tiles_or_entropy_coding_sync_idc is equal to 1 and num_entry_point_offsets is greater than 0, each subset contains all the coded bits of exactly one tile, and the number of subsets (i.e., the value of num_entry_point_offsets+1) is less than or equal to the number of tiles in the slice. When tiles_or_entropy_coding_sync_idc is equal to 1, each slice must contain an integer number of subsets of one tile or a complete tile (entry point case signaling is unnecessary).

[0076] When tiles_or_entropy_coding_sync_idc is equal to 2 and num_entry_point_offsets is greater than 0, each subset k, with k in the range 0 to num_entry_point_offsets-1 inclusive, contains exactly all coded bits of one row of the coding tree block, the last subset (with subset index equal to num_entry_point_offsets) contains all coded bits of the remaining coding blocks in the slice, which consist of either exactly one row of the coding tree block or a subset of one row of the coding tree block, and the number of subsets (i.e., the value of num_entry_point_offsets+1) is equal to the number of rows of the coding tree block in the slice, and the subsets of one row of the coding tree block in the slice are also counted.

[0077] When tiles_or_entropy_coding_sync_idc is equal to 2, a slice contains many rows of coding tree blocks and subsets of rows of coding tree blocks. For example, if a slice contains two and a half rows of coding tree blocks, the number of subsets (i.e., the value of num_entry_point_offsets+1) will be equal to 3.

[0078] Figure 9 shows the image parameter set RBSP syntax of [2] and the image parameter set RBSP semantics of [2], which are defined as follows:

[0079] dependent_slice_enabled_flag equal to 1 specifies the presence of the syntax element dependent_slice_flag in the slice header for the picture coded with reference to a picture parameter set. dependent_slice_enabled_flag equal to 0 specifies the absence of the syntax element dependent_slice_flag in the slice header for the picture coded with reference to a picture parameter set. When tiles_or_entropy_coding_sync_idc is equal to 3, the value of dependent_slice_enabled_flag shall be equal to 1.

[0080] tiles_or_entropy_coding_sync_idc equal to 0 indicates that there is only one tile in each picture referenced by the picture parameter set, and there is no specific synchronization process for the context change that is caused before decoding the first coding tree block in a row of coding tree blocks in each picture referenced by the picture parameter set, and the values ​​of cabac_independent_flag and dependent_slice_flag for coded pictures referenced by the picture parameter set shall not both be equal to 1.

[0081] When cabac_independent_flag and depedent_slice_flag are both equal to 1 for a slice, the slice is an entropy slice.

[0082] tiles_or_entropy_coding_sync_idc equal to 1 indicates that there are one or more tiles in each picture with reference to the picture parameter set, there will be no specific synchronization process for context changes drawn before decoding the first coding tree block in a row of coding tree blocks in each picture with reference to the picture parameter set, and the values ​​of cabac_independent_flag and dependent_slice_flag for the coded picture with reference to the picture parameter set will not both be equal to 1.

[0083] tiles_or_entropy_coding_sync_idc equal to 2 indicates that there is only one tile in each image with reference to the picture parameter set, a specific synchronization process for context changes is derived before decoding the first coding tree block in a row of coding tree blocks in each image with reference to the picture parameter set, a specific storage process for context changes is derived after decoding the second coding tree block in a row of coding tree blocks in each image with reference to the picture parameter set, and the values ​​of cabac_independent_flag and dependent_slice_flag for the coded image with reference to the picture parameter set shall not both be equal to 1.

[0084] tiles_or_entropy_coding_sync_idc equal to 3 clarifies that there is only one tile in each picture with reference to the picture parameter set, there will be no specific synchronization process for context changes drawn before decoding the first coding tree block in a row of coding tree blocks in each picture with reference to the picture parameter set, and the values ​​of cabac_independent_flag and dependent_slice_flag for the picture coded with reference to the picture parameter set will both be equal to 1.

[0085] When dependent_slice_enabled_flag is equal to 0, tiles_or_entropy_coding_sync_idc shall not be equal to 3.

[0086] It is a bitstream conformance requirement that the value of tiles_or_entropy_coding_sync_idc be common to all picture parameter sets that start in a coded video sequence.

[0087] For each slice referenced in the picture parameter set, if tiles_or_entropy_coding_sync_idc is equal to 2 and the first coding block of the slice is not the first coding tree block in a coding tree block row, then the last coding block of the slice belongs to the same coding tree block row as the first coding block of the slice.

[0088] num_tile_columns_minus1plus1 specifies the number of columns of tiles the image is divided into.

[0089] num_tile_rows_minus1plus1 specifies the number of rows of tiles the image is divided into. When num_tile_columns_minus1 is equal to 0, num_tile_rows_minus1 shall not be equal to 0.

[0090] uniform_spacing_flag equal to 1 specifies that column boundaries, and similarly row boundaries, are uniformly distributed across the image. uniform_spacing_flag equal to 0 specifies that column boundaries, and similarly row boundaries, are not uniformly distributed across the image, but are explicitly signaled using the syntax elements column_width[i] and row_height[i].

[0091] column_width[i] specifies the column width of the ith tile in units of coding tree blocks.

[0092] row_height[i] specifies the row height of the ith tile in coding tree blocks.

[0093] The vector colWidth[i] specifies the column width of the ith tile in units of CTB with column i ranging from 0 to num_tile_columns_minus1.

[0094] The vector CtbAddrRStoTS[ctbAddrRS] specifies the correspondence from a CTB address in raster scan order to a CTB address in tile scan order, with index ctbAddrRS ranging from 0 to (picHeightInCtbs*picWidthInCtbs)-1. The vector CtbAddrTStoRS[ctbAddrTS] specifies the correspondence from a CTB address in tile scan order to a CTB address in raster scan order, with the index ctbAddrTS ranging from 0 to (picHeightInCtbs*picWidthInCtbs)-1.

[0095] The vector TileId[ctbAddrTS] specifies the correspondence from CTB address to tile id in tile scan order, with ctbAddrTS ranging from 0 to (picHeightInCtbs*picWidthInCtbs)-1.

[0096] The values ​​of colWidth, CtbAddrRStoTS, CtbAddrTStoRS, and TileId are derived by calling the CTB raster and tile scanning interaction process, as specified in subclause 6.5.1, with PicHeightInCtbs and PicWidthInCtbs as input, and the output is assigned to colWidth, CtbAddrRStoTS, and TileId.

[0097] The value of ColumnWidthInLumaSamples[i], which defines the width of the i-th column of tiles in terms of luma samples, is set equal to colWidth[i] << Log2CtbSize.

[0098] For x ranging from 0 to picWidthInMinCbs - 1 and y ranging from 0 to picHeightInMinCbs - 1, the array MinCbAddrZS[x][y], which defines the mapping from the location (x, y) in terms of minimum Cbs in z-scan order to the minimum CB address, is derived by calling the z-scanning order array initialization process as specified in subclause 6.5.2 with Log2MinCbSize, Log2CtbSize, PicHeightInCtbs, PicWidthInCtbs, and the vector CtbAddrRStoTS, such that the input and output are assigned to MinCbAddrZS.

[0099] A loop_filter_across_tiles_enabled_flag equal to 1 indicates that the in-loop filtering operation is performed across tile boundaries. A loop_filter_across_tiles_enabled_flag equal to 0 indicates that the in-loop filtering operation is not performed across tile boundaries. The in-loop filtering operation includes deblocking filter, sample adaptive offset, and adaptive loop filter operations. When not present, the value of the loop_filter_across_tiles_enabled_flag is assumed to be equal to 1.

[0100] cabac_independent_flag equal to 1 indicates that the CABAC decoding of the coded blocks of the slice is independent of any aspect of previously decoded slices. cabac_independent_flag equal to 0 indicates that the CABAC decoding of the coded blocks of the slice is dependent on aspects of previously decoded slices. When not present, the value of cabac_independent_flag is inferred to be equal to 0.

[0101] The derivation process for the availability of coding blocks with minimal coding block addresses was explained as follows: The inputs to this process are: -z-scan-order minimum coded block address minCbAddrZS -z-scan-order current minimum coded block address currMinCBAddrZS

[0102] The output of this process is the availability of the coding block with the smallest coding block address cbAddrZS in the z-scan order cbAvailable.

[0103] Note that the meaning of validity is determined when this process is invoked.

[0104] Note that any coding block, regardless of its size, is associated with a minimum coding block address, which is the address of the coding block with the minimum coding block size in z-scan order.

[0105] - cbAvailable is set to FALSE if one or more of the following conditions are true: -minCbAddrZS is less than 0. -minCbAddrZS is greater than currMinCBAddrZS. - The coding block with the minimum coding block address minCbAddrZS belongs to a different slice than the coding block with the current minimum coding block address currMinCBAddrZS, and the dependent_slice_flag of the slice containing the coding block with the current minimum coding block address currMinCBAddrZS is equal to 0. The coding block with the minimum coding block address minCbAddrZS is contained in a different tile than the coding block with the current minimum coding block address currMinCBAddrZS. Otherwise, cbAvailable is set to TRUE.

[0106] The CABAC parsing process for the slice data in [2] was as follows: This process is invoked when parsing a syntax element with descriptor ae(v).

[0107] The inputs to this process are the value of a syntax element and a request for the value of a previously parsed syntax element.

[0108] The output of this process is the value of the syntax element.

[0109] The CABAC parser initialization process is called when parsing slice data for a slice.

[0110] The minimum coding block address, ctbMinCbAddrT, of the coding tree block containing spatial neighbor block T (Figure 10a) is derived using the location (x0, y0) of the top-left luma sample of the current coding tree block as follows: x=x0+2< <Log2CtbSize-1 y=y0-1 ctbMinCbAddrT=MinCbAddrZS[x>>Log2MinCbSize][y>>Log2MinCbSize]

[0111] The variable availableFlagT is obtained by calling the coding block validity derivation process with ctbMinCbAddrT as input.

[0112] When starting to parse the coding tree, the following steps are applied: 1. The arithmetic decoding engine is initialized as follows: - If CtbAddrRS is equal to slice_address, dependent_slice_flag is equal to 1, and entropy_coding_reset_flag is equal to 0, then the following applies: - The CABAC parsing process synchronization process is called with TableStateIdxDS and TableMPSValDS as input. -The decoding process for binary decisions before termination is called, followed by an initialization process for arithmetic decoding. - Otherwise, if tiles_or_entropy_coding_sync_idc is equal to 2 and CtbAddrRS%PicWidthInCtbs is equal to 0, the following applies: When -availableFlagT is equal to 1, the CABAC parser process synchronization process is invoked with TableStateIdxWPP and TableMPSValWPP as inputs. - The decoding process for binary decisions before termination is called, followed by an initialization process for the arithmetic decoding engine.

[0113] 2. When cabac_independent_flag is equal to 0 and dependent_slice_flag is equal to 1, or tiles_or_entropy_coding_sync_idc is equal to 2, the storage process is applied as follows: When tiles_or_entropy_coding_sync_idc is equal to 2 and CtbAddrRS%PicWidthInCtbs is equal to 2, the storage process of the CABAC parsing process is called with TableStateIdxWPP and TableMPSValWPP as outputs. When cabac_independent_flag is equal to 0, dependent_slice_flag is equal to 1, and end_of_slice_flag is equal to 1, the storage process of the CABAC parsing process is called with TableStateIdxDS and TableMPSValDS as outputs.

[0114] Parsing of a syntax element proceeds as follows:

[0115] For each requested value of the syntax element, a binarization is obtained.

[0116] The sequence of binarized and parsed bins for a syntax element determines the decoding process flow.

[0117] For each bin of the binarization of the syntax element indexed by the variable binIdx, a context index ctxIdx is obtained.

[0118] For each ctxIdx, an arithmetic decoding process is launched.

[0119] The resulting sequence of parsed bins (b0...bbinIdx) is compared with the set of bin strings given by the binarization process after decoding each bin. When a sequence matches a bin string in the given set, the corresponding value is assigned to the syntax element.

[0120] If a syntax element value request is processed for the syntax element pcm-flag and the decoded value of pcm_flag is equal to 1, the decoding engine initializes any pcm_alignment_zero_bit after decoding num_subsequent_pcm and all pcm_sample_luma and pcm_sample_chroma data.

[0121] In the design framework described so far, the following issues have arisen:

[0122] The timing of the decoding units must be known before encoding and transmitting data in low-latency scenarios, where NAL units are already sent by the encoder while the encoder is still encoding part of the picture, i.e., other sub-picture decoding units. This is because the NAL unit order in an access unit only allows SEI messages to precede VCL (video coding NAL units) in the access unit, and in such low-latency scenarios, non-VCL NAL units must already be communicated, i.e., transmitted, when the encoder starts encoding the decoding units. Figure 10b illustrates the structure of an access unit described in [2]. [2] has not yet identified the end of a sequence or stream, and therefore their presence in the access unit was provisional.

[0123] Also, the number of NAL units associated with a subpicture must be known in advance in low-latency scenarios, because the picture timing SEI message contains and must follow this information and must be sent before the encoder begins encoding the actual picture. Application designers are reluctant to insert filler data NAL units, because potentially no filler data matching the NAL unit number exists, but they are sent for each decoding unit in the picture timing SEI, and they need a means to transmit this information on the subpicture level, which holds for the subpicture timing currently determined in the presence of an access unit by the parameters given in the timing SEI message.

[0124] Furthermore, a further drawback of the draft specification [2] is the large amount of sub-picture level signaling required for specific applications, such as ROI signaling or tile dimension signaling. [Prior art documents] [Non-patent literature]

[0125] [Non-Patent Document 1] Thomas Wiegand, Gary J. Sullivan, Gisle Bjontegaard, Ajay Luthra, "Overview of the H.264 / AVC Video Coding Standard", IEEE Trans. Circuits Syst. Video Technol., vol. 13, N7, July 2003. [Non-patent document 2] JCT-VC, "High-Efficiency Video Coding (HEVC) text specification Working Draft 7", JCTVC-I1003, May 2012. [Non-patent document 3] ISO / IEC 13818-1: MPEG-2 Systems specification. [Non-patent document 4] IETF RFC 3550 - Real-time Transport Protocol. [Non-Patent Document 5] Y.-K. Wang et al., "RTP Payload Format for H.264 Video", IETF RFC6184, http: / / tools.ietf.org / html / [Non-patent document 6] S. Wenger et al., "RTP Payload Format for Scalable Video Coding", IETF RFC6190, http: / / tools.ietf.org / html / rfc6190 [Non-Patent Document 7] T. Schierl et al., "RTP Payload Format for High Efficiency Video Coding", IETF internet draft, http: / / datatracker.ietf.org / doc / draft-schierl-payload-rtp-h265 / Summary of the Invention [Problem to be solved by the invention]

[0126] The challenges outlined above are not unique to the HEVC standard. Rather, they arise in connection with other video codecs as well. More generally, FIG. 11 illustrates a video transmission scenario in which a pair of encoders 10 and decoders 12 are connected via a network 14 to transmit video 16 from the encoder 10 to the decoder 12 with a short end-to-end delay. The challenges already outlined above are as follows: The encoder 10 encodes a sequence of frames 18 of video 16 according to a specific decoding order that essentially, but not necessarily, follows a representation order 20 of the frames 18 and, within each frame 18, propagates through the frame domain of the frames 18 in some particular manner, such as raster scanning with or without a tile-sectioning method for the frames 18. The decoding order controls the availability of information for the coding techniques employed by the encoder 10, such as prediction and / or entropy coding, i.e., the availability of information about spatially and / or temporally neighboring portions of the video 16 that serves as a basis for prediction or context selection. Even if encoder 10 may be able to use parallel processing to encode frames 18 of video 16, encoder 10 inevitably needs some time to encode a particular frame 18, such as the current frame. Figure 11, for example, shows a moment when encoder 10 has already finished encoding portion 18a of current frame 18, while another portion 18b of current frame 18 has not yet been encoded. Because encoder 10 has not yet encoded portion 18b, encoder 10 cannot predict how the available bitrate for encoding current frame 18 should be spatially distributed throughout current frame 18 to achieve an optimum, for example, in terms of rate / distortion perception.Therefore, the encoder 10 simply has two choices: either the encoder 10 guesses a near-optimal distribution of the available bit rate for the current frame 18 over slices into which the current frame 18 is previously spatially subdivided and therefore acknowledges that the guess is wrong, or the encoder 10 finishes encoding the current frame 18 before sending packets containing the slices from the encoder 10 to the decoder 12. In any case, to be able to utilize any transmission of slice packets of the currently coded frame 18 before the end of its coding, the network 14 should signal the bitrate associated with each such slice packet in the form of a coded picture buffer search time. However, as noted above, although the encoder 10 can vary the bitrate distributed over a frame 18 by individually defining decoder buffer search times for subpicture regions according to current versions of HEVC, the encoder 10 must transmit or send such information to the decoder 12 over the network 14 at the beginning of each access unit that collects all data related to the current frame 18, thereby forcing the encoder 10 to choose between the two aforementioned options: one that provides low latency but with a worse rate / distortion, and the other that provides optimal rate / distortion but with increased end-to-end latency.

[0127] Thus, until now, no video codec has been available that achieves such low latency by allowing an encoder to begin transmitting packets for a portion 18a of a current frame prior to encoding the remaining portion 18b of the current frame, and by allowing a decoder to take advantage of this intermediate transmission of packets for the spare portion 18a over the network 16 according to the decoding buffer search timing conveyed in the video data stream sent from the encoder 12 to the decoder 14. Illustrative applications that utilize this type of low latency include industrial applications, such as workpiece or manufacturing monitoring for automation or inspection purposes. Until now, no satisfactory solution has been available for informing the decoder of the association of packets to tiles that are of interest to a region of interest (ROI) of the current frame, such that an intermediate network entity within the network 16 is enabled to gather this type of information from the data stream without having to examine deeply inside the packets, i.e., slice syntax.

[0128] It is therefore an object of the present invention to provide a video data stream coding concept that is more efficient in allowing for lower end-to-end delay and / or makes it easier to identify portions of the data stream in regions of interest or in specific tiles. [Means for solving the problem]

[0129] This object is achieved by the subject matter of the independent claims.

[0130] One idea on which this application is based is that decoder search timing information, ROI information, and tile identification information must be conveyed within the video data stream at a level that allows easy access by network entities such as MANEs or decoders, and to reach this level, this type of information must be conveyed within the video data stream via packets that are dispersed among packets of access units of the video data stream. According to an embodiment, the dispersed packets are among the removable packet types, i.e., removal of these dispersed packets maintains the decoder's ability to fully recover the video content carried via the video data stream.

[0131] According to an aspect of the present application, achieving low end-to-end delay is made more effective by using distributed packets to convey information on the decoder buffer lookup time for the decoding unit formed by the payload packets following each timing control packet of the video data stream within the current access unit. This measurement enables the encoder to determine the decoder buffer lookup time on the fly while encoding the current frame, and to continuously determine the bitrate to be spent on the portion of the current frame that has already been encoded and transmitted in the payload packets, or, on the other hand, to adapt the distribution of the remaining bitrate available for the current frame to the remaining portion of the current frame that has been preceded by the transmitted timing control packets and thus not yet encoded. This measurement allows the available bitrate to be utilized efficiently, and delay is nevertheless kept shorter because the encoder does not have to wait for the current frame to completely finish encoding.

[0132] According to a further aspect of the present application, packets distributed among the payload packets of an access unit are utilized to convey information about the region of interest, thereby allowing easy access of this information by network entities, as described above, without the need to inspect intermediate payload packets. Furthermore, the encoder does not need to determine the subdivision of the current frame into subportions and respective payload packets in advance, but is still able to determine packets belonging to the ROI on the fly while encoding the current frame. Furthermore, according to an embodiment in which the distributed packets are portable packet types, the ROI information can be ignored by recipients of the video data stream that are not interested in or cannot process the ROI information.

[0133] A similar idea is utilized in this application according to another aspect in which distributed packets convey information about which tile a particular packet within an access unit belongs to.

[0134] Advantageous embodiments of the invention are the subject matter of the dependent claims. The preferred embodiment of the present application is described in more detail with reference to the following figures, in which Figures 1 to 10b show the current state of HEVC. [Brief explanation of the drawings]

[0135] [Figure 1] Figure 1 shows the buffer period SEI message syntax. [Figure 2] Figure 2 shows the image timing SEI message syntax. [Figure 3a] Figure 3a shows the VUI parameter syntax. [Figure 3b] Figure 3b shows the VUI parameter syntax. [Figure 4] Figure 4 shows the HRD parameter syntax. [Figure 5] Figure 5 shows the filler data RBSP syntax. [Figure 6]FIG. 6 shows the structure of a byte stream and a NAL unit stream for HRD consistency checking. [Figure 7] Figure 7 shows the HRD buffer model. [Figure 8] Figure 8 shows the slice header syntax. [Figure 9] Figure 9 shows the image parameter set RBSP syntax. [Figure 10a] FIG. 10a shows a diagrammatic view of a spatially adjacent coding tree block T that may be used to trigger a coding tree block-based guidance process related to the current coding tree block. [Figure 10b] Figure 10b shows the definition of the structure of an access unit. [Figure 11] Figure 11 shows a diagram of a pair of encoders and decoders connected via a network to illustrate the challenges involved in video data stream transmission. [Figure 12] FIG. 12 shows a schematic block diagram of an encoder according to an embodiment using timing control packets. [Figure 13] FIG. 13 shows a flow diagram illustrating an operation mode of the encoder of FIG. 12 according to an embodiment. [Figure 14] FIG. 14 shows a block diagram of an embodiment of a decoder to explain its functioning in relation to the video data stream generated by the encoder according to FIG. [Figure 15] FIG. 15 shows a schematic block diagram illustrating an encoder, network entities, and a video data stream according to a further embodiment using ROI packets. [Figure 16] FIG. 16 shows a schematic block diagram illustrating an encoder, network entities, and a video data stream according to a further embodiment using tile identification packets. [Figure 17] 17 shows the structure of an access unit according to an embodiment, where the dotted lines reflect the case of an optional slice prefix NAL unit. [Figure 18] FIG. 18 illustrates the use of tiles for area of ​​interest signaling. [Figure 19] Figure 19 shows the first simple syntax / version 1. [Figure 20] Figure 20 shows the extended syntax / version 2, which includes tile_id signaling, decoding unit start identifier, slice prefix ID and slice header data, which differ from the SEI message concept. [Figure 21] Figure 21 shows NAL unit type codes and NAL unit type classes. [Figure 22] Figure 22 shows a possible syntax for a slice header, where certain syntax elements present in the slice header according to the current version are moved to a lower hierarchical syntax element called slice_header_data(). [Figure 23a] Figure 23a shows a table in which all syntax elements apart from the slice header are signaled by the syntax element slice header data. [Figure 23b] Figure 23b shows a table in which all syntax elements apart from the slice header are signaled by the syntax element slice header data. [Figure 23c] Figure 23c shows a table in which all syntax elements apart from the slice header are signaled by the syntax element slice header data. [Figure 24] Figure 24 shows the supplemental enhancement information message syntax. [Figure 25a] Figure 25a shows the SEI payload syntax adapted to introduce the new slice or sub-picture SEI message type. [Figure 25b] Figure 25b shows the SEI payload syntax adapted to introduce the new slice or sub-picture SEI message type. [Figure 26] Figure 26 shows an example for a subpicture buffer SEI message. [Figure 27]Figure 27 shows an example for a sub-picture timing SEI message. [Figure 28] Figure 28 shows what the sub-picture slice information SEI message looks like. [Figure 29] Figure 29 shows an example for a sub-picture tile information SEI message. [Figure 30] Figure 30 shows an example of the syntax for the Subpicture Tile Range Information SEI message. [Figure 31] FIG. 31 shows a first variation of an example syntax for a region of interest SEI message where each ROI is signaled in an individual SEI message. [Figure 32] FIG. 32 shows a second variation of an example syntax for a region of interest SEI message in which all ROIs are signaled in a single SEI message. [Figure 33] FIG. 33 shows a possible syntax for a timing control packet according to a further embodiment. [Figure 34] FIG. 34 illustrates a possible syntax for a Tile Identification Packet according to an embodiment. [Figure 35] FIG. 35 shows possible subdivisions of an image according to different subdivision settings according to an embodiment. [Figure 36] FIG. 36 shows possible subdivisions of an image according to different subdivision settings according to an embodiment. [Figure 37] FIG. 37 shows possible subdivisions of an image according to different subdivision settings according to an embodiment. [Figure 38] FIG. 38 shows possible subdivisions of an image according to different subdivision settings according to an embodiment. [Figure 39] FIG. 39 shows an example of a portion from a video data stream according to an embodiment using timing control packets interspersed between payload packets of an access unit. DETAILED DESCRIPTION OF THE INVENTION

[0136] 12, an encoder 10 according to an embodiment of the present application and its mode of operation is described. The encoder 10 is configured to encode video content 16 into a video data stream 22. The encoder is configured to do this in units of sub-portions of frames / images 18 of the video content 16, which may be, for example, slices 24 into which the images 18 are divided, or some other spatial segment, such as, for example, tiles 26 or WPP sub-streams 28, all of which are illustrated in FIG. 12 merely for illustrative purposes, rather than implying that the encoder 10 needs to be capable of supporting, for example, tile or WPP parallel processing, or that the sub-portions need to be slices.

[0137] In encoding the video content 16 in units of subportions 24, the encoder 10 may follow a decoding order—or encoding order—defined within the subportions 24, where the images 18 are distributed according to a raster scan order, e.g., not necessarily corresponding to the representation order 20 defined among the images 18, and where the images 18 are interleaved within each image 18 block, with the subportions 24 representing the sequential running of such blocks along the decoding order. In particular, the encoder 10 may be configured to follow this decoding order to determine the availability of spatially and / or temporally neighboring portions of the current portion to be coded, e.g., to use characteristics indicative of such neighboring portions in predictive and / or entropy coding, to determine the prediction and / or entropy context: only previously visited portions of the video are available for coding / decoding. Otherwise, the characteristics just mentioned may be set to default values, or some other alternative may be used.

[0138] On the other hand, encoder 10 need not encode sub-portions 24 consecutively in decoding order. Rather, encoder 10 may employ parallel processing to speed up the encoding process or perform more complex encoding in real time. Similarly, encoder 10 may or may not be configured to transmit or send data encoding sub-portions in decoding order. For example, encoder 10 may output / send encoded data in some other order, such as according to the order in which the encoding of sub-portions is completed by encoder 10 due to parallel processing, which may deviate from the just-mentioned decoding order.

[0139] To make the encoded version of the sub-portions 24 suitable for transmission over a network, the encoder 10 encodes each sub-portion 24 into one or more payload packets of a sequence of packets of the video stream 22. If the sub-portions 24 are slices, the encoder 10 may be configured to, for example, replace each slice data, e.g., each encoded data, with one payload packet, such as a NAL unit. This packetization may help make the video data stream 22 suitable for transmission over a network. Thus, a packet may represent the smallest unit from which the video data stream 22 occurs, i.e., that may be individually sent by the encoder 10 for transmission over a network to a recipient.

[0140] Additionally, there may be other types of packets such as payload packets and timing control packets interspersed among them, as well as other packets described below, such as fill data packets, picture or sequence parameter set packets, for conveying infrequently changing syntax elements or EOF (end of file) or AUE (end of access unit) packets.

[0141] The encoder performs the encoding into payload packets such that the sequence of packets is divided into a sequence of access units 30, each access unit collecting payload packets 32 related to one image 18 of the video content 16. That is, the sequence of packets 34 forming the video data stream 22 is subdivided into non-overlapping portions called access units 30, each associated with a respective one of the images 18. The sequence of access units 30 follows the decoding order of the images 18 to which the access units 30 are associated. FIG. 12, for example, shows that an access unit 30 located in the center of the illustrated portion of the data stream 22 includes one payload packet 32 ​​for each sub-portion 24 into which the image 18 is subdivided. That is, each payload packet 32 ​​carries a corresponding sub-portion 24. The encoder 10 distributes the payload packets 32 into a sequence 34 of timing control packets 36, whereby the timing control packets subdivide the access units 30 into decoder units 38, such that at least some access units 30 are subdivided into two or more decoder units 38, as shown in the center portion of FIG. 12, with each timing control packet signaling a decoder buffer lookup time for the decoder unit 38 whose payload packet 32 ​​follows the respective timing control packet in the sequence 34 of packets. In other words, the encoder 10 prefixes each subsequence of the sequence of payload packets 32 within one access unit 30 with a respective timing control packet 36, which signals a decoder buffer lookup time for each subsequence of payload packets that are preceded by a respective timing control packet 36 and form a decoder buffer lookup time for the decoder unit 38. FIG. 12 illustrates, for example, the case where every second packet 32 ​​represents the first payload packet of the decoder unit 38 of the access unit 30.As shown in FIG. 12, the amount of data or bit rate spent per decoding unit 38 varies, and the decoder buffer lookup time can be correlated to this bit rate variation among decoding units 38, in that the decoder buffer lookup time of a decoding unit 38 can follow the decoder buffer lookup time signaled by the timing control packet 36 of the immediately preceding decoding unit 38 and a time interval corresponding to the bit rate spent by this immediately preceding decoding unit 58.

[0142] That is, the encoder 10 may operate as shown in Figure 13. In particular, as described above, the encoder 10 may encode a current sub-portion 24 of a current image 18 in step 40. As already mentioned, the encoder 10 may cycle sequentially through the sub-portions 24 in the above-described decoding order, as indicated by arrow 42, or the encoder 10 may use some parallel processing, such as WPP and / or tile processing, to simultaneously encode several "current sub-portions" 24 in parallel. Regardless of the use of parallel processing, the encoder 10 forms a decoding unit from one or several sub-portions just encoded in step 40 and subsequent step 44, and the encoder 10 sets and transmits a decoder buffer search time to this decoding unit, which in turn signals the just-set decoding buffer search time for the decoding unit in a time control packet. For example, the encoder 10 can determine the decoder buffer search time in step 44 based on the bit rate consumed in encoding the subportions encoded in the payload packets that form the current decoding unit, including all further intermediate packets, if any, within this decoding unit, i.e., "prefix packets."

[0143] Then, in step 46, encoder 10 can adapt the available bitrate based on the bitrate consumed by the decoding unit just transmitted in step 44. For example, if the image content in the decoding unit just transmitted in step 44 is highly complex in terms of compression ratio, encoder 10 can reduce the available bitrate for the next decoding unit to comply with some externally set target bitrate determined, for example, based on the current bandwidth conditions encountered in connection with the network transmitting video data stream 22. Steps 40-46 are then repeated. According to this scaling, image 18 is encoded and transmitted, i.e., in units of decoding units, each preceded by a corresponding timing control packet.

[0144] In other words, while encoding a current image 18 of the video content 16, the encoder 10 encodes 40 a current sub-portion 24 of the current image 18 into a current payload packet 32 ​​of a current decoding unit 38, sets a decoder buffer search time signaled by a current timing control packet (36) in the data stream at a first instant, transmits 44 the current decoding unit 38 prefixed with the current timing control packet 36 in the data stream, and encodes 44 a further sub-portion 24 of the current image 18 at a second time instant—second time visit step 40—later than the first time instant—first time visit step 44, by returning to step 40 from step 46.

[0145] The encoder 10 can reduce end-to-end delay, since it can transmit a decoding unit prior to the encoding of the rest of the current image to which this decoding unit belongs. On the other hand, the encoder 10 does not have to waste available bitrate, since it can react to the specific nature of the content of the current image and its spatial distribution of complexity.

[0146] Meanwhile, intermediate network entities critical to the transmission of the video data stream from the encoder to the decoder can use timing control packets 36 to ensure that some decoders receiving the video data stream 22 receive the decoded units in time to benefit from the encoding and transmission of the decoded units by the encoder 10. See, for example, FIG. 14, which illustrates a decoder for decoding the video data stream 22. The decoder 12 receives the video data stream 22 in a coded picture buffer CPB 48 via the network through which the encoder 10 transmitted the video data stream 22 to the decoder 12. In particular, since the network 14 is assumed to be capable of supporting low-latency applications, the network 10 checks the decoder buffer lookup time to send a sequence 34 of packets of the video data stream 22 to the coded picture buffer 48 of the decoder 12, so that each decoded unit is present in the coded picture buffer 48 prior to the decoder buffer lookup time signaled by the timing control packet preceding the respective decoded unit. This measurement allows the decoder to use the decoder buffer search time of the timing control packets to empty the decoder's coded picture buffer 48 in units of decoding units rather than complete access units without stalling, i.e., without running out of payload packets available in the coded picture buffer 48. Figure 14, for example, illustrates an exemplary processing unit 50 connected to the output of the coded picture buffer 48 and having its input receiving the video data stream 22. Similar to the encoder 10, the decoder 12 can perform parallel processing using, for example, tile parallel processing / decoding and / or WPP parallel processing / decoding.

[0147] As will be described in more detail below, the decoder buffer search time is not necessarily related to the search time for the coded picture buffer 48 of the decoder 12. Rather, timing control packets can additionally or alternatively advance the search of already decoded image data in the corresponding decoded picture buffer of the decoder 12. FIG. 14 illustrates, for example, the decoder 12 including a decoder picture buffer in which decoded versions of video content obtained by decoding the video data stream 22 by the processing unit 50 are buffered and thus stored and output in units of decoded versions of decoded units. The decoder's decoded picture buffer 22 is thus connected between the output of the decoder 12 and the output of the processing unit 50. Having the ability to set the search time for outputting decoded versions of decoded units from the decoded picture buffer 52 gives the encoder 10 the opportunity to control playback on the fly, i.e., during encoding of the current picture, or between terminals, to achieve delays in the reproduction of the video content at the decoding end, even at a granularity smaller than the picture rate or frame rate. Obviously, excessive division of each image 18 into a large number of subportions 24 at the encoding side will have a negative impact on the bit rate for transmitting the video data stream 22. However, the time required to encode, transmit, decode, and output such a decoding unit is minimized, and thus end-to-end delay can be minimized. On the other hand, increasing the size of the subportions 24 increases the end-to-end delay. Therefore, a compromise must be found. Using the aforementioned decoder buffer search time to guide the output timing of the decoded version of the subportion 24 in terms of the decoding unit allows the encoder 10 or some other unit at the encoding side to apply this compromise spatially beyond the current image content. By this means, it is possible to control the end-to-end delay in such a way that it varies spatially across the current image content.

[0148] In implementing the above-described embodiment, it is possible to use packets of a removable packet type as timing control packets. Packets of the removable packet type are not necessary for recovering the video content at the decoding side. Hereinafter, this type of packet is referred to as an SEI packet. In addition, there are other types of removable packets as well, such as packets of the removable packet type, i.e., redundant packets, when transmitted in a stream. Alternatively, the timing control packet may be a packet of a specific removable packet type, but additionally carry a specific SEI packet type field. For example, the timing control packet may be an SEI packet, with each SEI packet carrying one or several SEI messages, and only those SEI packets containing a specific type of SEI message form the aforementioned timing control packet.

[0149] 12 to 14 may, according to further embodiments, be applied to the HEVC standard, thereby forming a possible concept for providing HEVC that is more effective in achieving lower end-to-end delay, where the aforementioned packets are formed by NAL units and the aforementioned payload packets are VCL NAL units of a NAL unit stream having slices that form the aforementioned sub-portions.

[0150] Before describing the more detailed embodiments, a further embodiment is presented that corresponds to the embodiment outlined above in that dispersed packets are used to transmit information representing a video data stream in an efficient manner, but the classification of information differs from the embodiment described above in that timing control packets transmit decoder buffer search timing information. In a further embodiment described below, the type of information transferred via dispersed packets dispersed to payload packets belonging to an access unit relates to region of interest (ROI) information and / or tile identification information. The further embodiment described below may or may not be combined with the embodiments described with reference to Figures 12-14.

[0151] FIG. 15 illustrates an encoder 10 that operates similarly to that described above with respect to FIG. 12, except for the distribution of timing control packets and functions described above with respect to FIG. 13, which are optional for the encoder 10 of FIG. 15. However, the encoder 10 of FIG. 15 is configured to encode video content 16 into a video data stream 22 in units of sub-portions 24 of images 18 of the video content 16, as described above with respect to FIG. 11. In encoding the video content 16, the encoder 10 is interested in conveying information about a region of interest (ROI) 60 to the decoding side along with the video data stream 22. The ROI 60 is a particular spatial sub-area of ​​the current image 18 to which the decoder should, for example, pay special attention. The spatial location of the ROI 60 can be input to the encoder 10 externally, as indicated by dotted line 62, such as by user input during the encoding of the current image 18, or can be automatically determined by the encoder 10 or some other entity on the fly. In any case, the encoder 10 faces the following challenge: Indicating the location of the ROI 60 is not, in principle, a problem for the encoder 10. Therefore, the encoder 10 can easily indicate the location of the ROI 60 in the data stream 22. However, to make this information easily accessible, the encoder 10 of FIG. 15 uses a distribution of ROI packets among the payload packets of an access unit, thereby allowing the encoder 10 to freely and continuously select the subportions 24 and / or the number of payload packets into which the subportions 24 are packetized that are spatially outside and spatially inside the ROI 60. Using distributed ROI packets, any network entity can easily identify the payload packets that belong to the ROI. On the other hand, using a removable packet type for these ROI packets allows them to be easily ignored by any network entity.

[0152] FIG. 15 illustrates an embodiment for distributing ROI packets 64 among payload packets 32 of an access unit 30. The ROI packets 64 indicate where within a sequence 34 of packets of the video data stream 22, coded data related to, i.e., encoding, the ROI 60 is contained. The manner in which the ROI packets 64 indicate the location of the ROI 60 can be implemented in a variety of ways. For example, the mere presence / occurrence of an ROI packet 64 indicates the incorporation of coded data related to the ROI 60 within one or more of the following payload packets 32 that follow, i.e., belong to the preceding payload packet in sequential order within the sequence 34. Alternatively, a syntax element within the ROI packet 64 indicates whether one or more subsequent payload packets 32 are related to, i.e., at least partially encode, the ROI 60. Variation also arises from possible variations regarding the "scope" of each ROI packet 64, i.e., the number of preceding payload packets preceded by one ROI packet 64. For example, the indication of inclusion or non-inclusion of some coded data related to ROI 60 within one ROI packet may be related to the payload packets 32 that follow in sequential order in sequence 34 until the occurrence of the next ROI packet 64, or simply the payload packets 32 that immediately follow, i.e., the payload packets 32 that immediately follow the respective ROI packet 64 in sequential order in sequence 34. In FIG. 15, graph 66 exemplarily illustrates the case where an ROI packet 64 indicates ROI association, i.e., the inclusion of any coded data related to ROI 60, or ROI non-association, i.e., the absence of any coded data related to ROI 60, relative to all payload packets 32 that occur downstream of the respective ROI packet 64 until the occurrence of the next ROI packet 64, or whatever occurs earlier along the packet sequence 34, relative to the end of the current access unit 30. In particular, FIG. 15 illustrates a case in which an ROI packet 64 has syntax elements therein that indicate whether or not subsequent payload packets 32 in the sequential order of the packet sequence 34 have encoded data relating to the ROI 60 therein.Such an embodiment is also described below. However, another possibility, as just mentioned, is for each ROI packet 64 to simply indicate by its presence in a packet sequence 34 that payload packets 32 that belong to the "scope" of the respective ROI packet 64 have the associated ROI 60 therein, i.e., data related to the ROI 60. According to an embodiment described in more detail below, the ROI packet 64 even indicates the location of the portion of the ROI 60 encoded in the payload packets 32 that belong to its "scope."

[0153] For example, any network entity 68 receiving video data stream 22 can utilize the ROI-related indication, as implemented using ROI packets 64, to treat ROI-related portions of sequence of packets 34, e.g., with higher priority than other portions of sequence of packets 34. Alternatively, network entity 68 can use ROI-related information to perform other operations related to the transmission of video data stream 22. Network entity 68 can be, for example, a MANE or decoder for decoding and playback of video content 60 as transmitted via video data streams 22, 28. In other words, network entity 68 can use the results of ROI packet identification to determine transmission operations related to the video data stream. Transmission operations can include retransmission requests for defective packets. Network entity 68 can be configured to treat regions of interest 70 with increased priority and assign higher priority to ROI packets 72 and their accompanying payload packets, i.e., those prepended thereby, that are signaled to cover the region of interest, compared to ROI packets and their accompanying payload packets that are signaled not to cover the ROI. Network entity 68 can first request the transfer of payload packets that have a higher priority assigned to it before requesting retransmission of some of the payload packets that have a lower priority assigned to it.

[0154] The embodiment of Figure 15 can be easily combined with the embodiments previously described with respect to Figures 12-14. For example, the ROI packet 64 described above may be an SEI packet having a particular type of SEI message included therein, i.e., an ROI SEI message. That is, an SEI packet may be, for example, a timing control packet and an ROI packet at the same time, i.e., where each SEI packet includes both timing control information and ROI display information. Alternatively, an SEI packet may be one of a timing control packet and an ROI packet, rather than the other, and may not be an ROI packet or a timing control packet.

[0155] In accordance with the embodiment illustrated in FIG. 16, the distribution of packets among the payload packets of an access unit is used to indicate in an easily accessible manner to the network entity 68 handling the video data stream 22, the tiles of the current image 18 to which the current access unit 30 is associated, each packet being overlaid with several subportions encoded into some of the payload packets 32 that serve as prefixes. In FIG. 16, for example, the current image 18 is shown subdivided into four tiles 70, here formed by the four quadrants of the current image 18 for illustrative purposes. The subdivision of the current image 18 into tiles 70 can also be signaled, for example, within the video data stream of units comprising the sequence of images, e.g., by VPS or SPS packets distributed among the sequence of packets 34. As will be explained in more detail below, the tile subdivision of the current image 18 can be a regular subdivision of the image 18 into columns and rows of tiles. The number of columns and rows, as well as the width of the columns and height of the rows of tiles, can vary. In particular, the width and height of columns / rows of tiles may be different for different rows and columns. FIG. 16 additionally illustrates an embodiment in which the subportions 24 are slices of the image 18. The slices 24 subdivide the image 18. As outlined in more detail below, the subdivision of the image 18 into slices 24 is subject to the constraint that each slice 24 can be entirely contained within one tile 70 or entirely cover two or more tiles 70. FIG. 16 illustrates a case in which the image 18 is subdivided into five slices 24. The first four slices 24 in the decoding order described above cover the first two tiles 70, while the fifth slice entirely covers the third and fourth tiles 70. Furthermore, FIG. 16 also illustrates a case in which each slice 24 is individually encoded into a respective payload packet 32. Furthermore, FIG. 16 also illustrates, by way of example, a case in which each payload packet 32 ​​is preceded by a preceding tile identification packet 72.Each tile identification packet 72 then indicates for its immediately following payload packet 32 ​​which of the tiles 70 the subportion 24 encoded in that payload packet 32 ​​covers. Thus, the first two tile identification packets 72 within an access unit 30 for the current image 18 indicate the first tile, while the third and fourth tile identification packets 72 indicate the second tile 70 of the image 18, and the fifth tile identification packet 72 indicates the third and fourth tiles 70. With respect to the embodiment of FIG. 16, for example, the same variations are possible as described above with respect to FIG. 15. That is, the "scope" of a tile identification packet 72 could, for example, only include the first immediately following payload packet 32 ​​or the immediately following payload packet 32 ​​until the occurrence of the next tile identification packet.

[0156] Encoder 10 can be configured to encode each tile 70 such that no spatial prediction or context selection occurs across tile boundaries for the tile. Encoder 10 can, for example, encode tiles 70 in parallel. Similarly, any decoder, such as network entity 68, can decode tiles 70 in parallel.

[0157] The network entity 68, which may be a MANE or a decoder or some other device between the encoder 10 and the decoder, can be configured to use the information conveyed by the tile identification packets 72 to determine specific transmission operations. For example, the network entity 68 can treat specific tiles of the current image 18 of the video 16 with a higher priority, i.e., forward each payload packet previously identified as relating to such tiles or using more secure FEC protection, etc. In other words, the network entity 68 can use the results of the identification to determine transmission operations related to the video data stream. Transmission operations can include retransmission requests for packets received in a defect state—i.e., exceeding any FEC protection, if any, of the video data stream. The network entity can, for example, treat different tiles 70 with different priorities. For this purpose, the network entity can assign a higher priority to tile identification packets 72 and their payload packets, i.e., precede tile identification packets 72 and their payload packets associated with lower priority tiles, thereby associated with higher priority tiles. Before requesting the transfer of any payload packets that have a lower priority assigned to them, network entity 68 may, for example, first request the transfer of payload packets that have a higher priority assigned to them.

[0158] The embodiments described so far can be incorporated into the HEVC framework as explained in the introductory part of the specification of this application, as explained below.

[0159] In particular, SEI messages can be assigned to slices of a decoding unit in the sub-picture CPB / HRD case. That is, buffering period and timing SEI messages can be assigned to NAL units containing slices of a decoding unit. This can be achieved by a new NAL unit type, which is a non-VCL NAL unit that can directly precede one or more slice / VCL NAL units in the decoder. This new NAL unit can be called a slice prefix NAL unit. Figure 17 illustrates the structure of an access unit, omitting any temporal NAL units for the ends of sequences and streams.

[0160] According to FIG. 17, the access unit 30 is structured as follows: In the sequential order of packets in the sequence of packets 34, the access unit 30 can begin with the occurrence of a special type of packet, namely, an access unit delimiter 80. Then, one or more SEI packets 82 of the SEI packet type for all access units can follow within the access unit 30. Both packet types 80 and 82 are optional; that is, packets of this type cannot occur within the access unit 30. Then, the sequence of decoding units 38 follows. Each decoding unit 38 optionally begins with a slice prefix NAL unit 84, which may contain, for example, timing control information, or, according to the embodiment of FIG. 15 or 16, ROI information or tile information or, more generally, a respective subpicture SEI message 86. Then, the actual slice data 88 of the respective payload packet or VCL NAL unit follows, as shown in 88. Thus, each decoding unit 38 includes a sequence of slice prefix NAL units 84 followed by respective slice data NAL units 88. The bypass arrow 90 in FIG. 17 that bypasses the slice prefix NAL unit indicates that in the case of a non-decoded unit subdivision of the current access unit 30, the slice prefix NAL unit 84 is not present.

[0161] As already mentioned above, all information signaled in the slice prefix and the accompanying sub-picture SEI message may be effective for all VCL NAL units of the access unit or up to the occurrence of the second prefix NAL unit, or for the following VCL-NAL units in decoding order, depending on the flags conveyed in the slice prefix NAL unit.

[0162] A slice VCL NAL unit for which the information signaled in the slice prefix is ​​valid is referred to below as a prefix slice. A prefix slice associated with a single slice it prefixes does not necessarily constitute a complete decoding unit, but can be part of it. However, a single slice prefix cannot be valid for multiple decoding units (subpictures), and the beginning of a decoding unit is signaled in the slice prefix. If no means for signaling are provided by the slice prefix syntax (such as the "simple syntax" / Version 1 shown below), the occurrence of a slice prefix NAL unit signals the beginning of a decoding unit. Only certain SEI messages (identified via payload types in the syntax description below) can be sent within a slice prefix NAL unit, exclusively at the subpicture level, while some SEI messages can be sent in slice prefix NAL units at the subpicture level or as regular SEI messages at the access unit level.

[0163] Additionally or alternatively, the Tile ID SEI message / Tile ID signaling can be implemented in a high-level syntax, as described above with respect to Figure 16. In earlier designs of HEVC, the slice header / slice data includes identifiers for the tiles contained in each slice. For example, slice data semantics may indicate the following:

[0164] tile_idx_minus_1 specifies the TileID in raster scan order. The first tile in an image has a TileID of 0. The value of tile_idx_minus_1 is in the range of 0 to (num_tile_columns_minus1+1)*(num_tile_rows_minus1+1)-1.

[0165] If tiles_or_entropy_coding_sync_idc is equal to 1, this parameter is not considered useful since this ID is easily derived from the slice address and from the slice dimension as signaled in the picture parameter set.

[0166] Although the tile ID can be derived implicitly during the decoding process, knowledge of this parameter at the application layer is important for different use cases, such as a videoconferencing scenario where different tiles have different priorities for playback (these tiles typically form a region of interest containing a speaker using a conversational use case) with one tile having a higher priority than another. In the event of loss of network packets in the transmission of multiple tiles, those network packets containing tiles representing the region of interest can be retransmitted with a higher priority to maintain a higher quality of experience at the receiver terminal than in the case of retransmitted tiles without any priority ordering. Other use cases allow tiles to be assigned to different screens if their dimensions and their positions are known, such as in a videoconferencing scenario.

[0167] In order for such an application layer to be able to handle tiles with a particular priority in a transmission scenario, the tile_id can be provided as a subpicture or slice-specific SEI message, or in a special NAL unit before one or more NAL units of the tile, or in a special header part of the NAL unit belonging to the tile.

[0168] As described above with respect to Figure 15, a region of interest SEI message may also or instead be provided. Such an SEI message may allow signaling of a region of interest (ROI), in particular in signaling the ROI to which a particular tile_id / tile belongs. The message may allow giving a priority of the region of interest in addition to the region of interest ID.

[0169] FIG. 18 illustrates the use of tiles for region of interest signaling.

[0170] In addition to what was described above, slice header signaling can be implemented. The slice prefix NAL unit can also contain slice headers for the following dependent slices, i.e., the slices prefixed by the respective slice prefix. If slice headers are only provided in the slice prefix NAL unit, the actual slice type needs to be derived by the NAL unit type of the NAL units containing the respective dependent slices, or by a flag in the slice prefix that signals whether the next slice data belongs to a slice type that serves as a random access point.

[0171] Additionally, the slice prefix NAL unit can carry slice or subpicture specific SEI messages to convey any information, such as subpicture timing or tile identifiers. Any subpicture specific messaging is not supported in the HEVC specification, as described in the introductory section of the specification of this application, but may be important for certain applications.

[0172] Next, possible syntaxes for implementing the above-mentioned concept of slice prefixing are described, in particular which changes can be made at the slice level when using the HEVC context outlined in the introductory part of the specification of this application as a basis.

[0173] In particular, in the following, two versions of the possible slice prefix syntax are given, one with functionality for SEI message dispatch only, and one with extended functionality for signaling part of the slice header for the following slice. The first simple syntax / version 1 is shown in Figure 19.

[0174] As a preliminary remark, Figure 19 thus illustrates a possible implementation for carrying out any of the embodiments described above with respect to Figures 11-16. The distributed packets shown therein can be interpreted as shown in Figure 19, which is described in more detail below by way of specific embodiments.

[0175] The extended syntax / version 2, which includes tile_id signaling, decoding unit start identifier, slice prefix ID and slice header data apart from the SEI message concept, is given in the table of Figure 20.

[0176] The semantics can be defined as follows: A rap_flag with a value of 1 indicates that the access unit containing the slice prefix is ​​a RAP picture. A rap_flag with a value of 0 indicates that the access unit containing the slice prefix is ​​not a RAP picture.

[0177] The decoding_unit_start_flag indicates the start of the decoding unit within the access unit, so that the following slice continues to the end of the access unit or to the start of another decoding unit that belongs to the same decoding unit.

[0178] single_slice_flag with a value of 0 indicates that the information given within the prefix slice NAL unit and the accompanying subpicture SEI message applies to all following VCL-NAL units until the next access unit, the occurrence of another slice prefix, or the beginning of another complete slice header. single_slice_flag with a value of 1 indicates that all information indicated in the slice prefix NAL unit and the accompanying subpicture SEI message is valid only for the next VCL-NAL unit in decoding order.

[0179] tile_idc indicates the amount of tiles present in the following slice. A tile_idc equal to 0 indicates that no tiles are used in the following slice. A tile_idc equal to 1 indicates that a single tile is used in the following slice, and that tile identifier is signaled accordingly. A tile_idc with a value of 2 indicates that multiple tiles are used within the following slice, and so the number of tiles and the first tile identifier are signaled accordingly.

[0180] prefix_slice_header_data_present_flag indicates that slice header data corresponding to the following slice in decoding order is signaled in the given slice prefix.

[0181] slice_header_data() is defined after the text. It contains the associated slice header information, and if dependent_slice_flag is set equal to 1, it is not covered by the slice header.

[0182] Note that separating the slice header and the actual slice data allows for more flexible transmission schemes for the header and slice data.

[0183] num_tiles_in_prefixed_slices_minus1 indicates the number of tiles used in the next decoding unit minus 1.

[0184] first_tile_id_in_prefixed_slices indicates the tile identifier of the first tile in the next decoding unit.

[0185] For slice prefix simple syntax / version 1, the following syntax elements, when not present, can be set to default values ​​as follows: -decoding_unit_start is equal to 1, i.e. the slice prefix always indicates the start of a decoding unit. -single_slice_flag is equal to 0, i.e., the slice prefix applies to all slices of the decoding unit.

[0186] The slice prefix NAL unit has 24 NAL unit types, and the NAL unit types are proposed to be overviewed in the table extended according to FIG.

[0187] Briefly, to summarize Figures 19-21, the syntax details found therein indicate that a particular packet type is attributed to the above-identified distributed packet, here exemplarily showing NAL unit type 24. Furthermore, in particular, the syntax example of Figure 20 reveals the "scope" of the distributed packets, the switching mechanism controlled by respective syntax elements within these distributed packets themselves, where single_slice_flag is used to control this scope, i.e., the above-mentioned options for switching between different options for this scope definition, respectively. Furthermore, it has been revealed that the above-described embodiments of Figures 1-16 can be extended in that the distributed packets include common slice header data for slices 24 contained in packets belonging to the "scope" of each distributed packet. That is, it can be said that there is a mechanism controlled by respective flags within these distributed packets that indicates whether common slice header data is included within each distributed packet.

[0188] Of course, as specified in the current version of HEVC, the just presented concept of which parts of the slice header data are moved to the slice header prefix requires changes to the slice header. The table in Figure 22 shows a possible syntax for such a slice header, where certain syntax elements present in the slice header according to the current version are shifted to a lower-level syntax element called slice_header_data(). This syntax for the slice header and slice header data only applies to the option where the extended slice header prefix NAL unit concept is used.

[0189] In Figure 22, slice_header_data_present_flag indicates that the slice header data for the current slice is predicted from the value signaled in the last slice prefix NAL unit of the access unit, i.e., the most recently occurring slice prefix NAL unit.

[0190] All syntax elements apart from the slice header are signaled by the syntax element slice header data as given in the table of Figure 23.

[0191] That is, transferring the concepts of Figures 22 and 23 to the embodiments of Figures 12-16, the distributed packets described therein can be extended by incorporating into these distributed packets a subportion 24 of the slice header syntax of the slices encoded in the payload packets, i.e., VCL NAL units. The incorporation may be optional. That is, each syntax element of the distributed packets can indicate whether such slice header syntax is included in the respective distributed packet. If included, the respective slice header data incorporated into the respective distributed packet can apply to all slices contained in packets belonging to the "scope" of the respective distributed packet. Whether the slice header data included in a distributed packet is introduced by a slice encoded in any of the payload packets belonging to the scope of this distributed packet can be signaled by a respective flag, e.g., slice_header_data_present_flag in Figure 22. This measure compacts the slice headers of the slices that are encoded into packets that belong to the "scope" of the respective distributed packet, and therefore, using the just-mentioned flags in the slice, any decoder that receives the slice header and video data stream, such as the network entities shown in Figures 12 to 16 above, responds to the just-mentioned flags in the slice header of the slice by copying the slice header data embedded in the distributed packet to the slice header of the respective flag within the scope of the slice that signals the movement of the slice header data to the slice prefix, i.e., for each distributed packet, into the slice header of the slice that is encoded into the payload packets that belong to the scope of this distributed packet.

[0192] Further continuing with the example syntax for implementing the embodiments of Figures 12-16, the SEI message syntax can be as shown in Figure 24. To introduce slice or sub-picture SEI message types, the SEI payload syntax can be configured as shown in the table of Figure 25. Only SEI messages with payloadType in the range of 180 to 184 can be sent exclusively at the sub-picture level within the slice prefix NAL unit. Furthermore, region of interest SEI messages with payloadType equal to 140 can be sent either in slice prefix NAL units on the sub-picture level or in regular SEI messages on the access unit level.

[0193] That is, in transferring the details shown in Figures 25 and 24 onto the embodiments described above with respect to Figures 12-16, the distributed packets shown in these embodiments of Figures 12-16 can be realized using slice prefix NAL units with a specific NAL unit type, e.g., 24, that includes a specific type of SEI message within the slice prefix NAL unit, e.g., signaled by payloadType at the beginning of each SEI message. In the specific syntax examples currently described, payloadType=180 and payloadType=181 result in a timing control packet according to the embodiment of Figures 11-14, while payloadType=140 results in an ROI packet according to the embodiment of Figure 15, and payloadType=182 results in a tile identification packet according to the embodiment of Figure 16. The specific syntax examples described hereinafter can include only one or a subset of the just-mentioned payloadType options. Beyond this, Figure 25 makes clear that any of the above-described embodiments of Figures 11-16 can be combined with each other. Furthermore, Figure 25 makes it clear that the embodiments of Figures 12 to 16 or a combination thereof can also be extended by distributed packets, which are then described with payloadType=184. As already mentioned above, the extensions described below with respect to payloadType=183 result in the possibility that any distributed packet could incorporate common slice header data for the slice headers of slices encoded in any payload packets belonging to its scope.

[0194] The following table defines the SEI messages that can be used at the slice or subpicture level. Areas of interest for SEI messages that can be used at the subpicture and access unit levels are also shown.

[0195] Figure 26, for example, shows an example for a subpicture buffer SEI message that occurs whenever a slice prefix NAL unit of NAL unit type 24 has an SEI message type 180 included therein, thus forming a timing control packet.

[0196] The semantics can be defined as follows: seq_parameter_set_id identifies the sequence parameter set that contains the sequence HRD attribute. The value of seq_parameter_set_id is equal to the value of seq_parameter_set_id of the picture parameter set referenced by the primary coded picture associated with the buffering period SEI message. The value of seq_parameter_set_id ranges from 0 to 31, including the following:

[0197] initial_cpb_removal_delay[SchedSelIdx] and initial_alt_cpb_removal_delay[SchedSelIdx] specifies the initial CPB removal delay for the SchedSelIdx-th CPB of a decoding unit (subpicture). The syntax element has a length in bits given by initial_cpb_removal_delay_length_minus1+1, in units of 90 kHz clock. The value of the syntax element shall not be equal to 0 and shall not exceed 90000 * (CpbSize[SchedSelIdx] ÷ BitRate[SchedSelIdx]), the time equivalent of the CPB size in 90 kHz clock units.

[0198] Across all coded video sequences, the sum of initial_cpb_removal_delay[SchedSelIdx] and initial_cpb_removal_delay_offset[SchedSelIdx] per decoding unit (subpicture) is constant for each value of SchedSelIdx, and the sum of initial_alt_cpb_removal_delay[SchedSelIdx] and initial_alt_cpb_removal_delay_offset[SchedSelIdx] is constant for each value of SchedSelIdx.

[0199] Figure 27 similarly shows an example for the Sub-picture Timing SEI message, and the semantics can be described as follows:

[0200] du_dpb_output_delay specifies how many clock ticks to wait after the removal from the CPB of the decoding unit (subpicture) associated with the most recent Subpicture Buffering Period SEI message of the preceding access unit, if any, before removing from the buffer the decoding unit (subpicture) data associated with the most recent Subpicture Timing SEI message of the preceding access unit for the same decoding unit (subpicture). This value is also used to calculate the earliest possible arrival time of the decoding unit (subpicture) data into the CPB for HSS (Hypothetical Stream Scheduler [2] 0). The syntax element is a fixed-length code whose length in bits is given by cpb_removal_delay_length_minus1 + 1. cpb_removal_delay is a modulo 2 (cpb_removal_delay_length_minus1+1) This is the remainder of the counter.

[0201] du_dpb_output_delay is used to calculate the DPB output time of a decoding unit (subpicture). It specifies how many clock ticks to wait after removal of a decoded decoding unit (subpicture) from the CPB before the decoding unit (subpicture) of the picture is output from the DPB.

[0202] Note that this allows for updating of sub-pictures. In such a scenario, the decoding units that are not updated remain unchanged in the final decoded image, and they remain visible.

[0203] 26 and 27 and transfer the specific details contained therein to the embodiments of FIGS. 12-14, it can be said that the decoder buffer search time for a decoding unit can be signaled in the associated timing control packet of the encoding method differentially, i.e., additionally relative to other decoder buffer search times. That is, to obtain the decoder buffer search time for a particular decoding unit, a decoder receiving a video data stream adds the decoder search time obtained from the timing control packet preceding the particular decoding unit to the decoder search time of the immediately preceding decoding unit, i.e., the one preceding the particular decoding unit and the decoding unit following it, in this method procedure. At the beginning of each several pictures or parts thereof of the encoded video sequence, the timing control packet may additionally or alternatively include a fully encoded decoder buffer search time value differentially compared to the decoder buffer search time value of the preceding decoding unit.

[0204] Figure 28 shows what the Sub-picture Slice Information SEI message looks like. The semantics can be defined as follows:

[0205] slice_header_data_flag with a value of 1 indicates that slice header data is present in the SEI message. The slice header data given in the SEI is valid for all slices that follow the decoding order up to the end of the access unit, occurrences of slice data in other SEI messages, slice NAL units, or slice prefix NAL units.

[0206] Figure 29 shows an example for a Subpicture Tile Information SEI message, the semantics can be defined as follows:

[0207] tile_priority indicates the priority of all tiles in the prefix slice that inherit the decoding order. The value of tile_priority is in the range 0 to 7 inclusive, with 7 indicating the highest priority.

[0208] A multiple_tiles_in_prefixed_slices_flag with a value of 1 indicates that there are more than one tile in the prefix slice to carry over the decoding order. A multiple_tiles_in_prefixed_slices_flag with a value of 0 indicates that the next prefix slice contains only one tile.

[0209] num_tiles_in_prefixed_slices_minus1 indicates the number of tiles in the prefix slice that inherit the decoding order.

[0210] first_tile_id_in_prefixed_slices indicates the tile_id of the first tile in the prefix slice that inherits the decoding order.

[0211] That is, the embodiment of FIG. 16 can be implemented using the syntax of FIG. 29 to realize the tile identification packet referred to in FIG. 16. As shown therein, a specific flag, here multiple_tiles_in_prefixed_slices_flag, may be used to indicate within a distributed tile identification packet whether only one tile or multiple tiles are covered by which subportion of the current image 18 encoded in any of the payload packets belonging to the range of the respective distributed tile identification packet. If the flag signals covering multiple tiles, a further syntax element is included in each distributed packet, here illustratively num_tiles_in_prefixed_slices_minus1, indicating the number of tiles covered by any subportion of any payload packet belonging to the range of the respective distributed tile identification packet. Finally, a further syntax element, here illustratively first_tile_id_in_prefixed_slices, indicates the ID of the tile among the number of tiles indicated by the current distributed tile identification packet, which is the first one according to decoding order. Transferring the syntax of Figure 29 to the example of Figure 16, the tile identification packet 72 preceding the fifth payload packet 32 ​​has all three just-mentioned syntax elements, for example, with multiple_tiles_in_prefixed_slices_flag set to 1, num_tiles_in_prefixed_slices_minus1 set to 1, thereby indicating that two tiles belong to the current scope, and first_tile_id_in_prefixed_slices set to 3, and indicating that the progression of tiles in decoding order belonging to the scope of the current tile identification packet 72 begins with the third tile (with tile_id = 2).

[0212] 29 also reveals that the tile identification packet 72 may also indicate a tile_priority, i.e., the priority of the tiles belonging to its scope. Similar to the ROI aspect, the network entity 68 may use such priority information to control transmission operations, such as requesting the transfer of specific payload packets.

[0213] Figure 30 shows an example syntax for the Subpicture Tile Range Information SEI message, and the semantics can be defined as follows:

[0214] A multiple_tiles_in_prefixed_slices_flag with a value of 1 indicates that there is more than one tile in the prefix slice to inherit the decoding order. A multiple_tiles_in_prefixed_slices_flag with a value of 0 indicates that the following prefix slice contains only one tile.

[0215] num_tiles_in_prefixed_slices_minus1 indicates the number of tiles in the prefix slice to inherit the decoding order.

[0216] tile_horz_start[i] indicates the horizontal start of the ith tile of pixels within the image.

[0217] tile_width[i] indicates the width of the ith tile in pixels within the image.

[0218] tile_vert_start[i] indicates the horizontal start of the ith tile of pixels within the image.

[0219] tile_height[i] indicates the height of the ith tile in pixels within the image.

[0220] Note that the tile range SEI message is used for display operations, e.g., to allocate tiles to screens in a multiple screen display scenario.

[0221] Figure 30 thus reveals that the example syntax implementation of Figure 29 for the Tile Identification Packet of Figure 16 may be varied in that tiles belonging to the scope of each Tile Identification Packet are indicated by their location within the current image 18 rather than their Tile ID. That is, rather than signaling the Tile ID of the first tile in decoding order covered by each subportion encoded in any of the payload packets belonging to the scope of each distributed Tile Identification Packet, its location can be signaled, for example, by the top left corner location of each tile i, here illustratively tile_horz_start and tile_vert_start, and the width and height of tile i, here illustratively tile_width and tile_height.

[0222] An example syntax for a region of interest SEI message is shown in Figure 31. For more precision, Figure 32 shows a first variant. In particular, the region of interest SEI message can be used, for example, at the access unit level or subpicture level to signal one or more regions of interest. According to the first variant of Figure 32, if multiple ROIs are within the current scope, each ROI is signaled once per ROI SEI message, rather than signaling all ROIs in the scope of each ROI packet within one ROI SEI message.

[0223] According to Figure 31, the region of interest SEI message signals each ROI individually. The semantics can be defined as follows:

[0224] roi_id indicates the identifier of the region of interest.

[0225] roi_priority indicates the priority of all tiles belonging to the region of interest in the prefix slice or all tiles belonging to all slices, depending on whether the SEI message is sent at the subpicture level or the access unit level, for inheritance of decoding order. The value of roi_priority ranges from 0 to 7, inclusive, where 7 indicates the highest priority. If both roi_priority in the roi information SEI message and tile_priority in the subpicture tile information SEI message are given, the highest value of both applies to the priority of the individual tile.

[0226] num_tiles_in_roi_minus1 indicates the number of tiles in the prefix slice that inherit the decoding order that belong to the region of interest.

[0227] roi_tile_id[i] indicates the tile_id of the ith tile belonging to the region of interest for the prefix slice that inherits the decoding order.

[0228] That is, FIG. 31 shows that an ROI packet such as that shown in FIG. 15 can signal the ID of the region of interest referenced by each ROI packet and the payload packets belonging to its scope. Optionally, an ROI priority index can be signaled along with the ROI ID. However, both syntax elements are optional. The syntax element num_tiles_in_roi_minus1 can then indicate the number of tiles in each ROI packet's scope that belong to the respective ROI 60. The roi_tile_id then indicates the tile-ID of the i-th tile belonging to the ROI 60. For example, if image 18 is subdivided into tiles 70 in the manner shown in FIG. 16, ROI 60 in FIG. 15, corresponding to the image corresponding to the left half of image 18, is formed by the first and third tiles in decoding order. The ROI packet can then be placed before the first payload packet 32 ​​of the access unit 30 in FIG. 16, followed by further ROI packets between the fourth and fifth payload packets 32 of this access unit 30. Then the first ROI packet has num_tile_in_roi_minus1 set to 0 and roi_tile_id[0] set to 0 (thereby directing attention to the first tile in decoding order), and the second ROI packet before the fifth payload packet 32 ​​has num_tiles_in_roi_minus1 set to 0 with roi_tile_id[0] set to 2 (thereby referring to the third tile in decoding order in the bottom left quarter of the image 18).

[0229] According to a second variant, the syntax of the region of interest in an SEI message can be as shown in Figure 32, where all ROIs in a single SEI message are signaled. In particular, the same syntax as described above with respect to Figure 31 is used, but the syntax element for each ROI of multiple ROIs referenced by each ROI SEI message or ROI packet is multiplied by a number signaled by a syntax element, here illustratively num_rois_minus1. Optionally, a further syntax element, here illustratively roi_presentation_on_separate_screen, can indicate for each ROI whether each ROI is suitable for being presented on a separate screen.

[0230] The semantics can be as follows:

[0231] num_rois_minus1 indicates the number of regular slices that inherit the ROI or decoding order of the prefix slice.

[0232] roi_id[i] indicates the identifier of the i-th region of interest.

[0233] roi_priority[i] indicates the priority of all tiles belonging to the i-th region of interest in the prefix slice or all tiles belonging to all slices, depending on whether the SEI message is sent at the subpicture level or the access unit level, for inheritance of decoding order. The value of roi_priority ranges from 0 to 7, inclusive, where 7 indicates the highest priority. If both roi_priority in the roi information SEI message and tile_priority in the subpicture tile information SEI message are given, the highest value of both applies to the priority of the individual tile.

[0234] num_tiles_in_roi_minus1[i] indicates the number of tiles in the prefix slice that inherit the decoding order belonging to the i-th region of interest.

[0235] roi_tile_id[i][n] indicates the tile_id of the nth tile belonging to the i-th region of interest for the prefix slice that inherits the decoding order.

[0236] roi_presentation_on_seperate_screen[i] indicates that the region of interest associated with the ith roi_id is suitable for presentation on a separate screen.

[0237] Thus, briefly summarizing the various embodiments described so far, a higher-level syntax signaling strategy was presented that allows for the application of SEI messages as well as higher-level syntax items beyond those contained in the per-slice-level NAL unit header. Thus, we described the slice prefix NAL unit. The syntax and semantics of the slice prefix and slice_level / subpicture SEI messages were described, along with use cases for low-latency / subpicture CPB operation, tile signaling, and ROI signaling. An extended syntax was additionally presented for signaling the slice header of slices following the slice prefix.

[0238] For completeness, Figure 33 shows a further example for a syntax that can be used for timing control packets according to the embodiments of Figures 12-14. The semantics can be as follows:

[0239] du_spt_cpb_removal_delay_increment specifies the period, in clock subticks, between the last decoding unit in the decoding order of the current access unit and the nominal CPB time of the decoding unit associated with the decoding unit information SEI message. As specified in Annex C, this value is also used to calculate the earliest possible time of arrival of decoding unit data at the CPB for HSS. The syntax element is represented by a fixed-length code whose length in bits is given by du_cpb_removal_delay_increment_length_minus1+1. When the decoding unit associated with the decoding unit information SEI message is the last decoding unit of the current access unit, the value of du_spt_cpb_removal_delay_increment is equal to 0.

[0240] dpb_output_du_delay_present_flag equal to 1 specifies the presence of the pic_spt_dpb_output_du_delay syntax element in the decoding unit information SEI message. dpb_output_du_delay_present_flag equal to 0 specifies the absence of the pic_spt_dpb_output_du_delay syntax element in the decoding unit information SEI message.

[0241] When SubPicHrdFlag is equal to 1, pic_spt_dpb_output_du_delay is used to calculate the DPB output time of the picture. It specifies how many subclock ticks to wait after removal of the last decoding unit of an access unit from the CPB before the decoded picture is output from the DPB. When not present, the value of pic_spt_dpb_output_du_delay is inferred to be equal to pic_dpb_output_du_delay. The length of the syntax element pic_spt_dpb_output_du_delay is given in bits by dpb_output_delay_du_length_minus1+1.

[0242] It is a requirement for bitstream consistency that all decoding unit information SEI messages associated with the same access unit apply to the same operation point and have a dpb_output_du_delay_present_flag equal to 1 and the same value as pic_spt_dpb_output_du_delay. The output time resulting from the pic_spt_dpb_output_du_delay of any picture output from an output timing synchronization decoder precedes the output times resulting from the pic_spt_dpb_output_du_delay of all pictures of any next CVS in decoding order.

[0243] The picture output order determined by the value of this syntax element is the same order established by the value of PicOrderCntVal.

[0244] For pictures that are not output by the "bumping" process because they precede, in decoding order, an IRAP picture whose NoRaslOutputFlag is 1 and whose no_output_of_prior_pics_flag is equal to or found to be equal to 1, the output time resulting from pic_spt_dpb_output_du_delay increases with increasing values ​​of PicOrderCntVal for all pictures in the same CVS. For any two pictures in a CVS, the difference in output times of the two pictures when SubPicHrdFlag is equal to 1 is the same as the difference when SubPicHrdFlag is equal to 0.

[0245] Furthermore, Figure 34 shows a further embodiment for signaling an ROI region using an ROI packet. According to Figure 34, the syntax of the ROI packet contains only one flag indicating whether all sub-portions of the image 18 encoded in several payload packets 32 belong to its scope. The "scope" extends until the occurrence of the ROI packet or region_refresh_info SEI message. If the flag is 1, the region is encoded in each subsequent payload packet; if it is 0, the opposite applies, i.e., it indicates that the respective sub-portion of the image 18 does not belong to the ROI 60.

[0246] Before discussing some of the above-mentioned embodiments again, i.e., explaining some of the terms used above, such as tile, slice, WPP, substream, and subdivision, it should be noted that the high-level signaling of the above-mentioned embodiments can alternatively be defined in a transport specification, such as [3-7]. In other words, the above-mentioned packets and forming sequences 34 are transport packets, some of which have application layer subportions, such as slices, contained therein as packetized whole or fragmented, and some of which are distributed between the latter and the above-mentioned destinations in a manner. In other words, the above-mentioned distributed packets are not limited to being defined in the application layer video codec, but can instead be special transport packets defined in the transport protocol, such as another type of NAL unit SEI message.

[0247] In other words, according to one aspect of the specification, the embodiment provides a video data stream having video content encoded in units of sub-portions (see coding tree blocks or slices) of images of the video content, each sub-portion being encoded into one or more payload packets (see VCL NAL unit) of a sequence of packets (NAL units) of the video data stream, the sequence of packets being divided into a sequence of access units, whereby each access unit collects payload packets for a respective image of the video content, the sequence of packets having distributed therein timing control packets (slice prefixes) whereby the timing control packets subdivide the access units into decoding units, whereby at least some access units are subdivided into two or more decoding units, each timing control packet signaling a decoder buffer lookup time for a decoding unit, a payload packet followed by each timing control packet of the sequence of packets.

[0248] As mentioned above, the domain in which video content is coded into a data stream in units of sub-portions of an image can cover syntax elements relating to predictive coding, e.g. coding modes (e.g. intra mode, inter mode, subdivision information, etc.), prediction parameters (e.g. motion vectors, extrapolation direction, etc.) and / or residual data (e.g. transform coefficient levels), which are associated with local portions of an image, e.g. coding tree blocks, prediction blocks and residual (e.g. transform, etc.) blocks, respectively.

[0249] As noted above, payload packets can each contain one or more slices (completely, respectively). Slices may be independently decodable or may exhibit interrelationships that prevent their independent decoding. For example, entropy slices may be independently entropy decodable, but prediction across slice boundaries may be prohibited. Dependent slices may be encoded / decoded using WPP processing, i.e., entropy and predictive coding across slice boundaries, with the ability to encode / decode overlapping, time-dependent slices in parallel, but allowing for an alternating start encoding / decoding procedure for individual dependent slices and the encoding / decoding of slices referenced by the dependent slice.

[0250] The order in which the access unit payload packets are arranged within each access unit is known to the decoder in advance, e.g., the encoding / decoding order can be specified within a sub-portion of the image, such as the scan order within the coding tree blocks in the example above.

[0251] For example, see the following figures: A currently encoded / decoded image 100 can be divided into tiles, which correspond to quadrants of an image 110 as an example in Figures 35 and 36 and are indicated by reference numerals 112a to 112d. That is, the entire image 110 can form one tile, as in Figure 37, or can be divided into multiple tiles. The tile division can be restricted to a regular one, where tiles are arranged only in columns and rows. Different examples are shown below.

[0252] Thus, image 110 is further subdivided into coding (tree) blocks (small boxes in the figure, referred to above as CTBs) 114, within which a coding order 116 is defined (here, raster scan order, but it may be different). The subdivision of the image into tiles 112a-d can be constrained so that the tiles are disjoint from the blocks 114. Furthermore, the blocks 114 and tiles 112a-d can be constrained to be regular arrays of columns and rows.

[0253] If there are tiles (i.e. more than one), the raster in encoding (decoding) order 116 scans the complete tile first, then - in raster scan tile order - transitions to the next tile in tile order.

[0254] Because tiles are encoded / decoded independently of each other due to non-crossing of tile boundaries through spatial prediction and context selection inferred from spatial neighbors, the encoder 10 and decoder 12 can encode / decode images subdivided into tiles 112 (originally shown as 70) independently and in parallel, except for, for example, in-loop or post-filtering where crossing of tile boundaries is allowed.

[0255] Image 110 can be further subdivided into slices 118a-d, 180—originally designated by reference numeral 24. A slice can contain only a portion of a tile, one complete tile, or multiple tiles. Thus, the division into slices can also result in a subdivision of tiles, as in the case of FIG. 35. Each slice completely contains at least one coding block 114 and consists of consecutive coding blocks 114 in coding order 116, such that the order is determined within the slices 118a-d assigned indexes in the figure. The slice divisions of FIGS. 35-37 are chosen solely for convenience of explanation. Tile boundaries could be signaled in the data stream. Image 110 can form a single tile, as illustrated in FIG. 37.

[0256] The encoder 10 and decoder 12 can be configured to follow tile boundaries, in that spatial prediction is not applied across tile boundaries. Context adaptation, i.e., probability adaptation of various entropy (arithmetic) contexts, can continue throughout all slices. However, as shown in Figure 36 for slices 118a and 118b, whenever a slice—along the coding order 116—crosses a tile boundary (if it is internal to the slice), the slice is then subdivided into subsections (substreams or tiles), with each subsection containing a pointer (cpentry_point_offset) indicating the beginning of that subsection. In the decoder loop, filters can cross tile boundaries. Such filters can include one or more of a deblocking filter, a Sample Adaptive Offset (SAO) filter, and an Adaptive Loop Filter (ALF). If activated, the latter can be applied across tile / slice boundaries.

[0257] Each optional second and subsequent subsections can have their beginning located byte-aligned within a slice with a pointer indicating the offset from the beginning of one subsection to the beginning of the next subsection. Subsections are located within slices in scan order 116. Figure 38 shows an example embodiment with slice 180c of Figure 37 being subdivided into subsections 119i as an example.

[0258] With respect to the figures, it should be noted that slices forming subparts of a tile do not necessarily have to be at the edge with a row of tiles 112a. See, for example, slice 118a in Figures 37 and 38.

[0259] The following diagram shows a typical portion of a data stream for an access unit associated with image 110 of FIG. 38 above. Here, each payload packet 122a-d—previously designated by reference numeral 32—applies, by way of example, to only one slice 118a. Two timing control packets 124a, b—previously designated by reference numeral 36—are shown as distributed across access unit 120 for illustrative purposes: 124a precedes packet 122a in packet order 126 (corresponding to the decoding / encoding time axis), and 124b precedes packet 122c. Thus, the access unit 120 is divided into two decoding units 128a,b - previously indicated by reference numeral 38 - the first of which consists of packets 122a,b (any filter data packets (following the first and second packets 122a,b, respectively) and any access unit following the SEI packet (preceding the first packet 122a)), and the second of which consists of packets 118c,d (any filter data packets (following packets 122c,d, respectively)).

[0260] As described above, each packet in a sequence of packets can be assigned to exactly one packet type from multiple packet types (nal_unit_type). Payload packets and timing control packets (and any filter data and SEI packets) may, for example, be of different packet types. The instantiation of packets of a particular packet type in a sequence of packets may be subject to certain limitations. These limitations may define the order among packet types (see FIG. 17) to be followed by packets within each access unit so that access unit boundaries 130a, b are detectable, and packets of any removable packet type remain in the same position in the sequence of packets even if they are removed from the video data stream. For example, payload packets are of a non-removable packet type. However, timing control packets, filter data packets, and SEI packets may, as described above, be of removable packet types, i.e., they may be non-VCL NAL units.

[0261] In the above example, the timing control packet is clearly illustrated above by the syntax slice_prefix_rbsp0.

[0262] Using such a distribution of timing control packets, the encoder is enabled to adjust buffer scheduling at the decoder side during the process of encoding each image of the video content. For example, the encoder is enabled to optimize buffer scheduling to minimize end-to-end delay. In this regard, the encoder is enabled to take into account the individual distribution of coding complexities across the image area of ​​the video content for each image of the video content. In particular, the encoder is enabled to distribute the timing control packets 122, 122a-d, 122a-d for each packet (i.e., it is output as soon as the current packet is finished). 1-3Using timing control packets, the encoder can adjust buffer scheduling from time to time at the decoder side in case some subportions of the current picture have already been coded into their respective payload packets, but the remaining subportions have not yet been coded.

[0263] Thus, an encoder for encoding video content of a video data stream of units of sub-portions (see coding tree blocks, tiles or slices) of pictures of the video content by encoding each sub-portion into one or more payload packets (see VCL NAL unit) of the sequence of packets (NAL units) of the video data stream, such that the sequence of packets is divided into a sequence of access units, each access unit collecting payload packets associated with a respective picture of the video content, is configured to distribute a sequence of timing control packets (slice prefixes) among the packets, whereby the timing control packets subdivide the access units into decoding units, whereby at least some access units are subdivided into a plurality of decoding units, each timing control packet signaling a decoding buffer lookup time for a decoding unit whose payload packets follow its respective timing control packet in the sequence of packets.

[0264] Any decoder receiving the video data stream just outlined is free to use the schedule information contained in the timing control packets. However, while a decoder is free to use the information, a decoder conforming to the codec level must be able to decode the data after the indicated timing. If used, the decoder will supply and empty its decoder buffer on a decoding unit-by-decode unit basis. The "decoder buffer" may include a decoded picture buffer and / or a coded picture buffer, as described above.

[0265] Thus, a decoder for decoding a video data stream having video content encoded in units of sub-portions (see coding tree blocks, tiles or slices) of images of the video content, such that each access unit collects payload packets associated with a respective image of the video content, each sub-portion being encoded into one or more payload packets (see VCL NAL unit) of a sequence of packets of the video data stream, the sequence of packets being divided into a sequence of access units, is configured to look for timing control packets interspersed among the sequence of packets that subdivided the access units into decoding units with the timing control packets, whereby at least some access units are subdivided into a plurality of decoding units, deriving from each timing control packet a decoder buffer lookup time for the decoding unit whose payload packets follow a respective timing control packet of the sequence of packets, and retrieving the decoding unit from the scheduled decoder buffer at a time specified by the decoder buffer lookup time for the decoding unit.

[0266] Looking for a timing control packet can involve the decoder examining the NAL unit header and syntax elements contained therein, i.e., nal_unit_type. If the value of the latter flag is equal to some value, i.e., 124 in the above example, then the currently examined packet is a timing control packet. That is, the timing control packet contains or conveys the information described above with respect to the pseudocode subpic_buffering as well as subpic_timing. That is, the timing control packet can convey or specify the initial CPB removal delay for the decoder, or specify how many clock ticks elapse after removal from the CPB of the respective decoder unit.

[0267] To allow repeated transmission of timing control packets without unintentionally further dividing the access unit into decoding units, a flag within the timing control packet explicitly signals whether the current timing control packet is involved in subdivision into coding units within the access unit (compare decoding_unit_start_flag=1, indicating the start of a decoding unit, and decoding_unit_start_flag=0, signaling the opposite situation).

[0268] The use of distributed decoding units associated with tile identification information differs from the use of distributed decoding units associated with timing control packets in that tile identification packets are distributed throughout the data stream. The timing control packets described above can additionally be distributed throughout the data stream, or the decoder buffer lookup time can be commonly conveyed within the same packet along with the tile identification information described below. Therefore, the details brought forward in the above section can be used to clarify issues in the following description.

[0269] Further aspects of the present specification that can be derived from the above embodiments disclose a video data stream having video content encoded therein, using prediction and entropy coding on a slice-by-slice basis into which images of the video content are spatially subdivided, using a coding order between the slices to limit the prediction and / or entropy coding of the predictive coding to within tiles into which the images of the video content are spatially subdivided, the sequence of slices in the coding order being packetized into payload packets from a sequence of packets (NAL units) in the video data stream in the coding order, the sequence of packets being divided into a sequence of access units, such that each access unit collects payload packets packetized in a slice for a respective image of the video content, the sequence of packets having tile identification packets identifying and distributing to one or more tiles immediately covered by the slice (potentially only one) packetized in one or more payload packets after each tile identification packet in the sequence of packets.

[0270] See, for example, the immediately preceding figure showing a data stream. Packets 124a and 124b represent the current tile identification packet. By explicit signaling (compare single_slice_flag=1), or per convention, the tile identification packet can only identify the tile covered by the slice packetized in the immediately following payload packet 122a. Alternatively, by explicit signaling, or per convention, the tile identification packet 124a can identify the tile covered by the slice packetized in one or more payload packets after the respective tile identification packet 124a in the sequence of packets up to the earlier one at the end 130b of the current access unit 120 and beginning the next decoding unit 128b, respectively. See, for example, FIG. 35: each slice 118a-d 1-3 Separately, each packet 122a-d1-3 If the packetization is to {122a 1-3}, {122b 1-3} and {122c 1-3 , 122d 1-3}, and the packets of the third decoding unit {122c 1-3 , 122d 1-3} into a slice {118c 1-3 , 118d 1-3} covers, for example, tiles 112c and 112d, and the corresponding slice prefixes are, for example, "c" and "d" when referring to the complete decoding unit, i.e., these tiles 112c and 112d.

[0271] Thus, the network entities described further below can use this explicit signaling or convention to correctly associate each tile identification packet with one or more payload packets immediately following the identification packet in the packet sequence. The way in which the identification can be signaled was exemplarily described above via the pseudocode subpic_tile_info. The associated payload packets were previously described as "prefix slices." Naturally, the embodiment can be modified. For example, the syntax element "tile_priority" can be omitted. Furthermore, the order among the syntax elements can be switched, and the coding principles of the descriptors and syntax elements with respect to possible bit lengths are merely illustrated.

[0272] A network entity receiving a video data stream (i.e., a video data stream having video content encoded therein, using prediction and entropy coding for images of the video content in units of slices into which the images of the video content are spatially subdivided, and using a coding order among the slices to limit the prediction of the predictive coding and / or entropy coding to within the tiles into which the images of the video content are spatially subdivided, wherein the sequence of slices in the coding order is packetized into payload packets in a sequence of packets (NAL units) in the video data stream in the coding order, the sequence of packets being divided into a sequence of access units, such that each access unit collects payload packets into which slices for a respective image of the video content are packetized, and the sequence of packets has tile identification packets distributed therein) can be configured to identify tiles covered by the slice packetization in one or more payload packets after each tile identification packet in the sequence of packets, based on the tile identification packets. The network entity can use the identification result to determine transmission operations. For example, the network entity can treat different tiles with different priorities for playback. For example, in case of packet loss, those payload packets for higher priority tiles are preferably retransmitted over the payload packets for lower priority tiles. That is, the network entity may first request retransmission of lost payload packets for higher priority tiles. If there is only enough time remaining (depending on the transmission rate), the network entity moves on to requesting retransmission of lost payload packets for lower priority tiles. However, the network entity may also be a playback unit that is able to allocate tiles or payload packets for a particular tile to different screens.

[0273] With respect to aspects that use distributed region of interest information, it should be noted that the following ROI packets can coexist either through the timing control packets and / or tile identification packets described above, by combining their information content within a common packet as described above with respect to the slice prefix, or in the form of separate packets.

[0274] In an embodiment using distributed region of interest information as described above, in other words, a video data stream has video content therein coded using prediction and entropy coding, wherein images of the video content are spatially subdivided using a coding order between slices, and the prediction and / or entropy coding of the predictive coding is restricted to the interior of the tiles into which the images of the video content are divided, wherein the sequence of slices in the coding order are packetized into payload packets of a sequence of packets (NAL units) of the video data stream in the coding order, wherein the sequence of packets is divided into a sequence of access units such that each access unit collects therein payload packets of a slice relating to a respective image of the video content, and wherein the sequence of packets each has ROI packets identifying and distributed among tiles of the image belonging to a ROI of the image.

[0275] With respect to ROI packets, similar comments are valid as those given previously with respect to tile identification packets: an ROI packet can identify image tiles that belong to the image ROI only among those tiles covered by slices contained in one or more payload packets to which it is associated because each ROI packet immediately precedes one or more payload packets as described above with respect to "prefix slices".

[0276] The ROI packet may make it possible to identify multiple ROIs per preceding slice by identifying the associated tile for each of these ROIs (cpnum_rois_minus1). Then, for each ROI, a priority may be transmitted that makes it possible to rank the ROI in terms of priority (cproi_priority[i]). To enable "tracking" of ROIs over time during a video image sequence, each ROI may be indexed by an ROI index so that the ROIs indicated in the ROI packet are related to each other across / across image boundaries, i.e., over time (cproi_id[i]).

[0277] A network entity receiving a video data stream (i.e., a video data stream having video content therein that is coded using prediction and entropy coding, wherein the probability adaptation of the entropy coding continues across slices while restricting the predictions of the predictive coding to the interiors of tiles into which the images of the video content are divided, in units of slices into which the images of the video content are spatially subdivided, using a coding order between slices, wherein the sequence of slices in coding order is packetized into payload packets of a sequence of packets (NAL units) of the video data stream in coding order, and the sequence of packets is divided into a sequence of access units such that each access unit collects payload packets therein for a slice relating to a respective image of the video content) is configured to identify packets that packetize slices that cover tiles that belong to a region of interest of the image, based on the tile identification packets.

[0278] A network entity can utilize the information conveyed by the ROI packet in a similar manner as described above with respect to the Tile Identification Packet in the previous section. Regarding the current section as well as the previous section, it should be noted that any network entity, such as a MANE or decoder, can ascertain which tiles are covered by the slice of the currently examined payload packet by examining the slice order of the slices of the image and the progress of the portion of the current image that these slices cover, with respect to the tile's position in the image either explicitly signaled in the data stream as described above or known to the encoder or decoder by convention. Alternatively, each slice (except the first one in the image in scan order) can be provided with an indication / index (slice_address, measured in units of coding tree blocks) of the first coding block (e.g., CTB) to which it refers (with the same code) so that the decoder can locate each slice (its reconstruction) in the image from this first coding block in the direction of the slice order. Therefore, if the index conveyed by a later Tile Identification Packet differs from the previous one by a difference of 1 or more, it may be sufficient if the above-mentioned Tile Information Packet simply consists of the index of the first tile (first_tile_id_in_prefixed_slices) that is superimposed by the slices of one or more payload packets that follow the respective Tile Identification Packet, since it will be apparent to the network entity upon encountering the next Tile Identification Packet according to which the payload packets between those two Tile Identification Packets cover the tile with the tile index between them.As mentioned above, if the tile subdivision and coding block subdivision are based on row / column-wise subdivision with raster scan order, for example, the tile index increases in this raster scan order, since the slices follow each other in slice order along this raster scan order within a coding block, and it is also true that the slices are consecutive between coding blocks in slice order along this raster scan order.

[0279] The packetized and distributed slice header signaling aspects derived from the above-described embodiments can be combined with any one of the above-described aspects derived from the described embodiments or any combination thereof. The slice prefix explicitly described above unifies all these aspects, e.g., by version 2. An effect of the present aspects is the possibility to make slice header data more readily available to network entities so that they can be conveyed in self-contained packets outside the prefixed slice / payload packets, and repeated transmission of slice header data is enabled.

[0280] Therefore, a further aspect of the present specification is an aspect of packetized and distributed slice header signaling, in other words, it can be considered to expose a video data stream having video content encoded in units of sub-portions (see coding tree blocks or slices) of images of the video content, each sub-portion being encoded into one or more payload packets (see VCL NAL unit) of a sequence of packets (NAL units) of the video data stream, the sequence of packets being divided into a sequence of access units such that each access unit collects payload packets for a respective image of the video content, and the sequence of packets having distributed slice header packets (slice prefixes) conveying missing slice header data between one or more payload packets following each slice header packet of the sequence of packets.

[0281] A network entity receiving a video data stream (i.e., a video data stream having video content encoded in units of sub-portions (see coding tree blocks or slices) of pictures of the video content, each sub-portion encoded into one or more payload packets (see VCL NAL unit) of a sequence of packets (NAL units) of the video data stream, the sequence of packets divided into a sequence of access units such that each access unit collects payload packets for a respective picture of the video content, and the sequence of packets having scattered slice header packets therein) is configured to read the slice headers by reading along the payload data for the slice from packets derived from the slice header data in the slice header packets, and for one or more payload packets following each slice header packet in the sequence of packets, but skipping the slice headers for one or more payload packets that apply a slice header derived from the slice header packet followed by one or more payload packets.

[0282] As was true with respect to the above-described aspects, a slice header packet, such as a slice header packet, may also function to indicate to any network entity, such as a MANE or decoder, the beginning of a decoding unit or the beginning of execution of one or more payload packets preceded by the respective slice. Thus, a network entity according to the present aspect, reading the slice header, can identify payload packets that should be skipped based on the above-described syntax elements of this packet, i.e., single_slice_flag, in conjunction with the decoding_unit_start_flag, in which a later flag, as described above, enables retransmission of a copy of a particular slice header packet within a decoding unit. This is useful, for example, because the slice headers of slices within a single decoding unit can vary along the slice sequence. Thus, the first slice header packet of a decoding unit can have the decoding_unit_start_flag set (equal to 1), while intervening slice header packets can have this flag not set, preventing any network entity from incorrectly reading the occurrence of this slice header packet as starting a new decoding unit.

[0283] Although some aspects have been described in the context of an apparatus, it will be apparent that these aspects also represent a description of a corresponding method, with blocks or apparatus corresponding to method steps or features of method steps. Similarly, aspects described in the context of a method step represent a description of a corresponding block or item or feature of a corresponding apparatus. Some or all method steps can be performed by (using) a hardware apparatus, such as a microprocessor, a programmable computer, or electronic circuitry. In some embodiments, one or more of the most important method steps can be performed by such an apparatus.

[0284] The inventive video data stream can be stored on a digital storage medium or can be transmitted over a transmission medium such as a wireless transmission medium or a wired transmission medium such as the Internet.

[0285] Depending on specific implementation requirements, embodiments of the present invention can be implemented in hardware or in software. Implementation can be performed using a digital storage medium having electronically readable control signals stored thereon, such as a floppy disk, DVD, Blu-ray, CD, ROM, PROM, EPROM, EEPROM, or flash memory, which cooperates (or can cooperate) with a programmable computer system so that the respective methods are performed. Thus, the digital storage medium may be computer-readable.

[0286] Some embodiments of the present invention include a data carrier having an electronically readable control signal that can cooperate with a programmable computer system to perform one of the methods described herein.

[0287] Typically, embodiments of the present invention can be realized as a computer program product having program code operable to perform one of the methods when the computer program product is run on a computer, the program code being for example stored on a machine-readable carrier.

[0288] Other embodiments comprise the computer program for performing one of the methods described herein, stored on a machine readable carrier.

[0289] In other words, an embodiment of the inventive method is, therefore, a computer program having a program code for performing one of the methods described herein, when the computer program runs on a computer.

[0290] A further embodiment of the inventive method is, therefore, a data carrier (or digital storage medium or computer-readable medium) comprising, recorded on it, a computer program for performing one of the methods described herein. The data carrier, digital storage medium or storage medium is typically tangible and / or non-transitory.

[0291] A further embodiment of the inventive method is, therefore, a data stream or a sequence of signals representing the computer program for performing one of the methods described herein, The data stream or the sequence of signals can be adapted to be transmitted via a data communication connection, for example the Internet.

[0292] A further embodiment comprises a processing means, for example a computer, or a programmable logic device, configured to or adapted to perform one of the methods herein.

[0293] A further embodiment comprises a computer having installed thereon the computer program for performing one of the methods described herein.

[0294] Further embodiments according to the present invention include an apparatus or system configured to transmit (e.g., electronically or optically) a computer program for performing one of the methods described herein to a receiver. The receiver may be, for example, a computer, a mobile device, a memory device, etc. The apparatus or system may include, for example, a file server for transmitting the computer program to the receiver.

[0295] In some embodiments, a programmable logic device (e.g., a field programmable gate array) can be used to perform some or all of the functionality of the methods described herein. In some embodiments, a field programmable gate array can cooperate with a microprocessor to perform one of the methods described herein. In general, the methods are preferably performed by any hardware device.

[0296] The above-described embodiments are merely illustrative for the principles of the present invention. It is understood that modifications and variations of the arrangements and details described herein will be apparent to others skilled in the art. It is therefore intended to be limited only by the scope of the impending patent claims, and not by the specific details set forth herein as descriptions and illustrations of the embodiments.

[0297] References [1] Thomas Wiegand, Gary J. Sullivan, Gisle Bjontegaard, Ajay Luthra, "Overview of the H.264 / AVC Video Coding Standard", IEEE Trans. Circuits Syst. Video Technol., vol. 13, N7, July 2003. [2] JCT-VC, "High-Efficiency Video Coding (HEVC) text specification Working Draft 7", JCTVC-I1003, May 2012. [3] ISO / IEC 13818-1: MPEG-2 Systems specification. [4] IETF RFC 3550 - Real-time Transport Protocol. [5] Y.-K. Wang et al. ,"RTP Payload Format for H.264 Video", IETF RFC6184, http: / / tools.ietf.org / html / [6] S. Wenger et al., "RTP Payload Format for Scalable Video Coding", IETF RFC6190, http: / / tools.ietf.org / html / rfc6190 [6] T. Schierl et al., "RTP Payload Format for High Efficiency Video Coding", IETF internet draft, http: / / datatracker.ietf.org / doc / draft-schierl-payload-rtp-h265 /

Claims

1. A video datastream having video content (16) encoded in units of sub-portions (24) of images (18) of the video content (16), each sub-portion (24) being encoded into one or more payload packets (32) of a sequence of packets (34) of a video datastream (22), the sequence of packets (34) being divided into a sequence of access units (30), whereby each access unit (30) collects payload packets (32) for a respective image (18) of the video content (16), a sequence of packets (34) having timing control packets (36) interspersed therein, whereby the timing control packets (36) subdivide the access units (30) into decoding units (38), whereby at least some of the access units (30) are subdivided into a plurality of decoding units (38), each timing control packet (38) signaling a decoder buffer lookup time for a decoding unit (38), whose payload packets (32) follow their respective timing control packets (38) in the sequence of packets (34).

2. 2. The video datastream of claim 1, wherein the sub-portions (24) are slices and each payload packet (32) contains one or more slices.

3. 3. The video data stream of claim 2, wherein the slices include independently decodable slices and dependent slices that enable decoding using entropy and predictive decoding across slice boundaries for WPP processing.

4. 4. The video data stream of claim 1, wherein each packet of the sequence of packets is assigned to exactly one packet type from a plurality of packet types, with the payload packets and timing control packets being of different packet types, and the occurrence of packets of the plurality of packet types in the sequence of packets is dependent on some constraint that defines an order between the packet types to be followed by the packets within each access unit, whereby an access unit boundary can be detected using the constraint by detecting instances, and even if a packet of any removable packet type is removed from the video data stream, the constraint is reversed and it remains in the same position within the sequence of packets, and the payload packets are of a non-removable packet type and the timing control packets are of a removable packet type.

5. 5. A video data stream as claimed in any preceding claim, wherein each packet includes a packet type indicating a syntax element portion.

6. 6. The video data stream of claim 5, wherein the packet type indicating syntax element portion includes a packet type field in the packet header of each packet, a different content between payload packets and timing control packets, and an SEI packet type field for timing control packets, distinguishing between timing control packets on the one hand and different types of SEI packets on the other hand.

7. An encoder for encoding video content (16) into a video data stream (22) in units of sub-portions (24) of images (18) of the video content (16), encoding each sub-portion (24) into one or more payload packets (32) of a sequence of packets (34) of the video data stream (22), whereby the sequence of packets (34) is divided into a sequence of access units (30), each access unit (30) collecting and encoding payload packets (32) for a respective image (18) of the video content (16). The encoder distributes timing control packets (36) among the sequence of packets (34), whereby the timing control packets (36) subdivide the access units (30) into decoding units (38), whereby at least some of the access units (30) are subdivided into multiple decoding units (38), and each timing control packet (36) signals a decoder buffer search time for a decoding unit (38), whose payload packets (32) follow their respective timing control packets (36) in the sequence of packets (34).

8. The encoder, while encoding a current image of the video content, encodes a current sub-portion (24) of the current image (18) into a current payload packet (32) of a current decoding unit (38); At a first instant in time, transmitting in the data stream a current decoding unit preceded by a current timing control packet (36) by setting a decoder buffer search time signaled by the current timing control packet (36); 8. An encoder as claimed in claim 7, configured to encode a further sub-portion of the current image at a second time instant later than the first time instant.

9. A method for encoding video content (16) into a video data stream (22) in units of sub-portions (24) of images (18) of the video content (16), comprising encoding each respective sub-portion (24) into one or more payload packets (32) of a sequence of packets (34) of the video data stream (22), whereby the sequence of packets (34) is divided into a sequence of access units (30), each access unit (30) collecting payload packets (32) associated with a respective image (18) of the video content (16), The method includes distributing timing control packets (36) among a sequence of packets (34), whereby the timing control packets (36) subdivide access units (30) into decoding units (38), whereby at least some access units (30) are subdivided into multiple decoding units (38), each timing control packet (36) signaling a decoder buffer search time for a decoding unit (38) whose payload packets (32) follow its respective timing control packet (36) in the sequence of packets (34).

10. 1. A decoder for decoding a video data stream (22) having video content (16) encoded therein in units of sub-portions (24) of images (18) of the video content (16), the decoder encoding each sub-portion into one or more payload packets (32) of a sequence of packets (34) of the video data stream (22), the sequence of packets (34) being divided into a sequence of access units (30), whereby each access unit (30) collects payload packets (32) associated with a respective image (18) of the video content (16), the decoder including a buffer for buffering the video data stream or a reconstruction of the video content obtained therefrom by decoding the video data stream, locating timing control packets (36) distributed among the sequence of packets, and subdividing the access units (30) into decoding units (38) at the timing control packets (36), whereby at least some of the access units are subdivided into a plurality of decoding units and configured to empty a buffer based on the decoding units.

11. 11. The decoder of claim 10, wherein the decoder is configured to, when searching for timing control packets (36), examine a packet type indicating syntax element portion in each packet, and recognize the respective packet as a timing control packet (36) if the value of the packet type indicating syntax element portion is equal to a predetermined value.

12. 1. A method for decoding a video data stream (22) having video content (16) encoded therein in units of sub-portions (24) of images (18) of the video content (16), the method comprising: encoding each sub-portion into one or more payload packets (32) of a sequence of packets (34) of the video data stream (22), the sequence of packets (34) being divided into a sequence of access units (30), whereby each access unit (30) collects payload packets (32) associated with a respective image (18) of the video content (16); the method using a buffer for buffering the video data stream or playback of the video content obtained therefrom by decoding the video data stream, locating timing control packets (36) interspersed among the sequence of packets, and subdividing the access units (30) into decoding units (38) at the timing control packets (36), whereby at least some access units are subdivided into a plurality of decoding units and emptying the buffer in units of decoding units.

13. A network entity for transmitting a video data stream (22) having video content (16) encoded therein in units of sub-portions (24) of images (18) of the video content (16), each sub-portion being encoded into one or more payload packets (32) of a sequence (34) of packets of the video data stream (22), the sequence (34) of packets being divided into a sequence of access units (30), whereby each access device (30) collects payload packets (32) associated with a respective image (18) of the video content (16), and a decoder extracts the sequence of packets. a network entity configured to search for timing control packets (36) distributed among the access units, subdivide the access units into decoding units with the timing control packets (36), whereby at least some of the access units (30) are subdivided into a plurality of decoding units (38), derive a decoder buffer lookup time for the decoding unit (38) from each timing control packet (36), and perform transmission of a video data stream whose payload packets (32) follow each timing control packet (36) in a sequence of packets (34) and are subject to the decoder buffer lookup time for the decoding unit (38).

14. A method for transmitting a video data stream (22) having video content (16) encoded therein in units of sub-portions (24) of images (18) of the video content (16), each sub-portion being encoded into one or more payload packets (32) of a sequence (34) of packets of the video data stream (22), the sequence (34) of packets being divided into a sequence of access units (30), whereby each access unit (30) collects payload packets (32) associated with a respective image (18) of the video content (16), the method comprising: and performing transmission of a video data stream, the video data stream being configured to look for timing control packets (36) distributed among the access units, subdividing the access units into decoding units with the timing control packets (36), whereby at least some of the access units (30) are subdivided into a plurality of decoding units (38), deriving a decoder buffer lookup time for the decoding unit (38) from each timing control packet (36), and the payload packets (32) of the video data stream following each timing control packet (36) in the sequence of packets (34) and being subject to the decoder buffer lookup time for the decoding unit (38).

15. A video data stream having video content (16) encoded into slices (24) into which images (18) of the video content (16) are spatially subdivided, using a coding order within the slices (24), wherein the prediction and / or entropy coding of the predictive coding is performed within tiles into which the images of the video content are spatially subdivided. or limiting the prediction of entropy coding, wherein a sequence of slices (24) is packetized in coding order into payload packets (32) of a sequence of packets of a video data stream, and the sequence of packets (34) is divided into a sequence of access units (30), whereby each access unit collects payload packets (32) into which slices (24) associated with a respective image (18) of the video content (16) are packetized, and wherein the sequence of packets (34) has tile identification packets (72) distributed therein among the payload packets of one access unit, identifying one or more tiles (70) covered by some of the slices (24) packetized in one or more payload packets (32) immediately following each tile identification packet (72) of the sequence of packets (34).

16. 16. The video data stream of claim 15, wherein a tile identification packet (72) identifies one or more tiles (70) that are covered by several slices (24) packetized in the payload packet that exactly follows.

17. 16. The video data stream of claim 15, wherein the tile identification packets (72) identify one or more tiles (70) covered by several slices (24) packetized in one or more payload packets (32) immediately following each tile identification packet (72) in the sequence of packets (34) until before the end of the current access unit (30), and each next tile identification packet (72) is in the sequence of packets (34).

18. A network entity configured to receive a video data stream as claimed in any one of claims 15 to 16 and to identify based on tile identification packets (72) tiles (70) covered by a slice (24) packetized in one or more payload packets (72) immediately following each tile identification packet (72) in the sequence of packets.

19. 20. The network entity of claim 18, wherein the network entity is configured to use a result of the identification to determine a transmission task related to the video data stream.

20. 20. A network entity according to claim 18 or claim 19, wherein the transmission task comprises a retransmission request relating to a bad packet.

21. 20. The network entity of claim 18 or 19, wherein the network entity handles different tiles (70) with different priorities by assigning a high priority to a tile identification packet (72) and a payload packet immediately following the respective tile identification packet (72) having a packetized slice covering the tile (70) identified by the respective tile identification packet, and therefore a low priority to a tile identification packet (72) and a payload packet immediately following the respective tile identification packet (72) having a packetized slice covering the tile (70) identified by the respective tile identification packet.

22. 22. The network entity of claim 21, wherein the network entity is configured to first request retransmission of payload packets having a higher assigned priority before any retransmission requests for payload packets having a lower assigned priority.

23. 17. A method configured to receive a video data stream as claimed in any one of claims 15 to 16 and identify based on tile identification packets (72) tiles (70) covered by a slice (24) packetized in one or more payload packets (72) immediately following each tile identification packet (72) in the sequence of packets.

24. A video data stream having video content (16) encoded therein in units of sub-portions (24) of images (18) of the video content (16), Each sub-portion (24) is encoded into one or more payload packets (32) of a sequence of packets (34) of the video data stream (22), and the sequence of packets (34) is divided into a sequence of access units (30), whereby each access unit (30) collects payload packets (32) associated with a respective image (18) of the video content (16), and at least some of the access units have a sequence of packets (34) having ROI packets (64) distributed therein, whereby timing control packets (64) are decoded. a coding unit (38) for subdividing the access units (64) so ​​that at least some of the access units (30) have ROI packets (64) interspersed among payload packets associated with the images of the respective access units, each ROI packet being associated with one or more subsequent payload packets in the sequence of packets (34), and determining whether a sub-portion (24) following each ROI packet and encoded in the one or more payload packets with which each ROI packet is associated covers a region of interest in the video content.

25. 25. The video data stream of claim 24, wherein the sub-portions are slices, and the video content is encoded into the video data stream using prediction and entropy coding constrained to prediction and / or entropy coding within tiles into which images of the video content are divided, and wherein each ROI packet further identifies the tile covering the region of interest, the sub-portion (24) coded into some of the one or more payload packets with which the respective ROI packet is associated.

26. 26. A video datastream according to claim 24 or claim 25, wherein each ROI packet relates only to the payload packet that immediately follows it.

27. 26. A video data stream as described in claim 24 or claim 25, wherein each ROI packet is associated with all payload packets that immediately follow it in the sequence of packets until the earlier of the end of the access unit in which the respective ROI packet is located or the start of the next ROI packet.

28. 28. A network entity configured to receive a video data stream according to any one of claims 24 to 27, and to ascertain an ROI of the video content based on the ROI packets.

29. 28. The network entity of claim 27, wherein the network entity is configured to use a result of the verification to determine a transmission task related to the video data stream.

30. 30. A network entity according to claim 28 or claim 29, wherein the transmission task comprises a retransmission request for a bad packet.

31. 30. The network entity of claim 28 or 29, wherein the network entity is configured to treat a region of interest (70) with increased priority by assigning a higher priority to a ROI packet (72) and one or more payload packets following the respective ROI packet (72) to which the respective ROI packet is associated and which signal region of interest overlap by a subportion (24) encoded in some of the one or more payload packets to which the respective ROI packet is associated than to the ROI packet and one or more payload packets following the respective ROI packet (72) to which the respective ROI packet is associated and in which the ROI packet signals no overlap in the ROI packet signal.

32. 32. The network entity of claim 31, wherein the network entity is configured to first request retransmission of payload packets that have a higher priority assigned thereto before any retransmission requests for payload packets that have a lower priority assigned thereto.

33. 27. A method for receiving a video data stream according to any one of claims 23 to 26, and ascertaining an ROI of the video content based on the ROI packets.

34. A computer program having a program code for performing the method according to claim 9, 12, 14, 23 or 33 when the computer program runs on a computer.

Citation Information

Patent Citations

  • IEC13818-1