Method, decoder, and computer-readable medium for transmitting a video data stream

By interspersing timing control packets into the video data stream, the problem of encoder transmitting video data streams under low latency is solved, and more efficient bit rate utilization and end-to-end delay reduction is achieved, supporting efficient transmission of ROI and tile identification information.

CN115442626BActive Publication Date: 2025-08-01DOLBY VIDEO COMPRESSION LLC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210894128.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2012-06-29
Filing Date
2013-07-01
Publication Date
2025-08-01
Estimated Expiration
2033-07-01

AI Technical Summary

Technical Problem

Existing video encoding technologies are difficult to effectively transmit video data streams under low latency conditions, especially before the encoder has not completed encoding of the current frame, resulting in an increase in end-to-end delay and unoptimized bit rate distribution.

Method used

By interspersing timing control packets in the video data stream, the encoder allows the decoder buffer acquisition time in real time during encoding the current frame and transmits the bit rate between the encoding units, and uses the interspersed packet to convey the decoder buffer acquisition time information to reduce end-to-end delay.

Benefits of technology

It realizes more efficient use of available bit rates during the encoding process, reduces end-to-end delays, and allows network entities to easily access ROI information and tile identification information, without in-depth verification of the internal packet, and improves the transmission efficiency and delay performance of the data stream.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115442626B_ABST
    Figure CN115442626B_ABST
Patent Text Reader

Abstract

The present invention relates to a video data stream, an encoder, a method for encoding video content, and a decoder. The video data stream has video content encoded therein in units of sub-parts of images of the video content, each sub-part being encoded into one or more payload packets of a packet sequence of the video data stream, the packet sequence being divided into a sequence of access units, so that each access unit collects the payload packets related to the respective images of the video content, wherein the packet sequence has timing control packets interspersed therein, so that the timing control packets subdivide the access units into decoding units, so that at least some of the access units are subdivided into two or more decoding units, wherein each timing control packet signals the decoder buffer fetch time of the decoding unit, and the payload packets of the decoding unit follow each timing control packet in the packet sequence.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of divisional application 201910661351.X of the invention application with the application number 201380034944.4 and the invention title of "Video Data Stream Concept Technology", which entered the national stage on December 29, 2014 for the international application with the international filing date of July 1, 2013 and the international application number of PCT / EP2013 / 063853. The entire content is incorporated herein by reference. Technical Field

[0002] This application relates to video data stream concept technologies. Specifically, these concept technologies are advantageous in low-latency applications. Background Art

[0003] HEVC [2] allows different means for high-level syntax signaling to the application layer. These means are NAL unit headers, parameter sets, and supplementary enhancement information (SEI) messages. SEI messages are not used during the decoding process. Other means of high-level syntax signaling originate from individual transport protocol specifications, such as the MPEG2 transport protocol [3] or the Real-Time Transport Protocol [4], and their payload-specific specifications, such as recommendations for H.264 / AVC [5], Scalable Video Coding (SVC) [6], or HEVC [7]. These transport protocols can introduce high-level signaling, and the structures and mechanisms they use are similar to those of high-level signaling in individual application layer codec specifications (such as HEVC [2]). An example of such signaling is the Payload Content Scalability Information (PACSI) NAL unit described in [6], which provides supplementary information to the transport layer.

[0004] HEVC [2] allows different means for high-level syntax signaling to the application layer. These means are NAL unit headers, parameter sets, and supplementary enhancement information (SEI) messages. SEI messages are not used during the decoding process. Other means of high-level syntax signaling originate from individual transport protocol specifications, such as the MPEG2 transport protocol [3] or the Real-Time Transport Protocol [4], and their payload-specific specifications, such as recommendations for H.264 / AVC [5], Scalable Video Coding (SVC) [6], or HEVC [7]. These transport protocols can introduce high-level signaling, and the structures and mechanisms they use are similar to those of high-level signaling in individual application layer codec specifications (such as HEVC [2]). An example of such signaling is the Payload Content Scalability Information (PACSI) NAL unit described in [6], which provides supplementary information to the transport layer.

[0005] In terms of parameter sets, HEVC includes Video Parameter Sets (VPSs), which compile the most important stream information to be used by the application layer at a single and central location. In earlier methods, this information needed to be collected from multiple parameter sets and NAL unit headers.

[0006] Prior to this application, the state of the standard regarding the operation of the coded picture buffer (CPB) of the Hypothetical Reference Decoder (HRD), and all relevant syntax provided in the Sequence Parameter Set (SPS) / Video Usability Information (VUI), Picture Timing SEI, Buffering Period SEI, and the definition of decoding units (which describe sub-pictures and the syntax of dependent slices as presented in slice headers and Picture Parameter Sets (PPSs)) are as follows.

[0007] To allow for low-latency CPB operation at the sub-picture level, sub-picture CPB operation has been proposed and incorporated into the HEVC standard draft 7JCTVC-I1003 [2]. In particular here, in Section 3 of [2], the decoding unit has been defined as:

[0008] Decoding Unit: An access unit or a subset of an access unit. If SubPicCpbFlag is equal to 0, the decoding unit is the access unit. Otherwise, the decoding unit consists of one or more VCL NAL units and associated non-VCL NAL units in the access unit. For the first VCL NAL unit in the access unit, the associated non-VCL NAL units are the filler data NAL unit (if present) immediately following the first VCL NAL unit and all non-VCL NAL units in the access unit that are before the first VCL NAL unit. For a VCL NAL unit that is not the first VCL NAL unit in the access unit, the associated non-VCL NAL unit is the filler data NAL unit (if present) immediately following that VCL NAL unit.

[0009] In the standards defined up to that point, "decoding unit removal timing and decoding unit decoding" had been described. To signal sub-picture timing, the buffering period SEI message and the picture timing SEI message, as well as the HRD parameters in the VUI, have been extended to support decoding units, such as sub-picture units.

[0010] [2]'s buffering period SEI message syntax is shown in Figure 1 in.

[0011] When NalHrdBpPresentFlag or VclHrdBpPresentFlag is equal to 1, the buffering period SEI message can be associated with any access unit in the bitstream, and the buffering period SEI message will be associated with every RAP access unit and every access unit associated with a recovery point SEI message.

[0012] For some applications, the frequent occurrence of buffering period SEI messages may be desirable.

[0013] The buffering period is defined as the set of access units between two instances of the buffering period SEI message in decoding order.

[0014] The semantics are as follows:

[0015] seq_parameter_set_id specifies the sequence parameter set containing the sequence HRD attributes. The value of seq_parameter_set_id shall be equal to the value of seq_parameter_set_id in the picture parameter set referred to by the primary coded picture associated with the buffering period SEI message. The value of seq_parameter_set_id shall be in the range 0 to 31, inclusive of 0 and 31.

[0016] rap_cpb_params_present_flag equal to 1 specifies the presence of the initial_alt_cpb_removal_delay[SchedSelIdx] and initial_alt_cpb_removal_delay_offset[SchedSelIdx] syntax elements. When not present, the value of rap_cpb_params_present_flag is inferred to be equal to 0. When the associated picture is neither a CRA picture nor a BLA picture, the value of rap_cpb_params_present_flag shall be equal to 0.

[0017] initial_cpb_removal_delay[SchedSelIdx] and initial_alt_cpb_removal_delay[SchedSelIdx] specify the initial CPB removal delay for the SchedSelIdx-th CPB. These syntax elements have a length in bits given by initial_cpb_removal_delay_length_minus1 + 1 and are in units of a 90 kHz clock. The values of these syntax elements shall not be equal to 0 and shall not exceed 90000 * (CpbSize[SchedSelIdx] ÷ BitRate[SchedSelIdx]), which is the time equivalent of the CPB size in units of a 90 kHz clock.

[0018] initial_cpb_removal_delay_offset[SchedSelIdx] and initial_alt_cpb_removal_delay_offset[SchedSelIdx] are used for the SchedSelIdx-th CPB to specify the initial delivery time of the coded data unit to that CPB. These syntax elements have a length in bits given by initial_cpb_removal_delay_length_minus1 + 1, and are in units of the 90 kHz clock. These syntax elements are not used by the decoder and are only required by the delivery scheduler (HSS).

[0019] Over the entire coded video sequence, the sum of initial_cpb_removal_delay[SchedSelIdx] and initial_cpb_removal_delay_offset[SchedSelIdx] will be constant for each value of SchedSelIdx, and the sum of initial_alt_cpb_removal_delay[SchedSelIdx] and initial_alt_cpb_removal_delay_offset[SchedSelIdx] will be constant for each value of SchedSelIdx.

[0020] [2] The picture timing SEI message syntax is shown in Figure 2 .

[0021] The syntax of the picture timing SEI message depends on the contents of the sequence parameter set that is valid for the coded picture associated with the picture timing SEI message. However, unless the picture timing SEI message of an IDR or BLA access unit is preceded by a buffering period SEI message within the same access unit, the start of the relevant sequence parameter set (and, for an IDR or BLA picture that is not the first picture in the bitstream, the determination that the coded picture is an IDR picture or a BLA picture) does not occur before the decoding of the first coded slice NAL unit of the coded picture. Since the coded slice NAL units of the coded picture follow the picture timing SEI message in NAL unit order, it may be the case that the decoder needs to store the RBSP containing the picture timing SEI message before determining the parameters of the sequence parameters that will be valid for the coded picture, and then perform the parsing of the picture timing SEI message.

[0022] The presence of the picture timing SEI message in the bitstream is specified as follows.

[0023] - If CpbDpbDelaysPresentFlag equals 1, then there will be a picture timing SEI message in each access unit of the encoded video sequence.

[0024] - Otherwise (CpbDpbDelaysPresentFlag equals 0), there will be no picture timing SEI message in any access unit of the encoded video sequence.

[0025] The semantics are defined as follows:

[0026] cpb_removal_delay specifies how many clock ticks to wait after cpb removal of the access unit associated with the most recent buffering period SEI message in the previous access unit and before removal of the access unit data associated with the picture timing SEI message from the buffer. This value is also used to calculate the earliest possible time for the access unit data to reach the CPB for the HSS. This syntax element is a fixed - length code, and its length in bits is given by cpb_removal_delay_length_minus1 + 1. cpb_removal_delay is the remainder of a modulo - 2 (cpb_removal_delay_length_minus1+1) counter.

[0027] The value of cpb_removal_delay_length_minus1 that determines the length (in bits) of the syntax element cpb_removal_delay is the value of cpb_removal_delay_length_minus1 encoded in the sequence parameter set that is valid for the primary encoded picture associated with the picture timing SEI message, but cpb_removal_delay specifies the number of clock ticks related to the removal time of the previous access unit containing the buffering period SEI message.

[0028] dpb_output_delay is used to calculate the DPB output time of a picture. It specifies how many clock ticks to wait after the last decoded unit in the access unit is removed from the CPB and before the decoded picture is output from the DPB.

[0029] The picture is removed from the DPB at its output time, when it is still marked as "for short - term reference" or "for long - term reference".

[0030] For a decoded picture, only one dpb_output_delay is specified.

[0031] The length of the syntax element dpb_output_delay is given in bits by dpb_output_delay_length_minus1 + 1. When sps_max_dec_pic_buffering[max_temporal_layers_minus1] equals 0, dpb_output_delay shall be equal to 0.

[0032] The output time derived from dpb_output_delay for any picture output by a decoder that is self-consistent with the output time shall be before, in decoding order, the output times derived from dpb_output_delay for all pictures in any subsequent coded video sequence.

[0033] The picture output order determined by the value of this syntax element shall be the same as that determined by the value of PicOrderCntVal.

[0034] For pictures that are not output by the "bumping" process (since such pictures are, in decoding order, before IDR or BLA pictures for which no_output_of_prior_pics_flag equals 1 or is inferred to equal 1), the output time derived from dpb_output_delay shall increase as the value of PicOrderCntVal increases for all pictures within the same coded video sequence.

[0035] num_decoding_units_minus1 + 1 specifies the number of decoding units in the access unit associated with the picture timing SEI message. The value of num_decoding_units_minus1 shall be in the range of 0 to PicWidthInCtbs * PicHeightInCtbs - 1, inclusive (including 0 and PicWidthInCtbs * PicHeightInCtbs - 1).

[0036] num_nalus_in_du_minus1[i] + 1 specifies the number of NAL units in the i-th decoding unit of the access unit associated with the picture timing SEI message. The value of num_nalus_in_du_minus1[i] shall be in the range of 0 to PicWidthInCtbs * PicHeightInCtbs - 1, inclusive (including 0 and PicWidthInCtbs * PicHeightInCtbs - 1).

[0037] The first decoding unit of an access unit consists of the first num_nalus_in_du_minus[0]+1 consecutive NAL units in decoding order in the access unit. The i-th (where i>0) decoding unit of an access unit consists of num_nalus_in_du_minus[i]+1 consecutive NAL units in decoding order immediately following the last NAL unit in the previous decoding unit of the access unit. There will be at least one VCL NAL unit in each decoding unit. All non-VCL NAL units associated with a VCL NAL unit will be included in the same decoding unit.

[0038] du_cpb_removal_delay[i] specifies how many sub-picture clock ticks to wait after cpb removal for the i-th decoding unit in the access unit related to the picture timing SEI message, after cpb removal for the first decoding unit of the access unit related to the most recent buffering period SEI message in the previous access unit. This value is also used to calculate the earliest possible time for the decoding unit data to reach the CPB for the HSS. This syntax element is a fixed length code, and its length in bits is given by cpb_removal_delay_length_minus1+1. du_cpb_removal_delay[i] is the remainder of modulo 2 (cpb _removal_delay_length_minus1+1) counter.

[0039] The value of cpb_removal_delay_length_minus1 that determines the length (in bits) of the syntax element du_cpb_removal_delay[i] is the value of cpb_removal_delay_length_minus1 encoded in the sequence parameter set valid for the coded picture related to the picture timing SEI message, but du_cpb_removal_delay[i] specifies the number of sub-picture clock ticks related to the removal time of the first decoding unit in the previous access unit containing the buffering period SEI message.

[0040] [2] contains some information in the VUI syntax. The VUI parameter syntax of [2] is shown in Figure 3A and Figure 3B in. The HRD parameter syntax of [2] is shown in Figure 4 in. The semantic definition is as follows:

[0041] When sub_pic_cpb_params_present_flag equals 1, it specifies that the sub-picture level CPB removal delay parameter exists and the CPB can operate at the access unit level or the sub-picture level. When sub_pic_cpb_params_present_flag equals 0, it specifies that the sub-picture level CPB removal delay parameter does not exist and the CPB operates at the access unit level. When the sub_pic_cpb_params_present_flag does not exist, it is inferred to have a value of 0.

[0042] num_units_in_sub_tick is the number of time units of a clock operating at a frequency of time_scale Hz, which corresponds to an increment of the sub-picture clock tick counter (referred to as a sub-picture clock tick). num_units_in_sub_tick will be greater than 0. The sub-picture clock tick is the smallest time interval that can be represented in the encoded data when sub_pic_cpb_params_present_flag equals 1.

[0043] When tiles_fixed_structure_flag equals 1, it indicates that each valid picture parameter set in the encoded video sequence has the same values for the following syntax elements: num_tile_columns_minus1, num_tile_rows_minus1, uniform_spacing_flag, column_width[i], row_height[i], and loop_filter_across_tiles_enabled_flag (when present). When tiles_fixed_structure_flag equals 0, it indicates that the tile syntax elements in different picture parameter sets may or may not have the same values. When the tiles_fixed_structure_flag syntax element does not exist, it is inferred to be equal to 0.

[0044] Signaling that tiles_fixed_structure_flag equals 1 is an assurance to the decoder that each picture in the encoded video sequence has the same number of tiles distributed in the same way, which may be useful for workload distribution in multi-threaded decoding scenarios.

[0045] Use Figure 5 the filler data RBSP syntax shown in [2] to signal the filler data.

[0046] The hypothetical reference decoder for [2] used to check bitstream and decoder compliance is defined as follows:

[0047] Two types of bitstreams are subject to HRD compliance checking for this proposed | international standard. The first type of such bitstream (referred to as type I bitstream) is a NAL unit stream that contains only VCL NAL units and filler data NAL units for all access units in the bitstream. The second type of such bitstream (referred to as type II bitstream) contains, in addition to VCL NAL units and filler data NAL units for all access units in the bitstream, at least one of the following:

[0048] - additional non-VCL NAL units other than filler data NAL units,

[0049] - all leading_zero_8bits, zero_byte, start_code_prefix_one_3bytes, and trailing_zero_8bits syntax elements, each of which forms a bitstream from the NAL unit stream.

[0050] Figure 6 Shows the types of bitstream compliance points for the HRD check of [2].

[0051] Two types of HRD parameter sets (NAL HRD parameters and VCL HRD parameters) are used. These HRD parameter sets are signaled via video usability information, which is part of the sequence parameter set syntax structure.

[0052] All sequence parameter sets, picture parameter sets, and corresponding buffering periods and picture timing SEI messages referred to in the VCL NAL units are conveyed to the HRD in a timely manner in the bitstream or by other means.

[0053] When these NAL units (or only some of them) are conveyed to the decoder (or conveyed to the HRD) by other means not specified by this proposed | international standard, the "presence" requirement for non-VCL NAL units is also met. For the purpose of bit counting, only the appropriate bits actually present in the bitstream are counted.

[0054] As an example, synchronization of non-VCL NAL units conveyed by means other than their presence in the bitstream with NAL units present in the bitstream can be achieved by indicating two points in the bitstream between which the non-VCL NAL unit would have been present in the bitstream (if the encoder had decided to convey the unit in the bitstream).

[0055] When the content of non-VCL NAL units is conveyed by some means other than their presence within the bitstream for application, it is not required that the representation of the content of non-VCL NAL units use the same syntax.

[0056] Note that when the bitstream contains HRD information, compliance with the requirements of this subclause for the bitstream may be verified based only on the information contained in the bitstream. When HRD information is not present in the bitstream (as in the case of all "stand-alone" type I bitstreams), compliance can be verified only when HRD data is supplied by some other means not specified in this Recommendation|International Standard.

[0057] HRD contains the coded picture buffer (CPB), the instantaneous decoding process, the decoded picture buffer (DPB), and output cropping, as Figure 7 shown in.

[0058] The CPB size (number of bits) is CpbSize[SchedSelIdx]. For each X in the range from 0 to sps_max_temporal_layers_minus1, inclusive (0 and sps_max_temporal_layers_minus1), the DPB size (number of picture storage buffers) for temporal layer X is sps_max_dec_pic_buffering[X].

[0059] The variable SubPicCpbPreferredFlag is specified by external means or is set to 0 when not specified by external means.

[0060] The variable SubPicCpbFlag is derived as follows:

[0061] SubPicCpbFlag = SubPicCpbPreferredFlag && sub_pic_cpb_params_present_flag

[0062] If SubPicCpbFlag equals 0, the CPB operates at the access unit level and each decoding unit is an access unit. Otherwise, the CPB operates at the sub-picture level and each decoding unit is a subset of the access units.

[0063] HRD operates as follows. Data related to the decoding units flowing into the CPB according to the specified arrival schedule is delivered by the HSS. At the CPB removal time, the data related to each decoding unit is instantaneously removed and decoded by the instantaneous decoding process. Each decoded picture is placed in the DPB. At a later DPB output time or at a time when inter-picture prediction references are no longer needed, the decoded pictures are removed from the DPB.

[0064] Initialize the HRD as specified by the buffering period SEI. The removal timing for removing the decoding units from the CPB and the output timing for outputting the decoded pictures from the DPB are specified in the picture timing SEI message. All timing information related to a particular decoding unit will arrive before the cpb removal time of the decoding unit.

[0065] Use HRD to check the compliance of the bitstream and the decoder.

[0066] Although compliance is guaranteed assuming that all frame rates and clocks used to generate the bitstream exactly match the values signaled in the bitstream, in an actual system, each of these frame rates and clocks may be different from the signaled or specified values.

[0067] All arithmetic is done with the actual values, so rounding errors do not propagate. For example, immediately before or after removal by the decoding unit, the number of bits in the CPB may not be an integer.

[0068] Variable t c is derived as follows and is called the clock tick:

[0069] t c = num_units_in_tick ÷ time_scale

[0070] Variable t c_sub is derived as follows and is called the sub-picture clock tick:

[0071] t c_sub = num_units_in_sub_tick ÷ time_scale

[0072] The following are specified to express the constraints:

[0073] - Let access unit n be the n-th access unit in decoding order, where the first access unit is access unit 0.

[0074] - Let picture n be the coded or decoded picture of access unit n.

[0075] - Let decoding unit m be the m-th access unit in decoding order, where the first decoding unit is decoding unit 0.

[0076] In [2], the slice header syntax allows so-called dependent slices.

[0077] Figure 8 Shows the slice header syntax of [2].

[0078] The slice header semantics are defined as follows:

[0079] A dependent_slice_flag equal to 1 specifies that the value of each non - existent slice header syntax element is equal to the value of the corresponding slice header syntax element in the previous slice pair containing the coding tree block, for which the coding tree block address is SliceCtbAddrRS - 1. When it does not exist, it is inferred that the value of dependent_slice_flag is equal to 0. When SliceCtbAddrRS is equal to 0, the value of dependent_slice_flag will be equal to 0.

[0080] slice_address specifies the address in the slice granularity resolution at which the slice starts. The length of the slice_address syntax element is (Ceil(Log2(PicWidthInCtbs*PicHeightInCtbs)) + SliceGranularity) bits.

[0081] The variable SliceCtbAddrRS that specifies the coding tree block (where the slice starts in the coding tree block raster scan order) is derived as follows.

[0082] SliceCtbAddrRS = (slice_address >> SliceGranularity)

[0083] The variable SliceCbAddrZS that specifies the address of the first coding block in the slice in the minimum coding block granularity and in the z - scan order is derived as follows.

[0084] SliceCbAddrZS = slice_address

[0085] << ((log2_diff_max_min_coding_block_size - SliceGranularity) << 1)

[0086] Slice decoding starts from the largest coding unit possible at the slice start coordinates.

[0087] first_slice_in_pic_flag indicates whether the slice is the first slice of the picture. If first_slice_in_pic_flag is equal to 1, both the variables SliceCbAddrZS and SliceCtbAddrRS are set to 0 and decoding starts from the first coding tree block in the picture.

[0088] The pic_parameter_set_id specifies the picture parameter set in use. The value of pic_parameter_set_id shall be in the range of 0 to 255, inclusive of 0 and 255.

[0089] The num_entry_point_offsets specifies the number of the entry_point_offset[i] syntax elements in the slice header. When tiles_or_entropy_coding_sync_idc is equal to 1, the value of num_entry_point_offsets shall be in the range of 0 to (num_tile_columns_minus1+1)*(num_tile_rows_minus1+1)-1, inclusive of 0 and (num_tile_columns_minus1+1)*(num_tile_rows_minus1+1)-1. When tiles_or_entropy_coding_sync_idc is equal to 2, the value of num_entry_point_offsets shall be in the range of 0 to PicHeightInCtbs-1, inclusive of 0 and PicHeightInCtbs-1. When absent, the value of num_entry_point_offsets is inferred to be equal to 0.

[0090] offset_len_minus1 plus 1 specifies the length in bits of the entry_point_offset[i] syntax element.

[0091] entry_point_offset[i] specifies the i-th entry point offset in bytes and is represented by offset_len_minus1 plus 1 bits. The encoded segment data after the segment header consists of num_entry_point_offsets + 1 subsets, where the subset index values range from 0 to num_entry_point_offsets (inclusive of 0 and num_entry_point_offsets). Subset 0 consists of bytes 0 to entry_point_offset[0] - 1 of the encoded segment data (inclusive of 0 and entry_point_offset[0] - 1), subset k (where k ranges from 1 to num_entry_point_offsets - 1 (inclusive of 1 and num_entry_point_offsets - 1)) consists of bytes entry_point_offset[k - 1] to entry_point_offset[k] + entry_point_offset[k - 1] - 1 of the encoded segment data (inclusive of entry_point_offset[k - 1] and entry_point_offset[k] + entry_point_offset[k - 1] - 1), and the last subset (where the subset index equals num_entry_point_offsets) consists of the remaining bytes of the encoded segment data.

[0092] When tiles_or_entropy_coding_sync_idc equals 1 and num_entry_point_offsets is greater than 0, each subset will contain all the encoded bits of exactly one tile, and the number of subsets (i.e., the value of num_entry_point_offsets + 1) will be equal to or less than the number of tiles in the segment.

[0093] When tiles_or_entropy_coding_sync_idc equals 1, each segment must include a subset of one tile (in which case signaling the entry point is not necessary) or an integer number of complete tiles.

[0094] When tiles_or_entropy_coding_sync_idc is equal to 2 and num_entry_point_offsets is greater than 0, each subset k (where k ranges from 0 to num_entry_point_offsets - 1, inclusive of 0 and num_entry_point_offsets - 1) will contain all the coded bits of exactly one column of coding tree blocks, and the last subset (where the subset index is equal to num_entry_point_offsets) will contain all the coded bits of the remaining coding blocks included in the slice, where the remaining coding blocks consist of exactly one column of coding tree blocks or a subset of a column of coding tree blocks, and the number of subsets (i.e., the value of num_entry_point_offsets + 1) will be equal to the number of columns of coding tree blocks in the slice, where subsets of a column of coding tree blocks in the slice are also counted.

[0095] When tiles_or_entropy_coding_sync_idc is equal to 2, the slice may include several columns of coding tree blocks and a subset of a column of coding tree blocks. For example, if the slice includes 2.5 columns of coding tree blocks, and the number of subsets (i.e., the value of num_entry_point_offsets + 1) will be equal to 3.

[0096] Figure 9 The picture parameter set RBSP syntax of [2] is shown, and the picture parameter set RBSP semantics of [2] are defined as follows:

[0097] A value of dependent_slice_enabled_flag equal to 1 specifies the presence of the syntax element dependent_slice_flag in the slice header of a coded picture that references this picture parameter set. A value of dependent_slice_enabled_flag equal to 0 specifies the absence of the syntax element dependent_slice_flag in the slice header of a coded picture that references this picture parameter set. When tiles_or_entropy_coding_sync_idc is equal to 3, the value of dependent_slice_enabled_flag will be equal to 1.

[0098] tiles_or_entropy_coding_sync_idc equal to 0 specifies that there will be only one tile in each picture that refers to this picture parameter set, no specific synchronization process for context variables will be called before decoding the first coding tree block of a column of coding tree blocks in each picture that refers to this picture parameter set, and the values of cabac_independent_flag and dependent_slice_flag of the coded pictures that refer to this picture parameter set will not both be equal to 1.

[0099] When cabac_independent_flag and dependent_slice_flag are both equal to 1 for a slice, that slice is an entropy slice.

[0100] tiles_or_entropy_coding_sync_idc equal to 1 specifies that there will be more than one tile in each picture that refers to this picture parameter set, no specific synchronization process for context variables will be called before decoding the first coding tree block of a column of coding tree blocks in each picture that refers to this picture parameter set, and the values of cabac_independent_flag and dependent_slice_flag of the coded pictures that refer to this picture parameter set will not both be equal to 1.

[0101] tiles_or_entropy_coding_sync_idc equal to 2 specifies that there will be only one tile in each picture that refers to this picture parameter set, a specific synchronization process for context variables will be called before decoding the first coding tree block of a column of coding tree blocks in each picture that refers to this picture parameter set, and a specific memory process for context variables will be called after decoding two coding tree blocks of a column of coding tree blocks in each picture that refers to this picture parameter set, and the values of cabac_independent_flag and dependent_slice_flag of the coded pictures that refer to this picture parameter set will not both be equal to 1.

[0102] tiles_or_entropy_coding_sync_idc equal to 3 specifies that there will be only one tile in each picture that refers to this picture parameter set, no specific synchronization process for context variables will be called before decoding the first coding tree block of a column of coding tree blocks in each picture that refers to this picture parameter set, and the values of cabac_independent_flag and dependent_slice_flag of the coded pictures that refer to this picture parameter set may both be equal to 1.

[0103] When the dependent_slice_enabled_flag is not equal to 0, the tiles_or_entropy_coding_sync_idc will not be equal to 3.

[0104] The requirement for bitstream conformance is that for all picture parameter sets initiated within the coded video sequence, the value of tiles_or_entropy_coding_sync_idc will be the same.

[0105] For each slice that refers to the picture parameter set, when tiles_or_entropy_coding_sync_idc is equal to 2 and the first coded block in the slice is not the first coded block in the first coded tree block of a column of coded tree blocks, the last coded block in the slice will be in the same column of coded tree blocks as the first coded block in the slice.

[0106] num_tile_columns_minus1 plus 1 specifies the number of tile rows that divide the picture.

[0107] num_tile_rows_minus1 plus 1 specifies the number of tile columns that divide the picture. When num_tile_columns_minus1 is equal to 0, num_tile_rows_minus1 will not be equal to 0.

[0108] uniform_spacing_flag being equal to 1 specifies that the row boundaries and column boundaries are distributed evenly across the picture. uniform_spacing_flag being equal to 0 specifies that the row boundaries and column boundaries are not distributed evenly across the picture, but rather the syntax elements column_width[i] and row_height[i] are used to explicitly signal the row boundaries and column boundaries.

[0109] column_width[i] specifies the width of the i-th tile row in units of coded tree blocks.

[0110] row_height[i] specifies the height of the i-th tile column in units of coded tree blocks.

[0111] The vector colWidth[i] specifies the width of the i-th tile row in units of CTBs, where row i is in the range from 0 to num_tile_columns_minus1 (including 0 and num_tile_columns_minus1).

[0112] The vector CtbAddrRStoTS[ctbAddrRS] specifies the conversion from the CTB address in raster scan order to the CTB address in tile scan order, where the index ctbAddrRS ranges from 0 to (picHeightInCtbs * picWidthInCtbs) - 1 (including 0 and (picHeightInCtbs * picWidthInCtbs) - 1).

[0113] The vector CtbAddrTStoRS[ctbAddrTS] specifies the conversion from the CTB address in tile scan order to the CTB address in raster scan order, where the index ctbAddrTS ranges from 0 to (picHeightInCtbs * picWidthInCtbs) - 1 (including 0 and (picHeightInCtbs * picWidthInCtbs) - 1).

[0114] The vector TileId[ctbAddrTS] specifies the conversion from the CTB address in tile scan order to the tile id, where ctbAddrTS ranges from 0 to (picHeightInCtbs * picWidthInCtbs) - 1 (including 0 and (picHeightInCtbs * picWidthInCtbs) - 1).

[0115] The values of colWidth, CtbAddrRStoTS, CtbAddrTStoRS, and TileId are derived by calling the CTB raster and tile scan conversion procedures as specified in subclause 6.5.1, which take PicHeightInCtbs and PicWidthInCtbs as inputs and assign the outputs to colWidth, CtbAddrRStoTS, CtbAddrTStoRS, and TileId.

[0116] The value of ColumnWidthInLumaSamples[i] (which specifies the width of the i-th tile row in luma samples) is set to be equal to colWidth[i] << Log2CtbSize.

[0117] The array MinCbAddrZS[x][y] (which specifies the conversion from the position (x,y) in units of the smallest CB to the smallest CB address in the z-scan order, where x ranges from 0 to picWidthInMinCbs-1 (including 0 and picWidthInMinCbs-1), and y ranges from 0 to picHeightInMinCbs-1 (including 0 and picHeightInMinCbs-1)) is derived by calling the Z-scan order array initialization process specified in Subclause 6.5.2, which takes Log2MinCbSize, Log2CtbSize, PicHeightInCtbs, PicWidthInCtbs, and the vector CtbAddrRStoTS as inputs and assigns the output to MinCbAddrZS.

[0118] The loop_filter_across_tiles_enabled_flag being equal to 1 specifies that loop filter operations are performed across tiles. The loop_filter_across_tiles_enabled_flag being equal to 0 specifies that loop filter operations are not performed across tiles. The loop filter operations include deblocking filtering, sample adaptive offset, and adaptive loop filtering operations. When not present, it is inferred that the value of the loop_filter_across_tiles_enabled_flag is equal to 1.

[0119] The cabac_independent_flag being equal to 1 specifies that the CABAC decoding of the coded blocks in a slice is independent of any state of the previously decoded slice. The cabac_independent_flag being equal to 0 specifies that the CABAC decoding of the coded blocks in a slice depends on the state of the previously decoded slice. When not present, it is inferred that the value of the cabac_independent_flag is equal to 0.

[0120] The process for deriving the availability of the coded block with the smallest coded block address is described as follows:

[0121] The inputs to this process are

[0122] - The smallest coded block address minCbAddrZS in the z-scan order

[0123] - The current smallest coded block address currMinCBAddrZS in the z-scan order

[0124] The output of this process is the availability cbAvailable of the coded block with the smallest coded block address in the z-scan order.

[0125] Note that when this procedure is called, the meaning of determining availability is considered.

[0126] Note that any coding block, regardless of its size, is related to the minimum coding block address, which is the address of the coding block with the minimum coding block size in the z-scan order.

[0127] - If one or more of the following conditions are true, set cbAvailable to false.

[0128] - minCbAddrZS is less than 0

[0129] - minCbAddrZS is greater than currMinCBAddrZS

[0130] - The coding block with the minimum coding block address minCbAddrZS belongs to a different slice from the coding block with the current minimum coding block address currMinCBAddrZS, and the dependent_slice_flag of the slice containing the coding block with the current minimum coding block address currMinCBAddrZS is equal to 0

[0131] - The coding block with the minimum coding block address minCbAddrZS is included in a different tile from the coding block with the current minimum coding block address currMinCBAddrZS.

[0132] - Otherwise, set cbAvailable to true.

[0133] The CABAC parsing procedure for slice data in [2] is as follows:

[0134] This procedure is called when parsing a syntax element with descriptor ae(v).

[0135] The input to this procedure is a request for the value of the syntax element and the values of previously parsed syntax elements.

[0136] The output of this procedure is the value of the syntax element.

[0137] When starting to parse the slice data of a slice, the initialization procedure of the CABAC parsing procedure is called.

[0138] Using the position (x0, y0) of the top-left luminance sample of the current coding tree block, the minimum coding block address of the coding tree block containing the spatially adjacent block T( Figure 10A ) is derived as follows.

[0139] x = x0 + 2 << Log2CtbSize - 1

[0140] y = y0 - 1

[0141] ctbMinCbAddrT = MinCbAddrZS[x >> Log2MinCbSize][y >> Log2MinCbSize]

[0142] The variable availableFlagT is obtained by invoking an encoding block availability derivation procedure with ctbMinCbAddrT as the input.

[0143] When starting the parsing of the coding tree, the following ordered steps apply.

[0144] 1. Initialize the arithmetic decoding engine as follows.

[0145] - If CtbAddrRS is equal to Slice_address, dependent_slice_flag is equal to 1, and entropy_coding_reset_flag is equal to 0, then the following applies.

[0146] - Invoke the synchronization procedure of the CABAC parsing process, which takes TableStateIdxDS and TableMPSValDS as inputs.

[0147] - Invoke the decoding procedure for binary decisions before termination, followed by the initialization procedure for arithmetic decoding.

[0148] - Otherwise, if tiles_or_entropy_coding_sync_idc is equal to 2 and CtbAddrRS % PicWidthInCtbs is equal to 0, then the following applies.

[0149] - When availableFlagT is equal to 1, invoke the synchronization procedure of the CABAC parsing process, which takes TableStateIdxWPP and TableMPSValWPP as inputs.

[0150] - Invoke the decoding procedure for binary decisions before termination, followed by the initialization procedure for the arithmetic decoding engine.

[0151] 2. When cabac_independent_flag is equal to 0 and dependent_slice_flag is equal to 1, or when tiles_or_entropy_coding_sync_idc is equal to 2, apply the memory process as follows.

[0152] - When tiles_or_entropy_coding_sync_idc is equal to 2 and CtbAddrRS % PicWidthInCtbs is equal to 2, call the memory process of the CABAC parsing process, which takes TableStateIdxWPP and TableMPSValWPP as inputs.

[0153] - When cabac_independent_flag is equal to 0, dependent_slice_flag is equal to 1, and end_of_slice_flag is equal to 1, call the memory process of the CABAC parsing process, which takes TableStateIdxDS and TableMPSValDS as outputs.

[0154] The parsing of syntax elements is performed as follows:

[0155] For each requested value of the syntax element, derive the binarization.

[0156] The binarization of the syntax element and the decoded process flow of the parsed binary number are determined.

[0157] For each binary number of the binarization of the syntax element (indexed by the variable binIdx), derive the context index ctxIdx.

[0158] For ctxIdx, call the arithmetic decoding process.

[0159] After decoding each binary number, compare the obtained sequence of the parsed binary number (b 0.. b binIdx ) with the set of binary number strings given by the binarization process. When the sequence matches a binary number string in the given set, assign the corresponding value to the syntax element.

[0160] When processing a request for the value of the syntax element pcm-flag and the decoded value of pcm_flag is equal to 1, initialize the decoding engine after decoding any pcm_alignment_zero_bit, num_subsequent_pcm, and all pcm_sample_luma and pcm_sample_chroma data.

[0161] In the design framework described so far, the following problems occur.

[0162] Before encoding and sending data in a low-latency situation, the timing of the decoding unit needs to be known, where the NAL unit will have been sent by the encoder while the encoder is still encoding parts of the image, i.e., other sub-image decoding units. That is, since the NAL unit order in the access unit only allows SEI messages to be before the VCL (Video Coding NAL unit) in the access unit, but in this low-latency situation, if the encoder starts encoding the decoding unit, the non-VCL NAL units need to already be online, i.e., sent. Figure 10B Illustrates the structure of an access unit as defined in [2]. [2] does not specify the end of a sequence or stream, so its presence in the access unit is assumed.

[0163] In addition, in a low-latency situation, the number of NAL units related to sub-images also needs to be known in advance, because the picture timing SEI message contains this information and the picture timing SEI message must be compiled and sent before the encoder starts encoding the actual image. Application designers who do not want to insert filler data NAL units (there may be no filler data to conform to the number of NAL units, such as signaled for each decoding unit in the picture timing SEI) need a means to signal this information at the sub-image level. This also applies to sub-image timing, which is currently fixed at the start of the access unit by the parameters given in the timing SEI message.

[0164] Another disadvantage of the draft specification [2] includes a large amount of signaling at the sub-image level, which is required for specific applications, such as ROI signaling or slice size signaling.

[0165] The problems outlined above are not specific to the HEVC standard. On the contrary, the same problems can also occur with other video coding decoders. Figure 11 More generally shows a video transmission scenario, where a pair of an encoder 10 and a decoder 12 are connected via a network 14 to transmit video 16 from the encoder 10 to the decoder 12 with a short end-to-end delay. The scenario outlined above is as follows. The encoder 10 encodes a sequence of frames 18 of the video 16 according to a specific decoding order that generally but not necessarily follows the reproduction order 20 of the frames 18, and traverses the frame area of each frame 18 in a defined manner within each frame 18, such as in a raster scan manner with or without tile partitioning of the frame 18. The decoding order controls the availability of information for the encoding techniques (such as prediction and / or entropy coding) used by the encoder 10, i.e., the availability of information related to spatially and / or temporally adjacent parts of the video 16, which can be used as a basis for prediction or context selection. Even though the encoder 10 may be able to use parallel processing to encode the frames 18 of the video 16, the encoder 10 definitely needs some time to encode a specific frame 18, such as the current frame. For example, Figure 11An example shows a moment when the encoder 10 has completed encoding a part 18a of the current frame 18, while another part 18b of the current frame 18 has not been encoded yet. Since the encoder 10 has not encoded the part 18b, the encoder 10 may not be able to predict how the available bitrate for encoding the current frame 18 should be spatially distributed over the current frame 18 to achieve, for example, an optimization in terms of rate / distortion. Therefore, the encoder 10 has only two options: the encoder 10 pre-estimates an almost optimal distribution of the available bitrate of the current frame 18 over segments (the current frame 18 is spatially subdivided into these segments), and accordingly accepts that this estimate may be incorrect; or the encoder 10 completes encoding the current frame 18 before transmitting packets containing these segments from the encoder 10 to the decoder 12. In any case, in order to be able to utilize any transmission of packets of segments of the current encoded frame 18 before completing the encoding of the current encoded frame 18, the bitrate associated with each such packet of segments should be notified to the segments 14 in the form of the capture time of the encoded image buffer. However, as indicated above, although the encoder 10 can change the bitrate distributed over the frame 18 by using the decoder buffer capture times defined separately for sub-image regions according to the current version of HEVC, the encoder 10 needs to transmit or emit this information via the network 14 at the beginning of each access unit that collects all the data related to the current frame 18, thereby forcing the encoder 10 to choose between the two alternative solutions outlined above, one resulting in lower latency but worse rate / distortion, and the other resulting in optimal rate / distortion but an increase in end-to-end latency.

[0166] Therefore, so far no video codec has allowed achieving such a low latency that the encoder would be allowed to start transmitting packets related to the part 18a of the current frame before encoding the remaining part 18b of the current frame, and the decoder could utilize this intermediate transmission of packets related to the preliminary part 18a via the network 16, which complies with the decoding buffer capture timing conveyed in the video data stream sent from the encoder 12 to the decoder 14. Examples of applications that would utilize such a low latency include industrial applications, such as workpiece or manufacturing monitoring for automation or inspection purposes or the like. So far, there has also been no satisfactory solution to notify the decoding side of the association of packets with the tiles constituting the current frame and the region of interest (region of concern) of the current frame, thus allowing intermediate network entities within the network 16 to collect this information from the data stream without delving into the internal inspection of the packets, i.e., the slice syntax.

[0167] Therefore, the object of the present invention is to provide a video data stream encoding concept technology that more efficiently allows for low end-to-end latency and / or makes it easier to identify parts of the data stream to regions of interest or to specific tiles.

[0168] This object is achieved by the subject matter described in the appended independent claims. Summary of the Invention

[0169] The present application relates to a video data stream having video content encoded therein in units of sub - parts of images of the video content, each sub - part being encoded into one or more payload packets of a packet sequence of the video data stream, the packet sequence being divided into a sequence of access units, so that each access unit collects the payload packets related to the respective images of the video content, wherein the packet sequence has timing control packets interspersed therein, so that the timing control packets subdivide the access units into decoding units, so that at least some access units are subdivided into two or more decoding units, wherein each timing control packet signals the decoder buffer fetch time of the decoding unit, and the payload packets of the decoding unit follow each timing control packet in the packet sequence.

[0170] One idea on which the present application is based is that decoder fetch timing information, ROI information, and tile identification information should be conveyed within the video data stream at a level that allows easy access by network entities such as MANE or the decoder; and to achieve this level, such types of information should be conveyed within the video data stream by packets interspersed in the packets of the access units of the video data stream. According to one embodiment, the interspersed packets are of a removable packet type, i.e., the removal of such interspersed packets maintains the ability of the decoder to fully reconstruct the video content conveyed via the video data stream.

[0171] According to one aspect of the present application, by using interspersed packets to convey information about the decoder buffer fetch time of decoding units (the decoding units are formed by payload packets that follow individual timing control packets of the video data stream) within the current access unit, achieving a low end - to - end delay is made more efficient. By this measure, the encoder is allowed to determine the decoder buffer fetch time on the fly during the encoding of the current frame. The encoder can then, while encoding the current frame, on the one hand continue to determine the bit rate of the part of the current frame that has already been encoded into payload packets and transmitted or signaled (prefixed by a timing control packet), and accordingly adjust the distribution of the remaining bit rate available for the current frame over the remaining non - encoded part of the current frame. By this measure, the available bit rate is utilized efficiently, and the delay remains short because the encoder does not need to wait for the encoding of the current frame to be completely finished.

[0172] According to another aspect of the present application, packets interspersed in the payload packets of an access unit are used to convey information about the region of interest, thereby allowing easy access to this information by network entities as outlined above, since these network entities do not need to examine the intermediate payload packets. In addition, the encoder can still determine in real time during the encoding of the current frame the packets belonging to the ROI without having to pre-determine the subdivision of the current frame into sub-parts and individual payload packets. Further, according to an embodiment (wherein the interspersed packets are of a removable packet type), video data stream receivers that are not interested in or unable to process the ROI information may ignore the ROI information.

[0173] According to another aspect, a similar idea is utilized in the present application, according to which the interspersed packets convey information about which tile a particular packet within an access unit belongs to. BRIEF DESCRIPTION OF THE DRAWINGS

[0174] Advantageous implementations of the invention are the subject matter of the dependent claims. The preferred embodiments of the present application are described in more detail below with respect to the figures, wherein:

[0175] Figures 1 to 10B The current state of HEVC is shown, wherein Figure 1 The syntax of the buffering period SEI message is shown, Figure 2 The syntax of the picture timing SEI message is shown, Figure 3A and Figure 3B The syntax of the VUI parameters is shown, Figure 4 The syntax of the HRD parameters is shown, Figure 5 The syntax of the filler data RBSP is shown, Figure 6 The structure of the bitstream and NAL unit stream for HRD compliance checking is shown, Figure 7 The HRD buffer model is shown, Figure 8 The slice header syntax is shown, Figure 9 The syntax of the picture parameter set RBSP is shown, Figure 10A A schematic diagram showing the spatial neighboring coding tree blocks T (which may be used to invoke the coding tree block availability derivation process related to the current coding tree block) is shown, and Figure 10B The definition of the structure of an access unit is shown;

[0176] Figure 11 A pair of encoder and decoder connected via a network is schematically shown to illustrate the problems that occur during the transmission of a video data stream;

[0177] Figure 12 A schematic block diagram of an encoder according to an embodiment using timing control packets is shown;

[0178] Figure 13Shows a flowchart illustrating the operation mode of an encoder according to an embodiment; Figure 12 of the encoder;

[0179] Figure 14 Shows a block diagram of an embodiment of a decoder for explaining the functionality of the decoder with respect to a video data stream generated by an encoder according to Figure 12 the encoder;

[0180] Figure 15 Shows a schematic block diagram illustrating an encoder, a network entity, and a video data stream according to another embodiment using ROI;

[0181] Figure 16 Shows a schematic block diagram illustrating an encoder, a network entity, and a video data stream according to another embodiment using tile identification packets;

[0182] Figure 17 Shows the structure of an access unit according to an embodiment. The dashed line reflects the case of a non-mandatory fragment prefix NAL unit;

[0183] Figure 18 Shows the use of tiles in signaling the region of interest;

[0184] Figure 19 Shows the first simple syntax / version 1;

[0185] Figure 20 Shows the extended syntax / version 2, which includes tile_id signaling, decoding unit start identifier, fragment prefix ID, and fragment header data in addition to the SEI message concept technology;

[0186] Figure 21 Shows the NAL unit type code and NAL unit type category;

[0187] Figure 22 Shows the possible syntax of the fragment header, where specific syntax elements present in the fragment header according to the current version are transformed into lower-level syntax elements called slice_header_data();

[0188] Figure 23A 、 Figure 23B and Figure 23C Shows a table of all syntax elements signaled via the syntax element fragment header data removed from the fragment header;

[0189] Figure 24 Shows the supplementary enhancement information message syntax;

[0190] Figure 25A and Figure 25BShows the adapted SEI payload syntax to introduce new tile or sub - picture SEI message types;

[0191] Figure 26 Shows an example of a sub - picture buffer SEI message;

[0192] Figure 27 Shows an example of a sub - picture timing SEI message;

[0193] Figure 28 Shows what a sub - picture segment information SEI message might look like;

[0194] Figure 29 Shows an example of a sub - picture tile information SEI message;

[0195] Figure 30 Shows a syntax example of a sub - picture segment size information SEI message;

[0196] Figure 31 Shows a first variant of a syntax example of a region - of - interest SEI message, where each ROI is signaled in an individual SEI message;

[0197] Figure 32 Shows a second variant of a syntax example of a region - of - interest SEI message, where all ROIs are signaled in a single SEI message;

[0198] Figure 33 Shows a possible syntax of a timing control packet according to another embodiment;

[0199] Figure 34 Shows a possible syntax of a tile identification packet according to an embodiment;

[0200] Figures 35 to 38 Shows a possible subdivision of an image according to different subdivision settings according to an embodiment; and

[0201] Figure 39 Shows an example of a part of a video data stream according to an embodiment using timing control packets that are interleaved between the payload packets of access units. Detailed Description

[0202] Regarding Figure 12 , describes an encoder 10 and its operating mode according to an embodiment of the present application. The encoder 10 is configured to encode video content 16 into a video data stream 22. The encoder is configured to perform this encoding in units of sub - parts of the frames / images 18 of the video content 16, where these sub - parts can be, for example, segments 24 into which the image 18 is divided, or some other spatial sections, such as tiles 26 or WPP sub - streams 28, all as described in Figure 12is described therein, and it is for illustrative purposes only and does not imply that the encoder 10 needs to be able to support, for example, tile or WPP parallel processing or that the sub-parts need to be segments.

[0203] When encoding video content 16 in units of sub-parts 24, the encoder 10 may follow the decoding order (or encoding order) defined between the sub-parts 24, which, for example, traverses the images 18 of video 16 according to the picture decoding order (which may not necessarily conform to the presentation order 20 defined between the images 18), and traverses the blocks into which each image 18 is divided within each image 18 according to the raster scan order, where the sub-parts 24 represent consecutive runs of such blocks along the decoding order. Specifically, the encoder 10 may be configured to: when determining the availability of the spatial and / or temporal neighboring parts of the currently to-be-encoded part for using the attributes of such neighboring parts to determine the prediction and / or entropy context in, for example, predictive coding and / or entropy coding, follow this decoding order. Only the previously visited (encoded / decoded) parts of the video are available. Otherwise, set the just-mentioned attributes to default values, or take some other alternative measures.

[0204] On the other hand, the encoder 10 does not need to encode the sub-parts 24 along the decoding order serially. Instead, the encoder 10 may use parallel processing to accelerate the encoding process, or be able to perform more complex encoding in real time. Similarly, the encoder 10 may or may not be configured to transmit or emit data encoding the sub-parts along the decoding order. For example, the encoder 10 may output / transmit the encoded data in some other order, such as according to the order in which the encoder 10 finishes encoding the sub-parts (which may deviate from the just-mentioned decoding order due to, for example, parallel processing).

[0205] To make the encoded version of the sub-parts 24 suitable for transmission via a network, the encoder 10 encodes each sub-part 24 into one or more payload packets of a packet sequence of the video data stream 22. In the case where the sub-parts 24 are segments, the encoder 10 may, for example, be configured to put each segment data (i.e., each encoded segment) into one payload packet (such as a NAL unit). This packetization can make the video data stream 22 suitable for transmission via a network. Thus, the packets may represent the smallest units in which the video data stream 22 may occur, i.e., the smallest units that can each be emitted by the encoder 10 for transmission via the network to a receiver.

[0206] In addition to the payload packets and the timing control packets interspersed therebetween and discussed hereinafter, other packets (i.e., other types of packets) may also exist, such as padding data packets, picture or sequence parameter set packets for conveying syntax elements that do not change frequently, or EOF (end of file) or AUE (end of access unit) packets or the like.

[0207] The encoder performs encoding into payload packets such that the packet sequence is divided into a sequence of access units 30, and each access unit collects the payload packets 32 related to one picture 18 of the video content 16. That is, the packet sequence 34 forming the video data stream 22 is subdivided into non-overlapping portions called access units 30, each non-overlapping portion being related to an individual one of the pictures 18. The sequence of access units 30 may follow the decoding order of the pictures 18 to which the access units 30 pertain. Figure 12 For example, an access unit 30 illustrated in the middle of the illustrated portion of the data stream 22 contains one payload packet 32 for each sub-portion 24, the picture 18 being subdivided into sub-portions 24. That is, each payload packet 32 carries a corresponding sub-portion 24. The encoder 10 is configured to interleave timing control packets 36 in the packet sequence 34, whereby the timing control packets subdivide the access units 30 into decoding units 38, such that at least some access units 30 (such as Figure 12 the middle access unit shown) are subdivided into two or more decoding units 38, each timing control packet signaling the decoder buffer fetch time of the decoding unit 38, the payload packet 32 of which follows the individual timing control packet in the packet sequence 34. In other words, the encoder 10 prefixes a subsequence of the sequence of payload packets 32 within an access unit 30 with an individual timing control packet 36, the timing control packet 36 signaling the decoder buffer fetch time for the individual subsequence of payload packets that are prefixed by the individual timing control packet 36 and form the encoding unit 38. Figure 12 For example, it is illustrated that every other packet 32 represents the first payload packet of a decoding unit 38 of an access unit 30. As Figure 12 illustrated, the amount of data or bit rate for each decoding unit 38 varies, and the decoder buffer fetch time may be related to this bit rate variation between the decoding units 38, since the decoder buffer fetch time of a decoding unit 38 may follow the decoder buffer fetch time signaled by the timing control packet 36 of the immediately preceding decoding unit 38 plus a time interval that corresponds to the bit rate for this immediately preceding decoding unit 38.

[0208] That is, the encoder 10 may, as Figure 13It operates as shown. Specifically, as mentioned above, the encoder 10 may subject the current sub - part 24 of the current image 18 to encoding in step 40. As already mentioned, the encoder 10 may sequentially loop through the sub - parts 24 in the decoding order as illustrated by arrow 42, or the encoder 10 may use some parallel processing, such as WPP and / or tile processing, to encode several "current sub - parts" 24 simultaneously. Whether parallel processing is used or not, the encoder 10 forms a decoding unit from one or several of the sub - parts just encoded in step 40, and proceeds to step 44, where the encoder 10 sets the decoder buffer fetch time of this decoding unit and transmits this decoding unit prefixed with a timing control packet that signals the just - set decoder buffer fetch time of this decoding unit. For example, the encoder 10 may determine the decoder buffer fetch time in step 44 based on the bitrate for encoding the sub - parts in the payload packets that have been encoded to form the current decoding unit (which includes all other intermediate packets (if any) within this decoding unit, i.e., "prefix packets")).

[0209] Then, in step 46, the encoder 10 may adjust the available bitrate based on the bitrate already used for the decoding unit just transmitted in step 44. For example, if the image content within the decoding unit just transmitted in step 44 is very complex in terms of the compression ratio, the encoder 10 may reduce the available bitrate for the next decoding unit in order to comply with a certain externally - set target bitrate that has been determined based on, for example, the current bandwidth situation faced by the network for transmitting the video data stream 22. Then steps 40 to 46 are repeated. By this measure, the image 18 is encoded and transmitted (i.e., signaled) in units of decoding units, with each decoding unit prefixed with a uniquely corresponding timing control packet.

[0210] In other words, the encoder 10 encodes 40 the current sub - part 24 of the current image 18 into the current payload packet 32 of the current decoding unit 38 during the process of encoding the current image 18 of the video content 16, transmits 44 the current decoding unit 38 prefixed with the current timing control packet 36 at a first moment in the data stream by setting the decoder buffer fetch time signaled by the current timing control packet (36), and encodes 44 another sub - part 24 of the current image 18 at a second moment (second visit to step 40) by looping back from step 46 to 40, and this second moment is later than the first moment (first visit to step 44).

[0211] Since the encoder is able to emit this decoding unit before the remainder of the current image to which the encoding / decoding unit belongs, the encoder 10 is able to reduce the end-to-end latency. On the other hand, the encoder 10 does not need to waste the available bitrate because the encoder 10 is able to react to the specific nature of the content of the current image and the spatial distribution of its complexity.

[0212] On the other hand, an intermediate network entity responsible for further transmitting the video data stream 22 from the encoder to the decoder can use the timing control packet 36 to ensure that any decoder receiving the video data stream 22 receives the decoding unit in time so as to be able to utilize the per-decoding-unit encoding and transmission performed by the encoder 10. For example, see Figure 14 , which shows an example of a decoder for decoding the video data stream 22. The decoder 12 receives the video data stream 22 at the encoded picture buffer CPB 48 via the network, where the encoder 10 transmits the video data stream 22 to the decoder 12 via the network. Specifically, since it is assumed that the network 14 can support low-latency applications, the network 10 checks the decoder buffer fetch time in order to feed the packet sequence 34 of the video data stream 22 into the encoded picture buffer 48 of the decoder 12, so that each decoding unit is present in the encoded picture buffer 48 before the decoder buffer fetch time signaled by the timing control packet prefixed to the individual decoding unit. By this measure, the decoder can use the decoder buffer fetch time in the timing control packet to empty the encoded picture buffer 48 of the decoder in units of decoding units rather than complete access units without stalling, that is, without running out of the available payload packets in the encoded picture buffer 48. For example, Figure 14 for illustrative purposes, a processing unit 50 is shown connected to the output of the encoded picture buffer 48, and the input of the encoded picture buffer 48 receives the video data stream 22. Similar to the encoder 10, the decoder 12 may be able to perform parallel processing such as using tile parallel processing / decoding and / or WPP parallel processing / decoding.

[0213] As will be outlined in more detail below, the decoder buffer fetch time mentioned so far does not necessarily relate to the fetch time of the encoded picture buffer 48 with respect to the decoder 12. Instead, the timing control packet may additionally or otherwise manipulate the fetch of the decoded picture data of the corresponding decoded picture buffer of the decoder 12. For example, Figure 14It shows that the decoder 12 includes a decoder picture buffer in which the decoded version of the video content is buffered (i.e., stored and output) in units of decoded units of the decoded version, such as the decoded version obtained by the processing unit 50 by decoding the video data stream 22. The decoded picture buffer 22 of the decoder can thus be connected between the output of the decoder 12 and the output of the processing unit 50. By being able to set the capture time for outputting the decoded version of the decoded unit from the decoded picture buffer 52, the encoder 10 has the opportunity to control in real time (i.e., during the process of encoding the current picture) the reproduction of the video content on the decoding side or the end-to-end delay of such reproduction, even with a granularity smaller than the picture rate or frame rate. Obviously, over-segmenting each picture 18 into a large number of sub-parts 24 on the encoding side will adversely affect the bit rate for transmitting the video data stream 22, but on the other hand, the end-to-end delay can be minimized because the time required to encode and transmit and decode and output this decoded unit will be minimized. On the other hand, increasing the size of the sub-parts 24 will increase the end-to-end delay. Therefore, a compromise must be found. Using the decoder buffer capture time mentioned just above to manipulate the output timing of the decoded version of the sub-parts 24 in units of decoded units allows the encoder 10 or some other unit on the encoding side to spatially adapt this compromise to the content of the current picture. By this measure, it will be possible to control the end-to-end delay in a manner that varies spatially across the content of the current picture.

[0214] When implementing the embodiments outlined above, it is possible to use packets of the removable packet type as timing control packets. Packets of the removable packet type are not necessary for reconstructing the video content on the decoding side. Hereinafter, these packets are referred to as SEI packets. Other packets of the removable packet type may also exist, i.e., another type of removable packet, such as (if transmitted in the stream) redundant packets. As another alternative, the timing control packet may be a packet of a specific removable packet type, which however additionally carries a specific SEI packet type field. For example, the timing control packet may be an SEI packet, where each SEI packet carries one or several SEI messages, and only those SEI packets containing a specific type of SEI message form the above-mentioned timing control packet.

[0215] Therefore, according to another embodiment, applying the embodiments described so far with respect to Figures 12 to 14 the described embodiments to the HEVC standard, a possible conceptual technique for making HEVC more efficient in achieving lower end-to-end delays is thus formed. In this way, the packets mentioned above are formed by NAL units, and the above-mentioned payload packets are the VCL NAL units of the NAL units, where the slices form the sub-parts mentioned above.

[0216] However, before this description of more detailed embodiments, other embodiments are described which are in line with the embodiments outlined above in that interspersed packets are used to convey information describing the video data stream in an efficient manner, but the classification of the information is different from the above embodiments, in which the timing control packets convey decoder buffer fetch timing information. In the embodiments further described below, the type of information passed via the interspersed packets interspersed in the payload packets belonging to the access unit is related to region of interest (ROI) information and / or tile identification information. The embodiments further described below may or may not be combined with the embodiments described with respect to Figures 12 to 14 the embodiments described.

[0217] Figure 15 An encoder 10 is shown which, apart from the interspersion of the timing control packets and the functionality described above with respect to Figure 13 which is optional for the Figure 15 encoder 10, operates similarly to the encoder explained above with respect to Figure 12 However, Figure 15 the encoder 10 is configured to encode the video content 16 into the video data stream 22 in units of sub-parts 24 of the images 18 of the video content 16, as explained above with respect to Figure 11 When encoding the video content 16, the encoder 10 is interested in conveying information about the region of interest ROI 60 to the decoding side together with the video data stream 22. The ROI 60 is a spatial sub-region of the current image 18 that the decoder should pay special attention to, for example. The spatial position of the ROI 60 can be input from the outside, such as by user input (illustrated by the dashed line 62), or can be automatically determined in real time by the encoder 10 or by some other entity during the encoding of the current image 18. In either case, the encoder 10 faces the following problem: The indication of the position of the ROI 60 is in principle not a problem for the encoder 10. To this end, the encoder 10 can easily indicate the position of the ROI 60 within the data stream 22. However, in order to make this information easily accessible, Figure the encoder 10 of

[0218] ​An example of interleaving ROI packets 64 among the payload packets 32 of the access unit 30 is shown. The ROI packet 64 indicates where in the packet sequence 34 of the video data stream 22 there is encoded data related to the ROI 60 (i.e., encoding the ROI 60). How the ROI packet 64 indicates the location of the ROI 60 can be implemented in a variety of ways. For example, the mere presence / occurrence of the ROI packet 64 can indicate that encoded data related to the ROI 60 is incorporated in one or more of the subsequent payload packets 32, which subsequent payload packets 32 follow in the sequential order of the sequence 34, i.e., are prefix payload packets. Alternatively, a syntax element in the ROI packet 64 can indicate whether one or more subsequent payload packets 32 are related to the ROI 60, i.e., at least partially encode the ROI 60. A large number of variations also stem from possible variations in the "scope" of an individual ROI packet 64 (i.e., the number of prefix payload packets prefixed by one ROI packet 64). For example, the indication of whether any encoded data related to the ROI 60 is incorporated or not incorporated in one ROI packet can be related to all payload packets 32 that follow in the sequential order of the sequence 34 until the next ROI packet 64 appears, or can be related only to the immediately subsequent payload packet 32 (i.e., the payload packet 32 that follows the individual ROI packet 64 in the sequential order of the sequence 34). In ​ FIG. 66 exemplarily illustrates a situation where the ROI packet 64 indicates ROI relevance (i.e., incorporation of any encoded data related to the ROI 60) or ROI irrelevance (i.e., absence of any encoded data related to the ROI 60) for all payload packets 32 that appear downstream of an individual ROI packet 64 until the next ROI packet 64 or the end of the current access unit 30 (whichever occurs earlier along the packet sequence 34). Specifically, ​ FIG. 64 illustrates that the ROI packet 64 has a block with a syntax element inside that indicates whether the subsequent payload packet 32 in the sequential order of the sequence 34 has any encoded data related to the ROI 60. This embodiment is also described below. However, as just mentioned, another possibility is that each ROI packet 64 indicates only by its presence in the packet sequence 34 whether the payload packets 32 belonging to the "scope" of the individual ROI packet 64 have ROI 60-related data inside, i.e., data related to the ROI 60. According to an embodiment described in more detail below, the ROI packet 64 even indicates the location of the portion of the ROI 60 encoded into the payload packets 32 belonging to its "scope".

[0219] Any network entity 68 that receives the video data stream 22 can utilize an indication of ROI correlation (such as implemented by using the ROI packet 64) to process the ROI-related portion of, for example, the packet sequence 34 with a higher priority than other portions of the packet sequence 34. Alternatively, the network entity 68 can use the ROI correlation information to perform other tasks related to the transmission of, for example, the video data stream 22. The network entity 68 can be, for example, a MANE or a decoder that is used to decode and play the video content 60 conveyed via the video data stream 22. In other words, the network entity 68 can use the identification result of the ROI packet to make decisions regarding the transmission tasks related to the video data stream. The transmission tasks can include retransmission requests for defective packets. The network entity 68 can be configured to process the region of interest 70 with an increased priority and assign a higher priority to the ROI packet 72 that is signaled to cover the region of interest and its associated payload packets (i.e., the payload packets prefixed by the ROI packet 72) compared to the ROI packets and their associated payload packets that are signaled not to cover the ROI. The network entity 68 can first request the retransmission of the payload packets assigned with a higher priority before requesting the retransmission of any payload packets assigned with a lower priority.

[0220] ​ Embodiments of can be easily combined with the embodiments previously described with respect to ​ For example, the ROI packet 64 mentioned above can also be an SEI packet that contains a specific type of SEI message, that is, an ROI SEI packet. That is, the SEI packet can be, for example, a timing control packet and at the same time an ROI packet, that is, in the case where an individual SEI packet contains both timing control information and ROI indication information. Alternatively, the SEI packet can be one of the timing control packet and the ROI packet but not the other, or can be neither an ROI packet nor a timing control packet.

[0221] According to ​ In the embodiment shown in, interspersing packets between the payload packets of an access unit is used to indicate, in a manner easily accessible by the network entity 68 that processes the video data stream 22, which tile or tiles of the current image 18 that the current access unit 30 pertains to are covered by any sub-portion encoded into any one of the payload packets 32 (with an individual packet serving as its prefix). In ​, for example, the current image 18 is shown to be subdivided into four tiles 70, which are here exemplarily formed by four quarters of the current image 18. The subdivision of the current image 18 into tiles 70 can be signaled, for example, within the video data stream in a unit containing the image sequence (such as, for example, in a VPS or SPS packet also interspersed in the packet sequence 34). As will be described in more detail below, the tile subdivision of the current image 18 can be a conventional subdivision of the image 18 into rows and columns of tiles. The number of tile rows and columns, as well as the row width and column height, can vary. In particular, the width and height of the rows / columns of tiles can be different for different columns and different rows, respectively. ​ Also shown are examples of sub-portions 24 being segments of image 18. Segments 24 subdivide image 18. As will be outlined in more detail below, the subdivision of image 18 into segments 24 may be subject to constraints such that each segment 24 may be completely contained within a single tile 70 or completely cover two or more tiles 70. ​ The example shows a case where the image 18 is subdivided into five segments 24. The first four of these segments 24 in the aforementioned decoding order cover the first two tiles 70, while the fifth segment completely covers the third and third tiles 70. In addition, ​ The case where each segment 24 is encoded in a separate payload packet 32 is exemplarily illustrated. Each tile identification packet 72 in turn indicates for the payload packet 32 immediately following it which of the tiles 70 the sub-portion 24 encoded in this payload packet 32 covers. Thus, while the first two tile identification packets 72 in the access unit 30 associated with the current picture 18 indicate the first tile, the third and fourth tile identification packets 72 indicate the second tile 70 of the picture 18, and the fifth tile identification packet 72 indicates the third and fourth tiles 70. ​ Embodiments, such as those described above with respect to ​ The same variations as described are possible. That is, the "scope" of a tile identification packet 72 may, for example, only include the first immediately following payload packet 32 or the immediately following payload packets 32 until the next tile identification packet occurs.

[0222] Regarding tiles, encoder 10 may be configured to encode each tile 70 such that no spatial prediction or context selection occurs across tile boundaries. Encoder 10 may, for example, encode tiles 70 in parallel. Similarly, any decoder, such as network entity 68, may decode tiles 70 in parallel.

[0223] The network entity 68 can be a MANE or a decoder or some other device between the encoder 10 and the decoder, and can be configured to use the information conveyed by the tile identification packet 72 to make decisions on specific transmission tasks. For example, the network entity 68 can process a specific tile of the current image 18 of the video 16 with a higher priority, that is, can forward the individual payload packets indicated to be related to this tile earlier or with more secure FEC protection or the like. In other words, the network entity 68 can use the recognition result to make decisions on the transmission tasks related to the video data stream. The transmission tasks can include retransmission requests for packets received in a defective state (i.e., in the case of any excessive FEC protection for the video data stream, if any). The network entity can, for example, process different tiles 70 with different priorities. To this end, compared with the tile identification packet 72 and its payload packets regarding lower-priority tiles, the network entity can assign a higher priority to the tile identification packet 72 and its payload packets related to higher-priority tiles (i.e., the payload packets prefixed with the tile identification packet 72). The network entity 68 can, for example, first request the retransmission of the payload packets assigned with a higher priority before requesting the retransmission of any payload packets assigned with a lower priority.

[0224] The embodiments described so far can be incorporated into the HEVC framework as described in the introductory part of the specification of this application, as described below.

[0225] Specifically, in the case of sub-picture CPB / HRD, the SEI message can be assigned to the segments of the decoding unit. That is, the buffering period and timing SEI message can be assigned to the NAL unit containing the segments of the decoding unit. This can be achieved by a new NAL unit type, which is a non-VCL NAL unit allowed to be immediately in front of one or more segments / VCL NAL units of the decoding unit. This new NAL unit can be called a segment prefix NAL unit. ​ Illustrates the structure of an access unit, which omits any hypothetical NAL units for the end of the sequence and the stream.

[0226] According to ​, the access unit 30 is understood as follows: In the sequential order of the packets of the packet sequence 34, the access unit may start from the occurrence of a special type of packet (i.e., the access unit delimiter 80). Then, within the access unit 30, one or more SEI packets 82 of one of the SEI packet types related to the entire access unit may follow. Both the packet types 80 and 82 are optional. That is, packets of this type may not appear in the access unit 30. Then, the sequence of decoding units 38 follows. Each decoding unit 38 optionally starts from a slice prefix NAL unit 84, which includes, for example, timing control information, or according to ​ or an embodiment of 16 includes ROI information or tile information, or even more generally includes individual sub - picture SEI messages 86. Then, the actual slice data 88 in individual payload packets or VCL NAL units follows, as indicated in 88. Thus, each decoding unit 38 contains a sequence of slice prefix NAL units 84, followed by individual slice data NAL units 88. ​ The skip arrow 90 that skips the slice prefix NAL unit in [reference] will indicate that, without subdivision of the decoding units of the current access unit 30, the slice prefix NAL unit 84 may not exist.

[0227] As already pointed out above, all information signaled in the slice prefix and related to the sub - picture SEI message may be valid for all VCL NAL units in the access unit or up to the occurrence of the second prefix NAL unit, or valid for subsequent VCL - NAL units in decoding order, depending on the flag given in the slice prefix NAL unit.

[0228] The slice VCL NAL units for which the information signaled in the slice prefix is valid are hereinafter referred to as prefix slices. A prefix slice related to a single slice of the prefix does not necessarily constitute a complete decoding unit, but may be part of a decoding unit. However, a single slice prefix cannot be valid for multiple decoding units (sub - pictures), and the start of a decoding unit is signaled by the occurrence of a slice prefix NAL unit. If the means for signaling is not given via the slice prefix syntax (as in the "simple syntax" / version 1 indicated below), the occurrence of a slice prefix NAL unit signals the start of a decoding unit. Only specific SEI messages (identified by payloadType in the following syntax description) can be uniquely sent within a slice prefix NAL unit at the sub - picture level, while some SEI messages can be sent within a slice prefix NAL unit at the sub - picture level or as regular SEI messages at the access unit level.

[0229] As above regarding ​As discussed, additionally or alternatively, tile ID SEI messages / tile ID signaling can be implemented in higher-level syntax. In an earlier design of HEVC, the slice header / slice data contained identifiers for the tiles contained in an individual slice. For example, the slice data semantics indicate that:

[0230] tile_idx_minus_1 specifies the TileID in raster scan order. The first tile in the picture will have a TileID of 0. The value of tile_idx_minus_1 will be in the range of 0 to (num_tile_columns_minus1 + 1)*(num_tile_rows_minus1 + 1) - 1.

[0231] However, this parameter is not considered useful because if tiles_or_entropy_coding_sync_idc equals 1, this ID can be easily derived from the slice address and slice size (signaled in the picture parameter set).

[0232] Although the tile ID can be implicitly derived during the decoding process, it is also important to know this parameter at the application layer for different usage scenarios, such as in a video conferencing situation where different tiles can have different playback priorities (these tiles usually form regions of interest that contain the speaker in a conversation scenario) and can have a higher priority than other tiles. In the case where network packets are lost during the transmission of multiple tiles, those network packets containing tiles representing regions of interest can be retransmitted with a higher priority in order to maintain a higher quality of experience at the receiver terminal than in the case of retransmitting without any priority order. Another usage scenario could be to assign tiles (if their size and position are known) to different screens, for example, in a video conferencing situation.

[0233] To allow this application layer to handle tiles with a specific priority in a transmission situation, the tile_id can be provided as a sub-picture or slice-specific SEI message, or in a special NAL unit in front of one or more NAL units of a tile or in a special header section of the NAL units belonging to a tile.

[0234] As described above regarding ​ Additionally or alternatively, a region of interest SEI message can also be provided. This SEI message can allow the signaling of a region of interest (ROI), specifically, the ROI to which a specific tile_id / tile belongs. This message can allow the giving of a region of interest ID plus a region of interest priority.

[0235] ​ Illustrates the use of tiles in region of interest signaling.

[0236] In addition to what is described above, slice header signaling may be implemented. A slice prefix NAL unit may also contain subsequent dependent slices (i.e., the slice headers of slices prefixed by individual slice prefixes). If the slice header is only provided in the slice prefix NAL unit, the actual slice type needs to be derived from the NAL unit type of the NAL unit containing the individual dependent slices or by a flag in the slice prefix, which signals whether the subsequent slice data belongs to a slice type that serves as a random access point.

[0237] In addition, a slice prefix NAL unit may carry slice- or sub-picture-specific SEI messages to convey non-mandatory information, such as sub-picture timing or tile identifiers. Non-mandatory sub-picture-specific message passing is not supported in the HEVC specification described in the introductory part of the specification of this application, but it is crucial for some applications.

[0238] The techniques for implementing the tile prefixing concept outlined above are described below. Specifically, it is described what changes may be sufficient at the tile level when using the HEVC state as outlined in the introductory part of the specification of this application as a basis.

[0239] Specifically, two versions of a possible slice prefix syntax are presented below, one version having only SEI message passing functionality and one version having extended functionality for signaling a part of the slice header of subsequent slices. ​ The first simple syntax / version 1 is shown.

[0240] As a preliminary note, ​ Thus, possible implementations for implementing any of the embodiments described above with respect to ​ are shown. The interleaved packets shown therein can be understood as shown in ​ and the above is described in more detail below using specific implementation examples.

[0241] In ​ a table, the extended syntax / version 2 is given, which includes tile_id signaling, decoding unit start identifier, slice prefix ID, and slice header data in addition to the SEI message concept techniques.

[0242] The semantics can be defined as follows:

[0243] A rap_flag with a value of 1 indicates that the access unit containing the slice prefix is a RAP picture. A rap_flag with a value of 0 indicates that the access unit containing the slice prefix is not a RAP picture.

[0244] The decoding_unit_start_flag indicates the start of a decoding unit within an access unit, and thus indicates that subsequent slices until the end of the access unit or the start of another decoding unit belong to the same decoding unit.

[0245] A single_slice_flag with a value of 0 indicates that the information provided in the prefix slice NAL unit and the associated sub-picture SEI message is valid for all subsequent VCL-NAL units until the start of the next access unit, the appearance of another slice prefix, or another complete slice header. A single_slice_flag with a value of 1 indicates that the information provided in the slice prefix NAL unit and the associated sub-picture SEI message is valid only for the next VCL-NAL unit in decoding order.

[0246] The tile_idc indicates the amount of tiles that will be present in the subsequent slice. A tile_idc equal to 0 indicates that no tiles are used in the subsequent slice. A tile_idc equal to 1 indicates that a single tile is used in the subsequent slice, and its tile identifier is signaled accordingly. A tile_idc with a value of 2 indicates that multiple tiles are used in the subsequent slice, and the number of tiles and the first tile identifier are signaled accordingly.

[0247] The prefix_slice_header_data_present_flag indicates that data corresponding to the slice header for the slice following in decoding order is signaled in the given slice prefix.

[0248] slice_header_data() is defined later in the text. It contains the relevant slice header information that is not covered by the slice header if the dependent_slice_flag is set equal to 1.

[0249] Note that decoupling the slice header from the actual slice data allows for a more flexible transmission scheme for the header and the slice data.

[0250] num_tiles_in_prefixed_slices_minus1 indicates the number of tiles used in the subsequent decoding unit minus 1.

[0251] first_tile_id_in_prefixed_slices indicates the tile identifier of the first tile in the subsequent decoding unit.

[0252] For the simple syntax / version 1 of the slice prefix, the following syntax elements can be set to the following default values (if not present):

[0253] decoding_unit_start is equal to 1, i.e., the slice prefix always indicates the start of a decoding unit.

[0254] single_slice_flag is equal to 0, ie the slice prefix is valid for all tiles in the decoding unit.

[0255] It is proposed that the segment prefix NAL unit have NAL unit type 24 and the NAL unit type summary table be based on ​ To expand.

[0256] That is, briefly summarize ​ , where the syntactic details shown indicate that a particular packet type may be considered to belong to the interleaved packets identified above, exemplarily here NAL unit type 24. In addition, in particular ​ The syntax example of shows that with respect to the above alternatives for the "category" of interleaved packets, a switching mechanism controlled by a separate syntax element (here exemplarily single_slice_flag) within these interleaved packets themselves can be used to control this category, i.e. to switch between the different alternatives defined for this category. Furthermore, it has been shown that the extensible ​ In the above embodiment, since the interspersed packets also include the common segment header data of the segments 24 contained in the packets belonging to the "scope" of the individual interspersed packets, there can be a mechanism controlled by a separate flag in these interspersed packets that indicates whether the individual interspersed packets contain the common segment header data.

[0257] Of course, the conceptual technique just presented, according to which part of the slice header data is transformed into a slice header prefix, requires changes to the slice header as specified in the current version of HEVC. ​ The table in shows a possible syntax for this slice header, where specific syntax elements present in the slice header according to the current version are transformed into a lower-level syntax element called slice_header_data(). This syntax of the slice header and the slice header data only applies to the following option, according to which the extended slice header prefix NAL unit concept technique is used.

[0258] ​ In the slice_header_data_present_flag, slice_header_data_present_flag indicates that the slice header data of the current slice is to be predicted from the value signaled in the last slice prefix NAL unit (ie, the most recently occurring slice prefix NAL unit) in the access unit.

[0259] Through ​ 、 ​ and ​ The syntax elements given in the table of The slice header data signals all syntax elements removed from the slice header.

[0260] That is, transfer the concepts of ​ and ​ , ​ and ​ to the embodiments of ​ . The interspersed packets described therein can be extended by the concept technology, which incorporates a part of the slice header syntax of the slice (sub-part) 24 encoded into the payload packet (i.e., VCL NAL unit) into these interspersed packets. This incorporation can be optional. That is, individual syntax elements in the interspersed packets can indicate whether this slice header syntax is contained in an individual interspersed packet. If incorporated, the individual slice header data incorporated into an individual interspersed packet can apply to all slices contained in the packets belonging to the "scope" of the individual interspersed packet. It can be signaled by an individual flag (such as ​ 's slice_header_data_present_flag) whether the slice encoded into any one of the payload packets belonging to the scope of the interspersed packet uses the slice header data contained in this interspersed packet. By this measure, the slice header of the slice encoded into the packets belonging to the "scope" of an individual interspersed packet can be correspondingly reduced using the aforementioned flag in the slice header, and any decoder receiving the video data stream (such as the network entity shown above in ​ ) will respond to the aforementioned flag in the slice header so as to copy the slice header data incorporated into the interspersed packet into the slice header of the slice encoded into the payload packet belonging to the scope of this interspersed packet in the case where the individual flag in the slice signals the movement of the slice header data to the slice prefix (i.e., the individual interspersed packet).

[0261] Continuing further with the syntax examples for implementing the embodiments of ​ , the SEI message syntax can be as shown in ​ . To introduce the slice or sub-picture SEI message type, the SEI payload syntax can be adapted as presented in the tables of ​ and ​ . Only SEI messages with payloadType in the range of 180 to 184 can be uniquely sent in the slice prefix NAL unit at the sub-picture level. Additionally, the Region of Interest SEI message with payloadType equal to 140 can be sent in the slice prefix NAL unit at the sub-picture level, or as a regular SEI message at the access unit level.

[0262] That is, in transferring the details shown in ​ and ​ and ​ to the above regarding ​When implemented on the described embodiments, it can be achieved by using a slice prefix NAL unit with a specific NAL unit type (e.g., 24). ​ The interspersed packets shown in these embodiments, including, for example, specific types of SEI messages signaled by payloadType at the beginning of each SEI message within the slice prefix NAL unit. In the specific syntax embodiments described now, payloadType = 180 and payloadType = 181 result in timing control packets according to ​ the embodiments of, while payloadType = 140 results in ROI packets according to ​ the embodiments of, and payloadType = 182 results in tile identification packets according to ​ the embodiments of. The specific syntax examples described herein below may contain only one or a subset of the payloadType options just mentioned. In addition, ​ and ​ indicate that ​ any of the above embodiments can be combined with each other. Further, ​ and ​ indicate that ​ any of the above embodiments or any combination thereof can be extended by another interspersed packet (subsequently explained by payloadType = 184). As already described above, the extension described below regarding payloadType = 183 ultimately results in the possibility that any interspersed packet may have incorporated therein common slice header data of the slice headers encoded into any payload packet within its scope.

[0263] The table definitions in the subsequent figures define SEI messages that can be used at the slice or sub - image level. Region of interest SEI messages that can be used at the sub - image and access unit levels are also presented.

[0264] For example, ​ an example of a sub - image buffer SEI message that appears as long as the slice prefix NAL unit of NAL unit type 24 has an SEI message type 180 included therein (thus forming a timing control packet) is shown.

[0265] The semantics can be defined as follows:

[0266] The seq_parameter_set_id specifies the sequence parameter set containing the sequence HRD properties. The value of seq_parameter_set_id shall be equal to the value of seq_parameter_set_id in the picture parameter set referred to by the primary coded picture associated with the buffering period SEI message. The value of seq_parameter_set_id shall be in the range of 0 to 31, inclusive of 0 and 31.

[0267] initial_cpb_removal_delay[SchedSelIdx] and initial_alt_cpb_removal_delay[SchedSelIdx] specify the initial CPB removal delay of the SchedSelIdx-th CPB of a decoding unit (slice). These syntax elements have a length in bits given by initial_cpb_removal_delay_length_minus1 + 1, and are in units of the 90 kHz clock. The values of these syntax elements shall not be equal to 0 and shall not exceed 90000 * (CpbSize[SchedSelIdx] ÷ BitRate[SchedSelIdx]), which is the time equivalent of the CPB size in units of the 90 kHz clock.

[0268] Over the entire coded video sequence, the sum of initial_cpb_removal_delay[SchedSelIdx] and initial_cpb_removal_delay_offset[SchedSelIdx], and the sum of initial_alt_cpb_removal_delay[SchedSelIdx] and initial_alt_cpb_removal_delay_offset[SchedSelIdx] for each decoding unit (slice) shall be constant for each value of SchedSelIdx.

[0269] ​ An example of a slice timing SEI message is also shown, the semantics of which can be described as follows:

[0270] The du_cpb_removal_delay specifies how many clock ticks to wait after removing the cpb of a decoding unit (sub - picture) related to the most recent sub - picture buffer cycle SEI message in the previous access unit in the same decoding unit (sub - picture) (if any), or otherwise related to the most recent buffer cycle SEI message in the previous access unit, before removing the decoding unit (sub - picture) data related to the sub - picture timing SEI message from the buffer. This value is also used to calculate the earliest possible time for the decoding unit (sub - picture) data to arrive at the CPB for the HSS (Hypothetical Scheduler [2]0). This syntax element is a fixed - length code, and its length in bits is given by cpb_removal_delay_length_minus1 + 1. cpb_removal_delay is the remainder of modulo 2 (cpb_removal_delay_length_minus1+1) counter.

[0271] The du_dpb_output_delay is used to calculate the dpb output time of a decoding unit (sub - picture). It specifies how many clock ticks to wait after removing the decoded decoding unit (sub - picture) from the CPB and before outputting the decoding unit (sub - picture) of the image from the DPB.

[0272] Note that this allows sub - picture updates. In this case, the non - updated decoding units can keep the last decoded image unchanged, i.e., they remain visible.

[0273] Summary ​ and ​ transferring the specific details contained therein to the ​ example, it can be said that the decoder buffer fetch time of a decoding unit can be signaled in a differentially - coded manner (i.e., incrementally) relative to another decoder buffer fetch time in the relevant timing control packet. That is, to obtain the decoder buffer fetch time of a specific decoding unit, the decoder receiving the video data stream adds the decoder fetch time obtained from the timing control packet prefixed to that specific decoding unit to the decoder fetch time of the immediately preceding decoding unit (i.e., the decoding unit preceding that specific decoding unit), and continues in this way for subsequent decoding units. At the start of an encoded video sequence of several pictures (each or parts thereof), the timing control packet may additionally or alternatively contain decoder buffer fetch time values that are encoded independently rather than differentially relative to the decoder buffer fetch time of any previous decoding unit.

[0274] ​ Shows what the sub - picture segment information SEI message might look like. The semantics can be defined as follows:

[0275] A slice_header_data_flag value of 1 indicates that slice header data is present in the SEI message. The slice header data provided in the SEI is valid for all slices that follow in decoding order until the end of the access unit, another SEI message, a slice NAL unit, or the occurrence of slice data in a slice prefix NAL unit.

[0276] ​ An example of an SEI message showing sub-picture tile information is presented, where the semantics can be defined as follows:

[0277] tile_priority indicates the priority of all tiles in the prefix slice that follows in decoding order. The value of tile_priority shall be in the range of 0 to 7 (inclusive of 0 and 7), where 7 indicates the highest priority.

[0278] A multiple_tiles_in_prefixed_slices_flag value of 1 indicates that there is more than one tile in the prefix slice that follows in decoding order. A multiple_tiles_in_prefixed_slices_flag value of 0 indicates that the subsequent prefix slice contains only one tile.

[0279] num_tiles_in_prefixed_slices_minus1 indicates the number of tiles in the prefix slice that follows in decoding order.

[0280] first_tile_id_in_prefixed_slices indicates the tile_id of the first tile in the prefix slice that follows in decoding order.

[0281] That is, using ​ the syntax to implement ​ the tile identification packet mentioned in ​ the embodiments can be implemented. As ​As shown, a specific flag (here, the multiple_tiles_in_prefixed_slices_flag) can be used to signal in an interlaced tile identification packet whether any sub - part of the current picture 18 encoded in any one of the payload packets belonging to the scope of an individual interlaced tile identification packet covers only one tile or more than one tile. If the flag signals coverage of more than one tile, another syntax element is present in the individual interlaced packet, here exemplarily the num_tiles_in_prefixed_slices_minus1, which indicates the number of tiles covered by any sub - part of any payload packet belonging to the scope of an individual interlaced tile identification packet. Finally, another syntax element (here exemplarily the first_tile_id_in_prefixed_slices) indicates the ID of the tile that is the first in decoding order among the number of tiles indicated by the current interlaced tile identification packet. ​ Transferring the ​ syntax to the embodiment of, the tile identification packet 72 pre - fixed to the fifth payload packet 32 may, for example, have all three of the syntax elements just discussed, where the multiple_tiles_in_prefixed_slices_flag is set to 1, the num_tiles_in_prefixed_slices_minus1 is set to 1, thereby indicating that two tiles belong to the current scope, and the first_tile_id_in_prefixed_slices is set to 3, which indicates that the run of tiles in decoding order belonging to the scope of the current tile identification packet 72 starts at the third tile (with tile_id = 2).

[0282] ​ It also shows that the tile identification packet 72 may also be able to indicate tile_priority, that is, the priority of the tiles belonging to its scope. Similar to the ROI aspect, the network entity 68 can use this priority information to control transmission tasks, such as requests for re - transmission of specific payload packets.

[0283] ​ Shows a syntax example of a sub - picture tile size information SEI message, where the semantics can be defined as follows:

[0284] A multiple_tiles_in_prefixed_slices_flag with a value of 1 indicates that there is more than one tile in the subsequent prefix segment in decoding order. A multiple_tiles_in_prefixed_slices_flag with a value of 0 indicates that the subsequent prefix segment contains only one tile.

[0285] num_tiles_in_prefixed_slices_minus1 indicates the number of tiles in the prefixed slices that follow in decoding order.

[0286] tile_horz_start[i] indicates the horizontal start of the i-th tile among the pixels in the image.

[0287] tile_width[i] indicates the width of the i-th tile among the pixels in the image.

[0288] tile_vert_start[i] indicates the horizontal start of the i-th tile among the pixels in the image.

[0289] tile_height[i] indicates the height of the i-th tile among the pixels in the image.

[0290] Note that size SEI messages can be used in the display operation, for example, to assign tiles to screens in a multi-screen display situation.

[0291] ​ Thus, it is shown that the implementation syntax example of the ​ tile identification packet can be changed because the tiles belonging to the scope of an individual tile identification packet are indicated by their position in the current image 18 rather than their tile ID. That is, instead of signaling the tile ID of the first tile in decoding order covered by an individual sub-part (which is encoded into any one of the payload packets belonging to the scope of an individual interleaved tile identification packet), for each tile belonging to the current tile identification packet, the position of the tile can be signaled by signaling, for example, the following: the top-left position of each tile i (demonstratively by tile_horz_start and tile_vert_start here), and the width and height of tile i (demonstratively by tile_width and tile_height here). ​

[0292] A syntax example of the region of interest SEI message is shown in ​ . For even more precision, ​ shows a first variant. Specifically, the region of interest SEI message can be used to signal one or more regions of interest, for example, at the access unit level or at the sub-image level. According to ​ 's first variant, each ROI SEI message signals an individual ROI once, rather than signaling all ROIs within the scope of an individual ROI packet (if multiple ROIs are within the current scope) within one ROI SEI message.

[0293] ​According to ​ , the SEI messages in the region of interest respectively signal each ROI. The semantics can be defined as follows:

[0294] roi_id indicates the identifier of the region of interest.

[0295] roi_priority indicates the priority of all tiles belonging to the region of interest in the subsequent prefix segment or all segments in the decoding order, which depends on whether the SEI message is sent at the sub-image level or the access unit level. The value of roi_priority will be in the range of 0 to 7 (including 0 and 7), where 7 indicates the highest priority. If both ROI_priority in the ROI information SEI message and tile_priority in the sub-image tile information SEI message are given, the higher value of the two is valid for the priority of individual tiles.

[0296] num_tiles_in_roi_minus1 indicates the number of tiles belonging to the region of interest in the subsequent prefix segment in the decoding order.

[0297] roi_tile_id[i] indicates the tile_id of the i-th tile belonging to the region of interest in the subsequent prefix segment in the decoding order.

[0298] That is,[[]] ​ shows that the ROI packet as shown in ​ can signal the ID of the region of interest therein, and the individual ROI packet and the payload packet belonging to its scope are about the region of interest. Optionally, the ROI priority index can be signaled together with the ROI ID. However, both syntax elements are optional. Then, the syntax element num_tiles_in_roi_minus1 can indicate the number of tiles belonging to the individual ROI 60 within the scope of the individual ROI packet. Then, roi_tile_id indicates the tile-ID of the i-th tile belonging to ROI 60. For example, imagine that image 18 will be divided into tiles 70 in the manner shown in ​ , and ​ ROI 60 will correspond to the left half of image 18 and will be formed by the first and third tiles in the decoding order. Then, the ROI packet in ​The first payload packet 32 of the access unit 30 may be preceded by another ROI packet, followed by another ROI packet between the fourth and fifth payload packets 32 of the access unit 30. The first ROI packet would then have num_tile_in_roi_minus1 set to 0 and roi_tile_id[0] to 0 (thus relating to the first tile in decoding order), wherein the second ROI packet preceding the fifth payload packet 32 would have num_tile_in_roi_minus1 set to 0 and roi_tile_id[0] to 2 (thus representing the third tile in decoding order, which is located in the lower left quarter of the picture 18).

[0299] According to the second variant, the syntax of the attention area SEI message can be as follows ​ Here, all ROIs are signaled in a single SEI message. In detail, the same ​ The same syntax is discussed, but the syntax elements for each of several ROIs (for which individual ROI SEI messages or ROI packets pertain) are multiplied by a number signaled by a syntax element (here exemplarily num_rois_minus1). Optionally, another syntax element (here exemplarily roi_presentation_on_separate_screen) can signal for each ROI whether the individual ROI is suitable for presentation on a separate screen.

[0300] The semantics can be as follows:

[0301] num_rois_minus1 indicates the number of ROIs in the following prefix or normal slice in decoding order.

[0302] roi_id[i] indicates the identifier of the i-th region of interest.

[0303] roi_priority[i] indicates the priority of all tiles belonging to the i-th region of interest in the preceding or all following slices in decoding order, depending on whether the SEI message is sent at the sub-picture level or the access unit level. The value of roi_priority shall be in the range of 0 to 7 (inclusive), with 7 indicating the highest priority. If both roi_priority in the roi_info SEI message and tile_priority in the sub-picture tile information SEI message are given, the highest of the two values is valid for the priority of an individual tile.

[0304] num_tiles_in_roi_minus1[i] indicates the number of tiles belonging to the i-th region of interest in the prefix slice that follows in decoding order.

[0305] roi_tile_id[i][n] indicates the tile_id of the n-th tile belonging to the i-th region of interest in the prefix slice that follows in decoding order.

[0306] roi_presentation_on_seperate_screen[i] indicates that the region of interest associated with the i-th roi_id is suitable for presentation on a separate screen.

[0307] Thus, to briefly summarize the various embodiments described so far, additional high-level syntax signaling strategies have been proposed that allow the application of SEI messages as well as additional high-level syntax items (beyond those included in the NAL unit header at the per-slice level). Thus, slice prefix NAL units have been described. Along with their usage for low-latency / sub-picture CPB operation, tile signaling, and ROI signaling, the syntax and semantics of slice prefix and slice-level / sub-picture SEI messages have been described. Extended syntax has been presented to additionally signal parts of the slice header of subsequent slices in the slice prefix.

[0308] For completeness, ​ Another example of the syntax is shown, which can be used for the timing control packet according to the ​ embodiment. The semantics can be:

[0309] du_spt_cpb_removal_delay_increment specifies the duration, in clock sub-ticks, between the nominal CPB time of the last decoded unit in decoding order in the current access unit and the nominal CPB time of the decoded unit associated with the decoded unit information SEI message. This value is also used to calculate the earliest possible time for the decoded unit data to arrive at the CPB for HSS. This syntax element is represented by a fixed-length code, the length of which in bits is given by du_cpb_removal_delay_increment_length_minus1 + 1. When the decoded unit associated with the decoded unit information SEI message is the last decoded unit in the current access unit, the value of du_spt_cpb_removal_delay_increment will be equal to 0.

[0310] The dpb_output_du_delay_present_flag being equal to 1 specifies the presence of the pic_spt_dpb_output_du_delay syntax element in the decoded unit information SEI message. The dpb_output_du_delay_present_flag being equal to 0 specifies the absence of the pic_spt_dpb_output_du_delay syntax element in the decoded unit information SEI message.

[0311] pic_spt_dpb_output_du_delay is used to calculate the dpb output time of the picture when SubPicHrdFlag is equal to 1. It specifies how many sub-clock ticks to wait after the last decoded unit in the access unit has been removed from the CPB and before the decoded picture is output from the DPB. When it is absent, it is inferred that the value of pic_spt_dpb_output_du_delay is equal to pic_dpb_output_du_delay. The length of the syntax element pic_spt_dpb_output_du_delay is given in bits by dpb_output_delay_du_length_minus1 + 1.

[0312] The bitstream conformance requirement is that all decoded unit information SEI messages related to the same access unit, applicable to the same operation point, and having dpb_output_du_delay_present_flag equal to 1 will have the same value of pic_spt_dpb_output_du_delay. The output time derived from pic_spt_dpb_output_du_delay of any picture output from a conforming output timing decoder will be before the output times derived from pic_spt_dpb_output_du_delay of all pictures in any subsequent CVS in decoding order.

[0313] The picture output order determined by the value of this syntax element will be the same as the order determined by the value of PicOrderCntVal.

[0314] For pictures that are not output by the "collision" process (since such pictures are in decoding order before IRAP pictures where no_output_of_prior_pics_flag equals 1 or is inferred to equal 1 and NoRaslOutputFlag equals 1), the output time derived from pic_spt_dpb_output_du_delay will increase as the PicOrderCntVal value increases for all pictures within the same CVS. For any two pictures in the CVS, the difference between the output times of the two pictures when SubPicHrdFlag equals 1 will be equal to the same difference when SubPicHrdFlag equals 0.

[0315] In addition, ​ Another example of signaling an ROI region using an ROI packet is shown. According to ​ , the syntax of the ROI packet contains only one flag that indicates whether all sub-parts of picture 18 encoded into any valid payload packet 32 within its scope belong to the ROI. The "scope" extends until the ROI packet or the region_refresh_info SEI message appears. If the flag is 1, it indicates that the region is encoded into individual subsequent payload packets, and if it is 0, the opposite is true, i.e., the individual sub-parts of picture 18 do not belong to the ROI 60.

[0316] Before discussing some of the above embodiments again (in other words, further explaining some of the terms used above, such as tiles, slices, and WPP sub-stream subdivisions), it should be noted that the high-level signaling of the above embodiments can alternatively be defined in transmission specifications such as [3-7]. In other words, the packets mentioned above and forming sequence 34 can be transmission packets, some of which have sub-parts (such as slices) of the application layer incorporated (such as fully encapsulated or split) into them, and some are interleaved between the latter in the manner and for the purposes discussed above. In other words, the interleaved packets mentioned above are not limited to SEI messages that are defined as other types of NAL units in a video codec decoder of the application layer, but can alternatively be additional transmission packets defined in the transmission protocol.

[0317] In other words, according to one aspect of the present specification, the above embodiment discloses a video data stream having video content encoded therein in units of sub-portions of pictures of the video content (see coding tree blocks or slices), each sub-portion being respectively encoded in one or more payload packets (see VCL NAL units) of a sequence of packets (NAL units) of the video data stream, the sequence of packets being divided into a sequence of access units, such that each access unit collects payload packets associated with a respective picture of the video content, wherein the sequence of packets has timing control packets (slice prefixes) interspersed therein, such that the timing control packets subdivide the access units into decoding units, such that at least some access units are subdivided into two or more decoding units, wherein each timing control packet signals a decoder buffer capture time of a decoding unit, and its payload packet follows the respective timing control packet in the sequence of packets.

[0318] As described above, the field related to encoding video content into a data stream in units of sub-portions of an image may include syntax elements related to predictive coding, such as coding modes (such as intra mode, inter mode, subdivision information, and the like), prediction parameters (such as motion vectors, extrapolation directions, and the like), and / or residual data (such as transform coefficient levels), wherein these syntax elements are respectively related to local portions of an image (such as coding tree blocks, prediction blocks, and residual (such as transform) blocks).

[0319] As described above, payload packets may each contain one or more (individually complete) segments. Segments may be independently decodable or may exhibit relationships that prevent independent decoding. For example, an entropy segment may be independently entropy decodable, but prediction across segment boundaries may be prevented. Dependent segments may allow for WPP processing, i.e., encoding / decoding using entropy and prediction coding across segment boundaries, with the ability to encode / decode dependent segments in parallel with temporal overlap, but with staggered start of the encoding / decoding process for each dependent segment and the segment(s) to which it refers.

[0320] The decoder may know in advance the sequential order in which the payload packets of an access unit are arranged within a respective access unit. For example, the encoding / decoding order may be defined between sub-parts of a picture, such as the scanning order between coding tree blocks in the above example.

[0321] For example, see the following figure. The image 100 currently being encoded / decoded can be divided into tiles, such as ​ and ​ exemplarily corresponds to four quadrants of the image 110 and is indicated by reference symbols 112a-112d. That is, the entire image 110 may form one tile, as shown in FIG. ​In some cases, it may be segmented into more than one tile. Tile segmentation may be limited to regular segmentation, where tiles are arranged only by rows and columns. Different examples are presented below.

[0322] As can be seen, image 110 is further subdivided into coded (tree) blocks (small squares in the figure, and previously referred to as CTBs) 114, and a coding order 116 is defined between these blocks (here it is a raster scan order, but it can also be different). Subdividing the image into tiles 112a-d can be restricted such that the tiles are a disjoint set of blocks 114. Additionally, both the blocks 114 and the tiles 112a-d can be limited to a regular arrangement by rows and columns.

[0323] If there are tiles (i.e., more than one), the coding (decoding) order 116 first performs a raster scan on the first complete tile, and then transitions to the next tile in tile order also in raster scan tile order.

[0324] Since the tiles can be encoded / decoded independently of each other due to non-intersecting tile boundaries, and this encoding / decoding is by spatial prediction and context selection inferred from the spatial neighborhood, the encoder 10 and decoder 12 can encode / decode the image subdivided into tiles 112 (previously indicated by 70) independently and in parallel of each other, except for example in-loop or post-filtering (which may be allowed to intersect the tile boundaries).

[0325] Image 110 can be further subdivided into segments 118a-d, 180 (previously indicated using reference symbol 24). A segment can contain only a part of a tile, a complete tile, or more than one complete tile. Thus, segmentation into segments can also be a subdivision of tiles, as in ​ In some cases. Each segment contains at least one complete coded block 114 and consists of consecutive coded blocks 114 in coding order 116, so an order is defined between segments 118a-d, and the indices in the figure are assigned following this order. The segment segmentation in ​ is chosen only for illustrative purposes. The tile boundaries can be signaled in the data stream. Image 110 can form a single tile, as depicted in ​ as shown.

[0326] The encoder 10 and decoder 12 can be configured to respect the tile boundaries because spatial prediction is not applied across tile boundaries. Context adaptation (i.e., probability adaptation) of various entropy (arithmetic) contexts can continue over the entire tile. However, as long as a segment intersects a tile boundary (if there is one within the segment) along the coding order 116, such as in ​Regarding segments 118a, 118b, the segments are further divided into sub - segments (sub - streams or tiles), where the segment contains an index (i.e., entry_point_offset) pointing to the start of each sub - segment. In the decoder loop, filters are allowed to intersect tile boundaries. Such filters can involve one or more of the de - blocking filter, sample adaptive offset (SAO) filter, and adaptive loop filter (ALF). The latter can be applied at the tile / segment boundary (if enabled).

[0327] Each optional second and subsequent sub - segment can have its start byte - aligned within the segment, where the index indicates the offset from the start of one sub - segment to the start of the next sub - segment. These sub - segments are arranged within the segment in scan order 116. ​ Shows ​ an example where segment 180c is divided into sub - segments 119 i .

[0328] Regarding the figures, note that the tiles forming the sub - parts of a segment do not have to end with the columns in tile 112a. For example, see ​ and ​ segment 118a in.

[0329] The following figure shows an exemplary part of the data stream regarding the access unit associated with the image 110 above ​ Here, each payload packet 122a - d (previously indicated by reference symbol 32) is exemplary applicable to only one segment 118a. For illustrative purposes, two timing control packets 124a, 124b (previously indicated by reference symbol 36) are shown interleaved in the access unit 120: 124a is before packet 122a in packet order 126 (corresponding to the decode / encode time axis), and 124b is before packet 122c. Thus, the access unit 120 is divided into two decoding units 128a, 128b (previously indicated by reference symbol 38), where the first decoding unit contains packets 122a, 122b (and optional padding data packets (after the first packet 122a and the second packet 122b respectively) and an optional access unit leading SEI packet (before the first packet 122a)), and the second decoding unit contains packets 118c, 118d (and optional padding data (after packets 122c, 122d respectively)).

[0330] As described above, each packet of a packet sequence can be assigned only one packet type (nal_unit_type) out of a plurality of packet types. Payload packets and timing control packets (and optional filler data and SEI packets) are, for example, different packet types. The instantiation of packets of a particular packet type in a packet sequence can be subject to particular restrictions. These restrictions can define an order between the packet types (see ​ ), and the packets within each access unit will adhere to this order such that the access unit boundaries 130a, 130b can be detected and remain in the same position within the packet sequence even if packets of any removable packet type are removed from the video data stream. For example, payload packets are non-removable packet types. However, timing control packets, filler data packets, and SEI packets can be removable packet types as discussed above, i.e., these packets can be non-VCL NAL units.

[0331] In the above example, the timing control packets have been explicitly illustrated by the syntax of slice_prefix_rbsp().

[0332] Using this interleaving of timing control packets enables the encoder to adjust the buffer scheduling at the decoder side during the encoding of individual images of the video content. For example, it enables the encoder to optimize the buffer scheduling to minimize the end-to-end delay. In this regard, the encoder can take into account the individual distribution of the image regions of the video content across the individual images of the video content for the encoding complexity. Specifically, the encoder can continuously output a sequence of packets 122, 122a-d, 122a-d 1-3 on a packet-by-packet basis (i.e., output the current packet as soon as it is finished). By using timing control packets, the encoder can adjust the buffer scheduling at the decoding side at a moment when some of the sub-parts in the current image have been encoded into individual payload packets, but the remaining sub-parts have not yet been encoded.

[0333] Accordingly, an encoder for encoding video content into a video data stream in units of sub-parts of images of the video content (see coding tree blocks, tiles, or segments), where each sub-part is encoded into one or more payload packets (see VCL NAL units) of a sequence of packets (NAL units) of the video data stream, such that the packet sequence is divided into a sequence of access units and each access unit collects the payload packets related to an individual image of the video content, can be configured to interleave timing control packets (slice prefixes) in the packet sequence, such that the timing control packets divide the access units into decoding units, such that at least some access units are divided into two or more decoding units, where each timing control packet signals the decoder buffer fetch time of the decoding unit, and its payload packets follow the individual timing control packets in the packet sequence.

[0334] Any decoder that receives the video data stream outlined above may or may not utilize the scheduling information contained in the timing control packet. However, while the decoder may utilize this information, a decoder compliant with the codec layer must be able to decode data that follows the indicated timing. If utilization occurs, the decoder feeds its decoder buffer in units of decoding units and empties its decoder buffer. As described above, the "decoder buffer" may refer to a decoded picture buffer and / or an encoded picture buffer.

[0335] Accordingly, a decoder for decoding a video data stream (having video content encoded therein in units of sub-parts of images of the video content (see coding tree blocks, tiles or slices), where each sub-part is encoded into one or more payload packets (see VCL NAL units) of a sequence of packets (NAL units) of the video data stream, so that the sequence of packets is divided into a sequence of access units and each access unit collects the payload packets related to an individual picture of the video content) may be configured to look for timing control packets interspersed in the sequence of packets, at which the access units are subdivided into decoding units, so that at least some access units are subdivided into two or more decoding units, derive the decoder buffer fetch time of the decoding units from each timing control packet, the payload packets of which are behind the individual timing control packet in the sequence of packets, and fetch the decoding units from the buffer of the decoder, this being scheduled at the time defined by the decoder buffer fetch time of the decoding units.

[0336] Looking for timing control packets may involve the decoder examining the NAL unit header and the syntax elements contained therein, i.e., nal_unit_type. If the value of the latter flag is equal to a certain value, i.e., 124 according to the above example, the currently examined packet is a timing control packet. That is, the timing control packet may contain or convey the information explained above regarding subpic_buffering and subpic_timing. That is, the timing control packet may convey or specify the initial CPB removal delay of the decoder, or specify how many clock ticks to wait after removing an individual decoding unit from the CPB.

[0337] To allow for the repeated transmission of timing control packets without inadvertently further subdividing the access units into other decoding units, a flag within the timing control packet may explicitly signal whether the current timing control packet participates in subdividing the access units into coding units (compare decoding_unit_start_flag = 1 indicating the start of a decoding unit and decoding_unit_start_flag = 0 signaling the opposite).

[0338] The manner of using the tile identification information related to the interleave decoding unit is different from the manner of using the timing control packet related to the interleave decoding unit in that the tile identification information is interleaved in the data stream. The above-mentioned timing control packet can be additionally interleaved in the data stream, or the decoder buffer fetch time can be conveyed together with the tile identification information explained below in the same packet. Therefore, the details presented in the above sections can be used to clarify the issues in the following description.

[0339] Another aspect of the present specification derivable from the above embodiments discloses a video data stream having video content encoded therein using prediction and entropy coding in units of segments into which an image of the video content is spatially subdivided, wherein the prediction coding and / or entropy coding prediction is limited within tiles into which an image of the video content is spatially subdivided, wherein a sequence of segments in coding order is packetized into payload packets of a sequence of packets (NAL units) of the video data stream, the packet sequence being divided into a sequence of access units, so that each access unit collects the payload packets related to an individual image of the video content, the payload packets having segments packetized therein, wherein the packet sequence has tile identification packets interleaved therein, the tile identification packets identifying the tile (possibly only one) covered by the segment (possibly only one), the segment being packetized into one or more payload packets immediately following the individual tile identification packet in the packet sequence.

[0340] For example, refer to the figure showing the data stream immediately preceding. Packets 124a and 124b will now represent the tile identification packets. By explicit signaling (compare single_slice_flag = 1) or by convention, the tile identification packet can identify only the tile covered by the segment packetized into the immediately following payload packet 122a. Alternatively, by explicit signaling or by convention, the tile identification packet 124a can identify the tile covered by the segments packetized into one or more payload packets immediately following the individual tile identification packet 124a in the packet sequence up to the earlier of the end 130b of the current access unit 120 and the start of the next decoding unit 128b, respectively. For example, refer to ​ , if each segment 118a-d 1-3 is separately packetized into individual packets 122a-d 1-3 wherein the subdivision into decoding units is such that the packets are grouped into three decoding units according to {122a 1-3},{122b 1-3} and {122c 1-3 ,122d 1-3}, then the packets packetized into the third decoding unit {122c 1-3,122d 1-3} in the fragment {118c 1-3 ,118d 1-3} would, for example, cover tiles 112c and 112d, and the corresponding slice prefix would (for example when referring to a complete decoding unit) indicate "c" and "d", ie these tiles 112c and 112d.

[0341] Therefore, network entities, as described further below, can use this explicit signaling or convention to correctly correlate each tile identification packet with the one or more payload packets that immediately follow it in the packet sequence. The manner in which this identification can be signaled has been described above by way of example using the pseudo code subpic_tile_info. The associated payload packets are referred to above as "prefix segments." This example can be modified in nature. For example, the syntax element "tile_priority" can be omitted. Furthermore, the order of the syntax elements can be switched, and the descriptions of possible bit lengths and coding principles for the syntax elements are merely illustrative.

[0342] A network entity receiving a video data stream (i.e., a video data stream having video content encoded therein using prediction and entropy coding in units of segments into which images of the video content are spatially subdivided, using a coding order between the segments, wherein prediction of the prediction coding and / or entropy coding is limited to within tiles into which images of the video content are spatially subdivided, wherein a sequence of segments in coding order is packetized in coding order into payload packets of a sequence of packets (NAL units) of the video data stream, wherein the sequence of packets is divided into a sequence of access units such that each access unit contains payload packets associated with a respective image of the video content, wherein the payload packets have tiles packetized therein, wherein the sequence of packets has tile identification packets interspersed therein) may be configured to: identify tiles covered by a segment based on the tile identification packets, wherein the tiles are packetized into one or more payload packets immediately following the respective tile identification packets in the packet sequence. The network entity may use the identification result to make a decision regarding transmission tasks. For example, the network entity may process different tiles with different playback priorities. For example, in the event of packet loss, it may be more willing to retransmit payload packets associated with tiles with higher priority than payload packets associated with tiles with lower priority. That is, the network entity may first request the retransmission of lost payload packets associated with tiles with higher priority. Only when there is sufficient time left (depending on the transmission rate) does the network entity continue to request the retransmission of lost payload packets associated with tiles with lower priority. However, the network entity may also be a playback unit that is capable of assigning tiles or payload packets associated with specific tiles to different screens.

[0343] Regarding the manner of using the interspersed region of interest (ROI) information, it should be noted that the ROI packets mentioned below can coexist with the timing control packets and / or tile identification packets mentioned above, and can be implemented in the following ways: combining the information content of these packets in a common packet as described above for the segment prefix, or in the form of separate packets.

[0344] In other words, the manner of using the interspersed ROI information as described above discloses a video data stream having video content encoded using prediction and entropy coding in units of segments into which an image of the video content is spatially subdivided, where the prediction of the predictive coding and / or the entropy coding is limited to within the tiles into which the image of the video content is divided, and where the sequence of segments in the coding order is packetized into the payload packets of a sequence of packets (NAL units) of the video data stream, and the packet sequence is divided into a sequence of access units, so that each access unit collects the payload packets related to an individual image of the video content, and these payload packets have segments packetized therein, and where the packet sequence has ROI packets interspersed therein, and these ROI packets identify the tiles of the image that respectively belong to the ROI of the image.

[0345] Regarding the ROI packets, the same comments as those previously provided for the tile identification packets are valid: the ROI packets can identify the tiles of the image only among those tiles covered by the segments contained in one or more payload packets, and these tiles belong to the ROI of the image, and the individual ROI packets relate to these one or more payload packets (by being immediately in front of the one or more payload packets), as described above for the "prefix segment".

[0346] The ROI packets can allow each prefix segment to identify more than one ROI, and for each of these ROIs (i.e., num_rois_minus1), the related tiles are identified. Then, for each ROI, a priority can be transmitted, thus allowing these ROIs to be sorted according to the priority (i.e., roi_priority[i]). To allow "tracking" of the ROI over time during the image sequence of the video, each ROI can be indexed with an ROI index, so that the ROIs indicated in the ROI packets that extend / span beyond the image boundaries (i.e., over time) can be related to each other (i.e., roi_id[i]).

[0347] A network entity that receives a video data stream (i.e., a video data stream having video content encoded therein in an encoding order of segments that are spatial subdivisions of images of the video content, using prediction and entropy coding on a per-segment basis, where prediction encoding and / or entropy-coded prediction is limited within tiles into which the images of the video content are divided, while entropy coding probability adaptation continues over the entire segment, and where the sequence of segments in the encoding order is packetized into the payload packets of a sequence of packets (NAL units) of the video data stream, the packet sequence being divided into a sequence of access units, so that each access unit collects the payload packets related to an individual image of the video content, the payload packets having segments packetized therein) may be configured to: identify packets based on tiles, and identify packets that continue the packetization of segments that cover tiles belonging to an ROI of an image.

[0348] The network entity may utilize the information conveyed by the ROI packets in a manner similar to that explained above in this previous section with respect to tile identification packets.

[0349] Regarding the current section as well as the previous sections, it should be noted that any network entity such as a MANE or a decoder can determine the tiles covered by the segments of the currently examined payload packet only by: examining the segment order of the segments of the image and examining the progress of the current image portion covered by these segments (with respect to the position of the tiles in the image), which may have been explicitly signaled in the data stream as explained above or may be known to the encoder and decoder by convention. Alternatively, each segment (except for the first segment in the scan order of the image) may have an indication / index (slice_address measured in terms of coding tree blocks) of the same reference (same code) as the first coding block (e.g., CTB), so that the decoder can place each segment (which it reconstructs) in the image starting from this first coding block and in the direction of the segment order. Thus, it suffices for the above-mentioned tile information packet to contain only the index (first_tile_id_in_prefixed_slices) of the first tile covered by any segment of one or more relevant payload packets that immediately follow an individual segment identification packet, because the network entity knows well when sequentially encountering the next tile identification packet: if the index conveyed by the latter tile identification packet differs from the previous one by more than one, then the payload packets between the two tile identification packets must cover tiles having tile indices in between. This holds in the case where, as mentioned, both the tile subdivision and the coding tree subdivision are, for example, based on a column / row subdivision with a raster scan order defined therebetween (for both tiles and coding blocks, which are, for example, column-wise), i.e., the tile indices increase in this raster scan order and the segments follow one another along this raster scan order between the coding blocks according to the segment order.

[0350] The packetized and interleaved slice header signaling aspects derived from the above embodiments can also be combined with any one or any combination of the above aspects. The slice prefixes described previously (e.g., according to version 2) unite all such aspects. The advantage of this aspect is the possibility of making it easier for network entities to access slice header data, since it is conveyed in a separate packet located outside the prefix slice / payload packet and allows for the repeated transmission of slice header data.

[0351] Accordingly, another aspect of the present specification is the aspect of packetized and interleaved slice header signaling, and in other words can be regarded as disclosing a video data stream having sub-parts of images of video content (see coding tree blocks or slices) encoded therein as units, each sub-part being encoded separately in one or more payload packets (see VCL NAL units) of a sequence of packets (NAL units) of the video data stream, the sequence of packets being divided into a sequence of access units, so that each access unit collects the payload packets related to an individual image of the video content, wherein the sequence of packets has slice header packets (slice prefixes) interleaved therein, the slice header packets conveying the slice header data for (and what is missing in) one or more payload packets following an individual slice header packet in the sequence of packets.

[0352] A network entity receiving a video data stream (i.e., a video data stream having sub-parts of images of video content (see coding tree blocks or slices) encoded therein as units, each sub-part being encoded separately in one or more payload packets (see VCL NAL units) of a sequence of packets (NAL units) of the video data stream, the sequence of packets being divided into a sequence of access units, so that each access unit collects the payload packets related to an individual image of the video content, wherein the sequence of packets has slice header packets interleaved therein) can be configured to: read the slice header and payload data of a slice from the packet, however, wherein the slice header data is derived from the slice header packet, and skip reading the slice header for one or more payload packets following an individual slice header packet in the sequence of packets, but instead use the slice header derived from the slice header packet following one or more payload packets.

[0353] If the aspects mentioned above hold true, it is possible that a packet (here a slice header packet) may also have the following functionality: to indicate to any network entity such as a MANE or a decoder the start of a decoding unit or the start of the run of one or more payload packets prefixed by an individual packet. Thus, a network entity according to this aspect can identify payload packets based on the above-mentioned syntax elements in this packet (i.e., single_slice_flag combined with, for example, decoding_unit_start_flag, where the latter flag allows retransmission of a copy of a specific slice header packet within a decoding unit as discussed above), for which the reading of the slice header must be skipped. This is useful, for example, because the slice headers of slices within a decoding unit can change along the slice sequence, and thus, while the slice header packet at the start of a decoding unit may have the decoding_unit_start_flag set (equal to 1), the slice header packets located in between may have this flag not set to prevent any network entity from misinterpreting the occurrence of this slice header packet as the start of a new decoding unit.

[0354] Although some aspects have been described in the context of a device, it is clear that these aspects also represent a description of a corresponding method, where a block or device corresponds to a method step or a feature of a method step. Similarly, the aspects described in the context of a method step also represent a description of a corresponding block or item or feature of a corresponding device. Some or all of the method steps can be performed by (or using) hardware devices, such as a microprocessor, a programmable computer, or an electronic circuit. In some embodiments, one or more of the most important method steps can be performed by such a device.

[0355] The encoded data stream of the present invention can be stored on a digital storage medium or can be transmitted over a transmission medium such as a wireless transmission medium or a wired transmission medium (such as the Internet).

[0356] Depending on certain implementation requirements, embodiments of the present invention can be implemented in hardware or software. A digital storage medium (e.g., a floppy disk, a DVD, a Blu-ray, a CD, a ROM, a PROM, an EPROM, an EEPROM, or a flash memory) storing an electronically readable control signal can be used to execute the implementation, which cooperates with (or is capable of cooperating with) a programmable computer system such that individual methods are executed. Thus, the digital storage medium can be computer-readable.

[0357] Some embodiments according to the present invention include a data carrier having an electronically readable control signal, which is capable of cooperating with a programmable computer system such that one of the methods described herein is executed.

[0358] Generally, embodiments of the present invention can be implemented as a computer program product having program code that, when run on a computer, operates to perform one of the methods. The program code can be stored, for example, on a machine-readable carrier.

[0359] Other embodiments include a computer program stored on a machine-readable carrier for performing one of the methods described herein.

[0360] In other words, an embodiment of the method of the present invention is thus a computer program having program code that, when run on a computer, is used to perform one of the methods described herein.

[0361] Another embodiment of the method of the present invention is thus a data carrier (or digital storage medium, or computer-readable medium) on which is recorded a computer program for performing one of the methods described herein. The data carrier, the digital storage medium, or the recorded medium is generally tangible and / or non-transitory.

[0362] Another embodiment of the method of the present invention is thus a data stream or signal sequence representing a computer program for performing one of the methods described herein. The data stream or signal sequence can be configured, for example, to be transmitted via a data communication connection (such as via the Internet).

[0363] Another embodiment includes a processing component, such as a computer or a programmable logic device, configured or adapted to perform one of the methods described herein.

[0364] Another embodiment includes a computer on which is installed a computer program for performing one of the methods described herein.

[0365] Another embodiment according to the present invention includes a device or a system configured to transmit (e.g., electronically or optically) a computer program for performing one of the methods described herein to a receiver. The receiver can be, for example, a computer, a mobile device, a memory device, etc. The device or system can include, for example, a file server for transmitting the computer program to the receiver.

[0366] In some embodiments, a programmable logic device (e.g., a field programmable gate array) can be used to perform some or all of the functionality of the methods described herein. In some embodiments, the field programmable gate array can cooperate with a microprocessor to perform one of the methods described herein. Generally, the methods are preferably performed by any hardware device.

[0367] In some embodiments of the present invention, the present invention can also be configured in the following manner:

[0368] 1. A video data stream having video content (16) encoded therein in units of sub - parts (24) of images (18) of the video content (16), each sub - part (24) being encoded into one or more payload packets (32) of a packet sequence (34) of the video data stream (22), the packet sequence (34) being divided into a sequence of access units (30), so that each access unit (30) collects payload packets (32) related to respective images (18) of the video content (16), wherein the packet sequence (34) has timing control packets (36) interspersed therein, so that the timing control packets (36) subdivide the access units (30) into decoding units (38), so that at least some access units (30) are subdivided into two or more decoding units (38), wherein each timing control packet (38) signals the decoder buffer fetch time of the decoding unit (38), and the payload packets (32) of the decoding unit follow each of the timing control packets (38) in the packet sequence (34).

[0369] 2. The video data stream according to item 1, wherein the sub - parts (24) are tiles, and each payload packet (32) contains one or more tiles.

[0370] 3. The video data stream according to item 2, wherein the tiles include independently decodable tiles and dependency segments, and the dependency segments allow decoding using entropy and prediction decoding beyond tile boundaries for WPP processing.

[0371] 4. The video data stream according to any one of items 1 to 3, wherein each packet of the packet sequence (34) can be assigned only one packet type out of a plurality of packet types, wherein payload packets (32) and timing control packets (36) are different packet types, and the occurrence of packets of the plurality of packet types in the packet sequence (34) is subject to specific restrictions that define an order between the packet types, and the packets within each access unit (30) will comply with the order, so that these restrictions can be used to detect access unit boundaries by detecting violations of the restrictions, and the access unit boundaries remain at the same position within the packet sequence even if packets of any removable packet type are removed from the video data stream, wherein payload packets (32) are of a non - removable packet type and timing control packets (36) are of the removable packet type.

[0372] 5. The video data stream according to any one of items 1 to 4, wherein each packet contains a packet type indicating a syntax element part.

[0373] 6. The video data stream according to item 5, wherein the packet type indicating the syntax element part is included in the packet type field in the header of each of the packets, the content of each of the packets is different between the payload packet and the timing control packet, and for the timing control packet, the SEI packet type field distinguishes between the timing control packets on the one hand and between different types of SEI packets on the other hand.

[0374] 7. An encoder for encoding video content (16) into a video data stream (22) in units of sub - parts (24) of images (18) of the video content (16), wherein each sub - part (24) is respectively encoded into one or more payload packets (32) of a packet sequence (34) of the video data stream (22), so that the packet sequence (34) is divided into a sequence of access units (30), and each access unit (30) collects the payload packets (32) related to each image (18) of the video content (16), the encoder is configured to interleave timing control packets (36) in the packet sequence (34), so that the timing control packets (36) subdivide the access units (30) into decoding units (38), so that at least some access units (30) are subdivided into two or more decoding units (38), wherein each timing control packet (36) signals the decoder buffer extraction time of the decoding unit (38), and the payload packets (32) of the decoding unit follow each of the timing control packets (36) in the packet sequence (34).

[0375] 8. The encoder according to item 7, wherein, here the encoder is configured to, during the process of encoding the current image of the video content,

[0376] encode the current sub - part (24) of the current image (18) into the current payload packet (32) of the current decoding unit (38);

[0377] at a first moment, within the data stream, transmit the current decoding unit (38) prefixed by the current timing control packet (36) by setting the decoder buffer extraction time signaled by the current timing control packet (36); and

[0378] at a second moment, encode another sub - part of the current image, the second moment being later than the first moment.

[0379] 9. A method for encoding video content (16) into a video data stream (22) in units of sub - parts (24) of images (18) of the video content (16), wherein each sub - part (24) is encoded into one or more payload packets (32) of a packet sequence (34) of the video data stream (22), so that the packet sequence (34) is divided into a sequence of access units (30), and each access unit (30) collects payload packets (32) related to respective images (18) of the video content (16). The method includes interspersing timing control packets (36) in the packet sequence (34), so that the timing control packets (36) subdivide the access units (30) into decoding units (38), so that at least some access units (30) are subdivided into two or more decoding units (38), wherein each timing control packet (36) signals the fetch time of the decoder buffer of the decoding unit (38), and the payload packets (32) of the decoding unit follow each of the timing control packets (36) in the packet sequence (34).

[0380] 10. A decoder for decoding a video data stream (22), the video data stream (22) having the video content (16) encoded therein in units of sub - parts (24) of images (18) of the video content (16), each sub - part being encoded into one or more payload packets (32) of a packet sequence (34) of the video data stream (22), the packet sequence (34) being divided into a sequence of access units (30), so that each access unit (30) collects payload packets (32) related to respective images (18) of the video content (16). The decoder includes a buffer for buffering the video data stream or the reconstructed video content obtained therefrom by decoding the video data stream, and the decoder is configured to search for timing control packets (36) interspersed in the packet sequence, subdivide the access units (30) into decoding units (38) at the timing control packets (36), so that at least some access units are subdivided into two or more decoding units, and empty the buffer in units of the decoding units.

[0381] 11. The decoder according to item 10, wherein the decoder is configured to: when searching for the timing control packet (36), check the packet type indicating the syntax element part in each packet, and if the value of the packet type indicating the syntax element part is equal to a predetermined value, identify the individual packet as the timing control packet (36).

[0382] 12. A method for decoding a video data stream (22), the video data stream (22) having the video content (16) encoded therein in units of sub - portions (24) of images (18) of the video content (16), each sub - portion being encoded into one or more payload packets (32) of a packet sequence (34) of the video data stream (22), the packet sequence (34) being divided into a sequence of access units (30), so that each access unit (30) collects the payload packets (32) related to respective images (18) of the video content (16), the method using a buffer to buffer the video data stream or the reconstruction of the video content obtained therefrom by decoding the video data stream, and the method comprising: finding timing control packets (36) interspersed in the packet sequence, subdividing the access units (30) into decoding units (38) at the timing control packets (36), so that at least some access units are subdivided into two or more decoding units, and emptying the buffer in units of the decoding units.

[0383] 13. A network entity for transmitting a video data stream (22), the video data stream (22) having the video content (16) encoded therein in units of sub - portions (24) of images (18) of the video content (16), each sub - portion being encoded into one or more payload packets (32) of a packet sequence (34) of the video data stream (22), the packet sequence (34) being divided into a sequence of access units (30), so that each access unit (30) collects the payload packets (32) related to respective images (18) of the video content (16), the decoder being configured to: find timing control packets (36) interspersed in the packet sequence, subdivide the access units into decoding units at the timing control packets (36), so that at least some access units (30) are subdivided into two or more decoding units (38), derive the decoding buffer fetch time of the decoding units (38) from each timing control packet (36), the payload packets (32) of the decoding units following each of the timing control packets (36) in the packet sequence (34), and perform the transmission of the video data stream according to the decoding buffer fetch time of the decoding units (38).

[0384] 14. A method for transmitting a video data stream (22), the video data stream (22) having video content (16) encoded therein in units of sub - parts (24) of images (18) of the video content (16), each sub - part being encoded into one or more payload packets (32) of a packet sequence (34) of the video data stream (22), the packet sequence (34) being divided into a sequence of access units (30), so that each access unit (30) collects the payload packets (32) related to respective images (18) of the video content (16), the method comprising: finding timing control packets (36) interspersed in the packet sequence, subdividing the access units into decoding units at the timing control packets (36), so that at least some access units (30) are subdivided into two or more decoding units (38), deriving a decoder buffer fetch time for the decoding units (38) from each timing control packet (36), the payload packets (32) of the decoding units following the respective timing control packets (36) in the packet sequence (34), and performing the transmission of the video data stream based on the decoder buffer fetch time of the decoding units (38).

[0385] 15. A video data stream having video content (16) encoded therein in units of segments (24) using prediction and entropy coding, the images (18) of the video content (16) being encoded using an encoding order between the segments (24), wherein the prediction for the prediction coding and / or entropy coding is limited within tiles (70) and the video content is spatially subdivided into the segments (24), the images of the video content being spatially subdivided into tiles (70), wherein the sequence of the segments (24) is packetized into the payload packets (32) of a packet sequence of the video data stream according to the encoding order, the packet sequence (34) being divided into a sequence of access units (30), so that each access unit collects the payload packets (32) having segments (24) related to respective images (18) of the video content (16) packetized therein, wherein the packet sequence (34) has tile identification packets (72) interspersed between the payload packets of one access unit, the tile identification packets identifying one or more tiles (70) covered by any segment (24) packetized into one or more payload packets (32) following the respective tile identification packet (72) in the packet sequence (34).

[0386] 16. The video data stream according to item 15, wherein the tile identification packet (72) identifies the one or more tiles (70) covered by any segment (24) packetized into only the immediately following payload packet (32).

[0387] 17. The video data stream according to item 15, wherein the tile identification packet (72) identifies one or more tiles (70) covered by any segment (24) in one or more payload packets (32) immediately following the respective tile identification packet (72) in the packet sequence (34), up to the earlier of the end of the current access unit (30) to which the packets are encapsulated and the next tile identification packet (72) in the packet sequence, respectively.

[0388] 18. A network entity configured to receive a video data stream according to any one of items 15 to 16, and identify tiles (70) covered by a segment (24) based on the tile identification packet (72), the segment (24) being encapsulated in one or more payload packets (72) immediately following the respective tile identification packet (72) in the packet sequence.

[0389] 19. The network entity according to item 18, wherein the network entity is configured to use the result of the identification to make a decision on a transmission task related to the video data stream.

[0390] 20. The network entity according to item 18 or 19, wherein the transmission task includes a retransmission request for a defective packet.

[0391] 21. The network entity according to item 18 or 19, wherein the network entity is configured to process different tiles (70) with different priorities by assigning a higher priority to a tile identification packet (72) and the payload packet immediately following each said tile identification packet (72), the payload packet (72) having a segment encapsulated therein that covers a tile with a higher priority identified by each said tile identification packet (72), compared to a tile identification packet (72) and the payload packet immediately following each said tile identification packet (72), the payload packet (72) having a segment encapsulated therein that covers a tile with a lower priority identified by each said tile identification packet.

[0392] 22. The network entity according to item 21, wherein the network entity is configured to: first request a retransmission of a payload packet assigned a higher priority before requesting a retransmission of any payload packet assigned a lower priority.

[0393] 23. A method, comprising: receiving a video data stream according to any one of items 15 to 16, and identifying tiles (70) covered by a segment (24) based on the tile identification packets (72), the segment (24) being packetized into one or more payload packets (72) immediately following each of the tile identification packets (72) in the packet sequence.

[0394] 24. A video data stream having the video content (16) encoded therein in units of sub - portions (24) of images (18) of the video content (16), each sub - portion (24) being encoded into one or more payload packets (32) of a packet sequence (34) of the video data stream (22), the packet sequence (34) being divided into a sequence of access units (30), so that each access unit (30) collects the payload packets (32) related to respective images (18) of the video content (16), wherein at least some access units (30) have the packet sequence (34) with ROI packets (64) interspersed therein, so that the timing control packet (36) subdivides the access unit (30) into decoding units (38), so that at least some access units (30) have ROI packets interspersed between the payload packets (32) related to the images of the individual access units, wherein each ROI packet is related to one or more subsequent payload packets following each of the ROI packets in the packet sequence (34), and identifies whether the sub - portion (24) encoded into any payload packet among the one or more payload packets that the ROI packet is about covers the region of interest of the video content.

[0395] 25. The video data stream according to item 24, wherein the sub - portion is a segment, and the video content is encoded into the video data stream using predictive and entropy coding, wherein the predictive and / or entropy coding is limited within the tiles into which the images of the video content are divided, and wherein each ROI packet further identifies the tile within which the sub - portion (24) encoded into any payload packet among the one or more payload packets that the ROI packet is about covers the region of interest.

[0396] 26. The video data stream according to item 24 or 25, wherein each ROI packet is uniquely related to the immediately following payload packet.

[0397] 27. The video data stream according to item 24 or 25, wherein each ROI packet is related to all payload packets immediately following each of the ROI packets in the packet sequence up to the end of the access unit and the earlier of the next ROI packet respectively, and each of the ROI packets is arranged within the access unit.

[0398] 28. A network entity configured to receive the video data stream according to any one of items 24 to 27, and identify the ROI of the video content based on the ROI packets.

[0399] 29. The network entity according to item 27, wherein the network entity is configured to use the result of the identification to make a decision on a transmission task related to the video data stream.

[0400] 30. The network entity according to item 28 or 29, wherein the transmission task includes a retransmission request for defective packets.

[0401] 31. The network entity according to item 28 or 29, wherein the network entity is configured to process the region of interest with increased priority by assigning a higher priority to the following ROI packets (72) and the one or more payload packets immediately following each of the ROI packets (72) to which each of the ROI packets relates, the ROI packets signaling the coverage of the region of interest by the sub - part (24) in any of the payload packets in the one or more payload packets to which each of the ROI packets relates, as compared to the following ROI packets and the one or more payload packets immediately following each of the ROI packets (72) to which each of the ROI packets relates, the ROI packets signaling no coverage.

[0402] 32. The network entity according to item 31, wherein the network entity is configured to: before requesting a retransmission of any payload packet assigned a lower priority, first request a retransmission of a payload packet assigned a higher priority.

[0403] 33. A method including receiving the video data stream according to any one of items 23 to 26, and identifying the ROI of the video content based on the ROI packets.

[0404] 34. A computer program having program code for performing the method according to item 9, 12, 14, 23 or 33 when run on a computer.

[0405] The embodiments described above merely illustrate the principles of the present invention. It should be understood that those skilled in the art will appreciate modifications and variations to the arrangements and details described herein. Accordingly, the intention of the present invention is limited only by the scope of the appended patent claims and not by the specific details presented by way of description and illustration of the embodiments herein.

[0406] References

[0407] [1] Thomas Wiegand, Gary J. Sullivan, Gisle Bjontegaard, Ajay Luthra, 「Overview of the H.264 / AVC Video Coding Standard」, IEEE Trans. Circuits Syst. Video Technol., vol. 13, N7, July 2003.

[0408] [2] JCT-VC, 「High-Efficiency Video Coding (HEVC) text specification Working Draft 7」, JCTVC-I1003, May 2012.

[0409] [3] ISO / IEC 13818-1: MPEG-2 Systems specification.

[0410] [4] IETF RFC 3550 - Real-time Transport Protocol.

[0411] [5] Y.-K. Wang et al., 「RTP Payload Format for H.264 Video」, IETF RFC6184, http: / / tools.ietf.org / html /

[0412] [6] S. Wenger et al., 「RTP Payload Format for Scalable Video Coding」, IETF RFC 6190, http: / / tools.ietf.org / html / rfc6190

[0413] [6]T. Schierl et al., "RTP Payload Format for High Efficiency Video Coding", IETF internet draft, http: / / datatracker.ietf.org / doc / draft-schierl-payload-rtp-h265 / .

Claims

1. A method for transmitting a video data stream, comprising: Transmitting a video data stream over a transmission medium, the video data stream having video content (16) encoded therein in units of segments (24) of images (18) of the video content (16) using prediction and entropy coding, each segment (24) being encoded into one or more payload packets (32) of a packet sequence (34) of the video data stream (22), the packet sequence (34) being divided into a sequence of access units (30), such that each access unit (30) collects payload packets (32) related to an individual image (18) of the video content (16), wherein, The packet sequence (34) has ROI packets (64) interspersed therein, and the ROI packets (64) identify tiles of the image that belong to the ROI of the image, wherein the ROI packets can identify tiles of the image that belong to the ROI of the image only in tiles covered by segments contained in one or more payload packets related to an individual ROI packet, wherein the one or more payload packets related to the individual ROI packet are the payload packets immediately preceding the individual ROI packet.

2. A decoder, comprising a processor, wherein, The processor is configured to: Decode a video data stream, the video data stream having the video content (16) encoded therein in units of segments (24) of images (18) of the video content (16) using prediction and entropy coding, and respectively encode each segment (24) into one or more payload packets (32) of a packet sequence (34) of the video data stream (22), the packet sequence (34) being divided into a sequence of access units (30), so that each access unit (30) collects the payload packets (32) related to an individual image (18) of the video content (16), wherein the packet sequence (34) has ROI packets (64) interspersed therein, and the ROI packets (64) identify tiles of the image that belong to the ROI of the image, wherein the ROI packets can identify tiles of the image that belong to the ROI of the image only in tiles covered by segments contained in one or more payload packets related to an individual ROI packet, wherein the one or more payload packets related to the individual ROI packet are the payload packets immediately preceding the individual ROI packet; and Packetize the segments that will cover the tiles belonging to the ROI of the image based on the identification by the ROI packets.

3. An encoder, comprising a processor, wherein, The processor is configured to: Encoded video data stream, the video data stream having video content (16) encoded therein in units of segments (24) of images (18) of the video content (16) using prediction and entropy coding, each segment (24) being encoded into one or more payload packets (32) of a packet sequence (34) of the video data stream (22), the packet sequence (34) being divided into a sequence of access units (30), so that each access unit (30) collects payload packets (32) related to individual images (18) of the video content (16), wherein the packet sequence (34) has ROI packets (64) interspersed therein, and the ROI packets (64) identify tiles of the image that belong to the ROI of the image, wherein the ROI packets can identify tiles of the image that belong to the ROI of the image only in tiles covered by segments contained in one or more payload packets related to an individual ROI packet, wherein the one or more payload packets related to an individual ROI packet are the payload packets immediately preceding the individual ROI packet, so that packets encapsulating segments that will cover the tiles belonging to the ROI of the image are identified based on the ROI packet identification.

4. A decoding method, comprising: Decoding a video data stream, the video data stream having video content (16) encoded therein in units of segments (24) of images (18) of the video content (16) using prediction and entropy coding, each segment (24) being encoded into one or more payload packets (32) of a packet sequence (34) of the video data stream (22), the packet sequence (34) being divided into a sequence of access units (30), so that each access unit (30) collects payload packets (32) related to individual images (18) of the video content (16), wherein the packet sequence (34) has ROI packets (64) interspersed therein, and the ROI packets (64) identify tiles of the image that belong to the ROI of the image, wherein the ROI packets can identify tiles of the image that belong to the ROI of the image only in tiles covered by segments contained in one or more payload packets related to an individual ROI packet, wherein the one or more payload packets related to an individual ROI packet are the payload packets immediately preceding the individual ROI packet; and Identifying, based on the ROI packet identification, packets encapsulating segments that will cover the tiles belonging to the ROI of the image.

5. An encoding method, comprising: Encoded video data stream, the video data stream having video content (16) encoded therein in units of segments (24) of images (18) of the video content (16) using prediction and entropy coding, each segment (24) being encoded into one or more payload packets (32) of a packet sequence (34) of the video data stream (22), the packet sequence (34) being divided into a sequence of access units (30), so that each access unit (30) collects the payload packets (32) related to an individual image (18) of the video content (16), wherein the packet sequence (34) has ROI packets (64) interspersed therein, and the ROI packets (64) identify the tiles of the image that belong to the ROI of the image, wherein the ROI packets can identify the tiles of the image that belong to the ROI of the image only in the tiles covered by the segments contained in one or more payload packets involved in an individual ROI packet, wherein the one or more payload packets involved in the individual ROI packet are the payload packets immediately preceding the individual ROI packet, so that the packets encapsulating the segments that will cover the tiles belonging to the ROI of the image are identified based on the ROI packet identification.

6. A non-transitory computer-readable medium having a computer program stored thereon, the computer program having program code for performing the method according to claim 1, 4 or 5 when run on a computer.

Citation Information

Patent Citations

  • Video data stream, encoder, method for encoding video content, and decoder

    CN110536136A

  • Video coding device using a forced intra update and aMethod thereof

    KR1020080001156A

  • Video decoding implementations for a graphics processing unit

    US20090002379A1