Low-delay picture coding
By implementing dependent slices and optimizing WPP entry points, the challenges of achieving high compression efficiency and low latency in video encoding are addressed, resulting in improved encoding efficiency and reduced delay.
Patent Information
- Application Number
- JP2025025636
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2012-06-29
- Filing Date
- 2025-02-20
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2033-04-15
AI Technical Summary
Current video encoding technologies, including HEVC, struggle to achieve high compression efficiency while maintaining low latency, particularly in wavefront parallel processing (WPP) concepts.
The introduction of dependent slices, which allow for mutual dependencies across slice boundaries, and the use of the start syntax part of a slice to determine the position of a WPP entry point, enhance the efficiency of wavefront parallel processing and reduce end-to-end delay.
This approach enables increased encoding efficiency, reduced encoding overhead, and lower end-to-end delay, facilitating parallel decoding and improving the robustness of the encoding process.
Smart Images

Figure 2025093939000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to low-latency coding of images.
Background Art
[0002] In current HEVC design slices, entropy slices (previously lightweight slices), tiles, and WPP (wavefront parallel processing) are included as tools for parallelization.
[0003] For parallelization of video encoders and decoders, picture-level partitioning has several advantages compared to other approaches. In conventional video codecs, such as H.264 / AVC [1], picture partitioning was only possible with regular slices that had a high cost in terms of coding efficiency. For scalable parallel H.264 / AVC decoding, it was necessary to combine macroblock-level parallel processing for picture reconstruction and frame-level parallel processing for entropy decoding. However, this approach results in limited reduction in picture latency and high memory usage. To overcome these limitations, a new picture partitioning strategy is included in the HEVC codec. The current reference software version (HM-6) includes four different approaches: regular or normal slices, entropy slices, wavefront parallel processing (WPP) substreams, and tiles. Typically, those picture partitions include one set of the largest coding units (LCUs), or, in synonymous terms, coding tree units (CTUs) as defined in HEVC or subsets thereof.
[0004] FIG. 1 shows an image 898 illustratively arranged in a regular slice 900 for each row 902 of LCUs or macroblocks in the image. Regular or normal slices (as defined in H.264 [1]) have the maximum coding penalty such that they break the entropy decoding and prediction dependencies. An entropy slice, like a slice, breaks the entropy decoding dependency but allows prediction (and filtering) to cross slice boundaries.
[0005] In WPP, the image partitions are row interleaved, and furthermore, both entropy decoding and prediction are enabled to use data from blocks in other partitions. Thus, the coding loss is minimized and at the same time, wavefront parallel processing can be utilized. However, interleaving disrupts the bitstream causality such that the next partition is required for the conventional partition to decode.
[0006] FIG. 2 illustratively shows an image 898 divided into two rows 904a, 904b of tiles 906 partitioned horizontally. The tiles define horizontal 908 and vertical boundaries 910 that partition the image 898 into tile columns 912a, b, c and rows 904a, b. Similar to the regular slice 900, the tile 906 breaks the entropy decoding and prediction dependencies but does not require a header for each tile.
[0007] For each of these techniques, the number of partitions can be freely selected by the encoder. Generally, having more partitions results in higher compression loss. However, in WPP, the loss propagation is not so high, and thus, the number of image partitions can even be fixed to 1 for each row. This also brings several advantages. First, for WPP, bitstream causality is guaranteed. Second, decoder implementation can assume that a certain amount of parallel processing is available, which also increases with the resolution. Furthermore, finally, neither context selection nor prediction dependency needs to be disrupted when decoding in wavefront order, resulting in relatively low encoding loss.
[0008] However, so far, all parallel encodings in the transform concept have not been able to provide high compression efficiency while maintaining low latency. This is also the case for the WPP concept. A slice is the minimum unit of transfer in the encoding pipeline, and furthermore, some WPP sub-streams still have to be transferred continuously.
Summary of the Invention
Problems to be Solved by the Invention
[0009] Therefore, an object of the present invention is to provide an image encoding concept that can increase the encoding efficiency with increased efficiency, for example, by further reducing the end-to-end delay or reducing the encoding overhead incurred, and enabling parallel decoding, for example, according to wavefront parallel processing.
Means for Solving the Problems
[0010] This object is achieved by the subject matter of the independent claims.
[0011] The basic finding of the present invention is that the concept of normal slices, where slices are encoded / decoded completely independently of the image regions outside each slice or at least independently of the regions outside each slice as far as entropy coding is concerned, is abandoned in favor of supporting different modes of slices, namely, for example, dependent slices that allow for mutual dependencies across slice boundaries, and others called normal slices that do not allow for such dependencies. In this case, parallel processing concepts such as wavefront parallel processing can be realized with a reduced end-to-end delay.
[0012] A further basic finding of the present invention, which can be combined with the first one or used individually, is that when the start syntax part of a slice is used to determine the position of a WPP entry point, the WPP processing concept can be made more efficient.
[0013] Preferred embodiments of the present application are described below with reference to the figures, and advantageous embodiments are the subject matter of the dependent claims.
Brief Description of the Drawings
[0014]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11a
Figure 11b
Figure 11c
Figure 12a
Figure 12b
Figure 13a
Figure 13b
Figure 13c
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
Figure 20
Figure 21
Figure 22
Figure 23
Figure 24
Figure 25
Figure 26
Figure 27
Figure 28
Figure 29
Figure 30
Figure 31
[0015] In the following, the description starts with an explanation of current concepts for enabling parallel image processing and low-latency encoding, respectively. The problems that arise when desiring to have both capabilities are outlined. In particular, as will be apparent from the following description, the WPP substream concept as has been taught so far somehow conflicts with the desire to have low latency due to the need to transmit the WPP substream by grouping it into one slice. The following embodiments represent parallel processing concepts, such as the WPP concept, that can be applied to applications that require even less latency by expanding the slice concept, i.e., by introducing another type of slice, later referred to as a dependent slice.
[0016] Minimizing end-to-end video latency from capture to display is one of the main objectives in applications such as video conferencing.
[0017] The signal processing chain for digital video transmission consists of a camera, a capturing device, an encoder, encapsulation, transmission, a demultiplexer, a decoder, a renderer, and a display. Each of these stages contributes to the end-to-end latency by buffering the image data before its serial transmission to the subsequent stage.
[0018] Some applications require minimization of such latency, for example, remote operation of objects in dangerous areas where the objects being operated on are not directly visible, or minimally invasive surgery. Even short latency can cause significant problems with proper operation or result in significant errors.
[0019] In many cases, all video frames are buffered within the processing stage, for example, to enable intra-frame processing. Some stages collect data to form packets to be sent to the next stage. Generally, there is a lower bound on the latency resulting from local processing requirements. This is analyzed in more detail for each individual stage below.
[0020] Processing within the camera does not necessarily require intra-frame signal processing, as the minimum latency is given by the integration time of the sensor, which is limited by the frame rate, and some design choices by the hardware manufacturer. The camera output is typically associated with a scan order that starts processing in the upper left corner, moves to the upper right corner, and then continues to the lower right corner for each line. As a result, it takes approximately one frame period for all data to be transferred from the sensor to the camera output.
[0021] The capturing device can send camera data immediately after reception, but it typically buffers some data and further generates bursts to optimize data access to memory or storage. Further, the connection between the camera / capture and the computer's memory typically limits the bitrate for sending the captured image data to the memory for further processing (encoding). Typically, the camera is connected via USB2.0 or, soon, USB3.0, which always involves partially transferring the image data to the encoder. This limits the parallelization ability on the encoder side in extreme low-latency scenarios, i.e., the encoder tries to start encoding as soon as the data becomes available from the camera in, for example, raster scan order from the top to the bottom of the image.
[0022] In the encoder, there is some degree of freedom to trade off encoding efficiency for reduction of processing delay with respect to the data rate required for a particular video fidelity.
[0023] The encoder uses data that has already been sent to predict the image to be encoded later. Generally, the difference between the actual image and the prediction can be encoded with fewer bits than would be required without prediction. This predicted value needs to be made available at the decoder, and thus the prediction is based on a previously decoded part of the same image (intra-frame prediction) or another previously processed image (inter-frame prediction). The pre-HEVC video encoding standard uses only parts of the image that are above the same line or to the left within the same line, which have been previously encoded for intra-frame prediction, motion vector prediction, and entropy coding (CABAC).
[0024] In addition to optimizing the prediction structure, the impact of parallel processing can be considered. Parallel processing requires the identification of image regions that can be processed independently. For practical reasons, continuous regions such as horizontal or vertical rectangles are selected, which are often referred to as "tiles". In the case of low latency constraints, those regions should enable the parallel encoding of data input from capture to memory as soon as possible. Assuming a raster scan memory transfer, a vertical partition of the raw data makes sense to start encoding immediately. Divide the image into vertical partitions (see the figure below), within such tiles, intra prediction, motion vector prediction, and entropy encoding (CABAC) can bring considerable encoding efficiency. To minimize latency, only the part of the image starting from the top is transferred to the encoder's frame memory, and further, parallel processing should be started in the vertical tiles.
[0025] Another way to enable parallel processing is to use WPP within a regular slice compared to a tile, where the "rows" of the tile are included in a single slice. The data within the slice can be encoded in parallel using the WPP substream within the slice. The separation of the image into slice 900 and tile / WPP substream 914 is shown in the form of Figure 3 / Figure 1 of the example.
[0026] Thus, Figure 3 shows the assignment of parallel encoded partitions such as 906 or 914 to, for example, slices or network transfer segments (a single network packet or multiple network 900 packets).
[0027] The encapsulation of the encoded data into Network Abstraction Layer (NAL) units adds some headers to the data blocks, if applicable, as defined in H.264 or HEVC, before transmission or during the encoding process, which enables the identification of each block and the reordering of the blocks. In the case of the standard, additional signaling is not required since the order of the encoding elements is always the decoding order, which is given by the implicit assignment of the position of the tiles or general encoding fragments.
[0028] When parallel processing is considered for an additional transport layer for low-latency parallel transmission, i.e., the transport layer means sending fragments as shown in FIG. 4 as they are encoded, the image partitions for the tiles can be reordered to enable low-latency transmission. Those fragments may be slices that are not fully encoded, they may be subsets of slices, or they may be included in dependent slices.
[0029] When creating additional fragments, there is a trade-off between the highest efficiency in large data blocks due to the header information adding a certain number of bytes and the delay due to the large data blocks of the parallel encoder needing to be buffered before transmission. If the encoded representation of the vertical tile 906 is separated into a number of fragments 916 that are transmitted as soon as they are fully encoded, the overall delay can be reduced. The size of each fragment can be determined with respect to a fixed image area, such as a macroblock, LCU, etc., or with respect to the maximum data as shown in FIG. 4.
[0030] Thus, FIG. 4 shows the general fragmentation of a frame with a tile encoding approach for minimum end-to-end delay.
[0031] Similarly, FIG. 5 shows the fragmentation of a frame with a WPP encoding approach for minimum end-to-end delay.
[0032] Transmission can add additional delay, for example, when additional block-oriented processing is applied, such as when a forward error correction code increases the robustness of the transmission. Also, the network infrastructure (such as routers) or the physical link can add delay, which is typically known as the delay time for the connection. In addition to the delay time, the transmission bit rate determines the time (delay) for transferring data from participant a to participant b in a conversation as shown in FIG. 6 using a video service.
[0033] If the encoded data blocks are transmitted out of order, a delay reordering must be considered. Decoding can start as soon as the data unit arrives, assuming that there are no other data units that must be decoded before this one.
[0034] In the case of tiles, there is no dependency between tiles, so the tiles can be decoded immediately. If the fragments are made of tiles, such as separate slices for each fragment as shown in FIG. 4, the fragments can be directly transferred as soon as they are each encoded and their contained LCU or CU is encoded.
[0035] The renderer assembles the output of the parallel decoding engines and further sends the combined image line by line to the display.
[0036] The display does not necessarily add any delay, but in practice, it can perform some intra-frame processing before the image data is actually displayed. This depends on the design choice of the hardware manufacturer.
[0037] In summary, we can affect stage encoding, encapsulation, transmission, and decoding in order to achieve minimum end-to-end latency. When we use parallel processing, tiling, and fragmentation within tiles, we can significantly reduce the overall latency as compared to a commonly used processing chain that adds a latency of approximately one frame at each of these stages as shown in FIG. 8, as shown in FIG. 7.
[0038] In particular, FIG. 7 shows encoding, transmission, and decoding for tiles with a generic subset at minimum end-to-end latency, while FIG. 8 shows the end-to-end latency achieved in common.
[0039] HEVC enables the use of slice partitioning, tile partitioning in the following further ways.
[0040] Tile: An integer number of tree blocks that occur simultaneously in one column and one row, and are sequentially ordered continuously in the tile's tree block raster scan. The partitioning of each image into tiles is a partition. Tiles in an image are sequentially ordered continuously in the image's tile raster scan. Although a slice contains tree blocks that are continuous in the tile's tree block raster scan, these tree blocks are not necessarily continuous in the image's tree block raster scan.
[0041] Slice: An integer number of tree blocks that are sequentially ordered continuously in a raster scan. The partitioning of each image into slices is a partition. The tree block address is derived from the first tree block address in the slice (as represented in the slice header).
[0042] Raster scan: The mapping of a rectangular two-dimensional pattern to a one-dimensional pattern in which the first entry in a one-dimensional pattern is from the topmost first row of a two-dimensional pattern scanned from left to right, and subsequent rows (descending) of the pattern scanned from left to right follow in a similar manner.
[0043] Tree block: An NxN block of luma samples and two corresponding blocks of chroma samples of an image having three sample arrays, or an NxN block of samples of a monochrome image or an image encoded using three separate color planes. The division of a slice into tree blocks is a partition division.
[0044] Partition division: The division of a set into subsets such that each element of the set is in exactly one of the subsets.
[0045] Quad tree: A tree that can divide a parent node into four child nodes. A child node may become a parent node for another division into four child nodes.
[0046] In the following, the spatial re-division of images, slices, and tiles is described. In particular, the following description specifies how an image is partitioned into slices, tiles, and encoding tree blocks. An image is divided into slices and tiles. A slice is a series of encoding tree blocks. Similarly, a tile is a series of encoding tree blocks.
[0047] Samples are processed in units of coding tree blocks. The luma array size per tree block in a sample in both width and height is CtbSize. The width and height of the chroma array per coding tree block are CtbWidthC and CtbHeightC, respectively. For example, an image can be divided into two slices as shown in the following figure. As another example, an image can be divided into three tiles as shown in the second following figure.
[0048] Unlike a slice, a tile is always rectangular and further always contains an integer number of coding tree blocks in a coding tree block raster scan. A tile can consist of coding tree blocks included in more than one slice. Similarly, a slice can consist of coding tree blocks included in more than one tile.
[0049] Figure 9 shows an image 898 having an 11×9 coding tree block 918 partitioned into two slices 900a, b.
[0050] Figure 10 shows an image having a 13×8 coding tree block 918 partitioned into three tiles.
[0051] Each coding tree block 918 of each coding 898 is assigned a partition that signals to identify the block size for intra or inter prediction and for transform coding. The partitioning is a recursive quad tree partitioning. The root of the quad tree is associated with the coding tree block. The quad tree is divided until it reaches a leaf, which is called a coding block. A coding block is the root node of two trees, a prediction tree and a transform tree.
[0052] The prediction tree specifies the position and size of prediction blocks. Prediction blocks and associated prediction data are called prediction units.
[0053] FIG. 11 shows an exemplary sequence parameter set RBSP syntax.
[0054] The transform tree specifies the position and size of transform blocks. The transform blocks and associated transform data are referred to as transform units.
[0055] The partitioning information for luma and chroma is the same as the prediction tree and may or may not be the same as the transform tree.
[0056] The coded blocks, associated coded data, associated prediction and transform units together form a coding unit.
[0057] The process for converting the coded tree block addresses in raster order to tile scan order may be as follows. The output of this process is - an array CtbAddrTS[ctbAddrRS] having ctbAddrRS in the range from 0 to PicHeightInCtbs*PicWidthInCtbs-1, - an array TileId[ctbAddrTS] having ctbAddrTS in the range from 0 to PicHeightInCtbs*PicWidthInCtbs-1. The array CtbAddrTS[] is derived as follows. TIFF2025093939000002.tif62155
[0058] The array TileId[] is derived as follows: TIFF2025093939000003.tif29134
[0059] The corresponding exemplary syntax is shown in FIGS. 11, 12 and 13, where FIG. 12 has an exemplary picture parameter set RBSP syntax. FIG. 13 shows an exemplary slice header syntax.
[0060] In the syntax example, the following semantics may apply:
[0061] An entropy_slice_flag equal to 1 indicates that the value of a non-existing slice header syntax element is inferred to be equal to the value of the slice header syntax element in the current slice, and the current slice is defined as a slice containing a coded tree block having a position (SliceCtbAddrRS-1). The entropy_slice_flag is equal to 0 when SliceCtbAddrRS is equal to 0.
[0062] A tiles_or_entropy_coding_sync_idc equal to 0 indicates that there is only one tile in each picture in the coded video sequence, and further, a specific synchronization process for context variables is not called before decoding the first coded tree block in a row of coded tree blocks.
[0063] A tiles_or_entropy_coding_sync_idc equal to 1 indicates that there may be more than one tile in each picture in the coded video sequence, and further, a specific synchronization process for context variables is not called before decoding the first coded tree block in a row of coded tree blocks.
[0064] A tiles_or_entropy_coding_sync_idc equal to 2 indicates that there is only one tile in each picture in the coded video sequence, a specific synchronization process for context variables is called before decoding the first coded tree block in a row of coded tree blocks, and further, a specific storage process for context variables is called after decoding two coded tree blocks in a row of coded tree blocks.
[0065] The value of tiles_or_entropy_coding_sync_idc is in the range of 0 to 2.
[0066] num_tile_columns_minus1 + 1 indicates the number of tile columns that partition the image.
[0067] num_tile_rows_minus1 + 1 indicates the number of tile rows that partition the image. When num_tile_columns_minus1 is equal to 0, num_tile_rows_minus1 is not equal to 0.
[0068] One or both of the following situations are satisfied for each slice and tile: - All the encoded blocks in a slice belong to the same tile. - All the encoded blocks in a tile belong to the same slice.
[0069] Note - In the same image, there may be both slices containing multiple tiles and tiles containing multiple slices.
[0070] A uniform_spacing_flag equal to 1 indicates that the column boundaries and likewise the row boundaries are uniformly distributed across the image. A uniform_spacing_flag equal to 0 indicates that the column boundaries and likewise the row boundaries are not uniformly distributed across the image but are explicitly indicated using the syntax elements column_width[i] and row_height[i].
[0071] column_width[i] indicates the width of the i-th tile column in units of coded tree blocks.
[0072] row #The value of height[i], which specifies the height of the i-th tile row in units of coding tree blocks, the value of ColumnWidth[i], which specifies the width of the i-th tile column in units of coding tree blocks, and the value of ColumnWidthInLumaSamples[i], which specifies the width of the i-th tile column in units of luma samples, are derived as follows: TIFF2025093939000004.tif40155
[0073] The value of RowHeight[i], which specifies the height of the i-th tile row in units of coding tree blocks, is derived as follows: TIFF2025093939000005.tif34161
[0074] The value of ColBd[i], which specifies the position of the left column boundary of the i-th tile column in units of coding tree blocks, is derived as follows: TIFF2025093939000006.tif13140
[0075] The value of RowBd[i], which specifies the position of the top row boundary of the i-th tile row in units of coding tree blocks, is derived as follows: TIFF2025093939000007.tif12140
[0076] num_substreams_minus1 + 1 specifies the maximum number of subsets included in a slice when tiles_or_entropy_coding_sync_idc is equal to 2. When it does not exist, the value of num_substreams_minus1 is presumed to be equal to 0.
[0077] num_entry_point_offsets indicates the number of entry_point_offset[i] syntax elements in the slice header. When tiles_or_entropy_coding_sync_idc is equal to 1, the value of num_entry_point_offsets is in the range from 0 to (num_tile_columns_minus1+1)*(num_tile_rows_minus1+1)-1. When tiles_or_entropy_coding_sync_idc is equal to 2, the value of num_entry_point_offsets is in the range from 0 to num_substreams_minus1. When it does not exist, the value of num_entry_point_offsets is presumed to be equal to 0.
[0078] offset_len_minus1+1 indicates the length of the entry_point_offset[i] syntax element in bits.
[0079] entry_point_offset[i] indicates the i-th entry point offset in bytes and is further represented by offset_len_minus1+1 bits. The coded slice NAL unit consists of num_entry_point_offsets+1 subsets having subset index values in the range from 0 to num_entry_point_offsets. Subset 0 consists of bytes from byte 0 to entry_point_offset[0]-1 of the coded slice NAL unit, subset k having k in the range from 1 to num_entry_point_offsets-1 consists of bytes from entry_point_offset[k-1] to entry_point_offset[k]+entry_point_offset[k-1]-1 of the coded slice NAL unit, and further, the last subset (having a subset index equal to num_entry_point_offsets) consists of the remaining bytes of the coded slice NAL unit.
[0080] Note - The NAL unit header and the slice header of the encoded slice NAL unit are always included in subset 0.
[0081] When tiles_or_entropy_coding_sync_idc is equal to 1 and further num_entry_point_offsets is greater than 0, each subset contains all the encoded bits of one or more complete tiles, and further the number of subsets is less than or equal to the number of tiles in the slice. When tiles_or_entropy_coding_sync_idc is equal to 2 and further num_entry_point_offsets is greater than 0, subset k contains all the bits used during the initialization process for the current bitstream pointer k for each of all possible k values.
[0082] Regarding slice data semantics, the following may apply.
[0083] An end_of_slice_flag equal to 0 indicates that another macroblock follows in the slice. An end_of_slice_flag equal to 1 indicates the end of the slice and further no additional macroblocks follow.
[0084] entry_point_marker_two_3bytes is a 3-byte fixed value sequence equal to 0x000002. This syntax element is called the entry marker prefix.
[0085] tile_idx_minus_1 indicates the TileID in raster scan order. The first tile in the image has a TileID of 0. The value of tile_idx_minus_1 is in the range from 0 to (num_tile_columns_minus1 + 1)*(num_tile_rows_minus1 + 1)-1.
[0086] The CABAC parsing process for slice data can be as follows:
[0087] This process is called when parsing a syntax element having a descriptor ae(v).
[0088] The input to this process is a request for the value of the syntax element and the values of previously parsed syntax elements.
[0089] The output of this process is the value of the syntax element.
[0090] When starting to parse the slice data of a slice, an initialization process of the CABAC parsing process is called. When tiles_or_entropy_coding_sync_idc is equal to 2 and furthermore num_substreams_minus1 is greater than 0, a mapping table BitStreamTable having num_substreams_minus1 + 1 entries for specifying the bitstream pointer table for use in subsequent current bitstream pointer derivation is derived as follows. - BitStreamTable[0] contains the initialized bitstream pointer. - For all indexes i greater than 0 and further less than num_substreams_minus1 + 1, BitStreamTable[i] contains a bitstream pointer to entry_point_offset[i] bytes after BitStreamTable[i - 1]. The current bitstream pointer is set to BitStreamTable[0].
[0091] The minimum coded block address of the coded tree block containing the spatially adjacent block T, ctbMinCbAddrT, is derived, for example, as follows, using the position (x0, y0) of the top - left luma sample of the current coded tree block. x = x0 + 2 << Log2CtbSize - 1 y = y0 - 1 ctbMinCbAddrT = MinCbAddrZS[x >> Log2MinCbSize][y >> Log2MinCbSize]
[0092] The variable availableFlagT is obtained by calling an appropriate coded block availability derivation process having ctbMinCbAddrT as input. Start analyzing the coding tree, and further, when tiles_or_entropy_coding_sync_idc is equal to 2 and further num_substreams_minus1 is greater than 0, the following applies. - When CtbAddrRS % PicWidthInCtbs is equal to 0, the following applies. - When availableFlagT is equal to 1, the synchronization process of the CABAC analysis process is called as specified in the dependent section "Synchronization Process for Context Variables". - The decoding process for the binary decision before termination is called, followed by the initialization process for the arithmetic decoding engine. - The current bitstream pointer is set to indicate BitStreamTable[i] having an index i derived as follows. i = (CtbAddrRS / PicWidthInCtbs) % (num_substreams_minus1 + 1) - Otherwise, when CtbAddrRS % PicWidthInCtbs is equal to 2, the storage process of the CABAC analysis process is called as specified in the dependent section "Storage Process for Context Variables".
[0093] The initialization process can be as follows:
[0094] The output of this process is the initialized CABAC internal variables.
[0095] Its special process is called when starting the analysis of the slice data of the slice or when starting the analysis of the data of the coding tree and the coding tree is the first coding tree in the tile.
[0096] The storage process for the context variables can be as follows:
[0097] The input of this process is the CABAC context variable indexed by ctxIdx.
[0098] The output of this process is the variables TableStateSync and TableMPSSync that contain the values of the variables m and n used in the initialization process of the context variables assigned to the syntax elements except for the slice end flag. For each context variable, the corresponding entries n and m in the tables TableStateSync and TableMPSSync are initialized to the corresponding pStateIdx and valMPS.
[0099] The synchronization process for the context variables can be as follows:
[0100] The input of this process is the variables TableStateSync and TableMPSSync that contain the values of the variables n and m used in the storage process of the context variables assigned to the syntax elements except for the slice end flag.
[0101] The output of this process is the CABAC context variable indexed by ctxIdx.
[0102] For each context variable, the corresponding context variables pStateIdx and valMPS are initialized to the corresponding entries n and m in the tables TableStateSync and TableMPSSync.
[0103] In the following, low-latency encoding and transmission using WPP are described. In particular, the following description reveals how low-latency transmission as depicted in FIG. 7 can be applied to WPP.
[0104] First and foremost, it is important that a subset of the image can be sent before the completion of the entire image. Usually, this can be achieved using slices, as already shown in FIG. 5.
[0105] To reduce latency compared to tiles, as shown in the following figures, it is necessary to apply a single WPP sub-bitstream per row of LCU and further enable separate transmission of each of those rows. To keep the coding efficiency high, slices per row / sub-stream cannot be used. Therefore, in the following, so-called dependent slices as defined in the next section are introduced. This slice has fields that are used, for example, not in all fields of a complete HEVC slice header but for entropy slices. Further, there may be a switch to turn off the destruction of CABAC between rows. In the case of WPP, the use of CABAC contexts (arrows in FIG. 14) and prediction of rows are enabled to maintain the coding efficiency gain of WPP on the tile.
[0106] In particular, FIG. 14 illustrates image 10 for WPP for regular slice 900 (regular SL) and for low-latency processing for dependent slice (DS) 920.
[0107] Currently, the current HEVC standard provides two types of partitionings with respect to slices. There are regular (normal) slices and entropy slices. A regular slice is a completely independent image partition except for some dependencies that can be used to unblock the filter process at the slice boundary. An entropy slice is independent only with respect to entropy coding. The idea of Figure 14 is to generalize the slicing concept. Thus, the current HEVC standard should provide two general types of slices: independent (regular) or dependent. Therefore, a new type of slice, the dependent slice, is introduced.
[0108] A dependent slice is a slice that has a dependency on the previous slice. The dependency is specific data that can be used between slices in the entropy decoding process and / or the pixel reconstruction process.
[0109] In Figure 14, the concept of a dependent slice is exemplarily shown. The image always starts, for example, with a regular slice. It should be noted that in this concept, the regular slice behavior is slightly changed. Typically, in a standard such as H264 / AVC or HEVC, a regular slice is a completely independent partition and furthermore, after decoding, does not need to retain any data except for some data for unblocking the filter process. However, the processing of the next dependent slice 920 is only possible by referring to the data of the above-mentioned slice, here the regular slice 900 in the first row. To establish this, the regular slice 900 should retain the data of the last CU row. This data is - CABAC coding engine data (the context model state of one CU from which the entropy decoding process of the dependent slice can be initialized), - all decoded syntax elements of the CU for the regular CABAC decoding process of the dependent CU, - Data for intra and motion vector prediction is included.
[0110] Consequently, each dependent slice 920 should perform the same procedure of keeping data for the next dependent slice in the same picture.
[0111] In fact, these additional steps should not be a problem because the decoding process is generally always forced to store some data such as syntax elements.
[0112] In the following section, possible changes to the HEVC standard syntax that are required to enable the concept of dependent slices are shown.
[0113] FIG. 5 shows possible changes, for example, in the picture parameter set RBSP syntax.
[0114] The picture parameter set semantics for dependent slices can be as follows:
[0115] A dependent_slices_present_flag equal to 1 indicates that the picture contains dependent slices, and further, the decoding process of each (regular or dependent) slice should store the entropy decoding state and the data for intra and motion vector prediction for the next slice which may be a dependent slice following a regular slice. The following dependent slices can refer to the stored data.
[0116] FIG. 16 shows a possible slice_header syntax with changes related to the current situation of HEVC.
[0117] A dependent_slice_flag equal to 1 indicates that the value of a non-existing slice header syntax element is assumed to be equal to the value of the slice header syntax element in a progressive (regular) slice, and a progressive slice is defined as a slice that includes coded tree blocks having positions (SliceCtbAddrRS-1). The dependent_slice_flag is equal to 0 when SliceCtbAddrRS is equal to 0.
[0118] A no_cabac_reset_flag equal to 1 indicates CABAC initialization from the saved state of a previously decoded slice (having no initial value). Otherwise, i.e., when equal to 0, CABAC initialization is specified independently of the state of any previously decoded slice, i.e., having an initial value.
[0119] A last_ctb_cabac_init_flag equal to 1 indicates CABAC initialization from the saved state of the last coded tree block of a previously decoded slice (e.g., for a tile that is always equal to 1). Otherwise (when equal to 0), the initialization data is referenced from the saved state of the second coded tree block of the last (adjacent) ctb-row of the previously decoded slice when the first coded tree block of the current slice is the first coded tree block in the row (i.e., in WPP mode), otherwise, the CABAC initialization is preformed from the saved state of the last coded tree block of the previously decoded slice.
[0120] A comparison of dependent slices and other partitioning schemes (information) is provided below.
[0121] In Figure 17, the differences between normal and dependent slices are shown.
[0122] As shown with respect to FIG. 18, the possible encoding and transmission of WPP sub-streams in the dependent slice (DS) is comparable to the encoding for low-latency transmission of tiles (left) and WPP / DS (right). The thick continuous crosses depicted in FIG. 18 indicate the same time points for the two methods assuming that the encoding of the WPP rows takes the same time as the encoding of a single tile. Due to encoding dependencies, only the first row of the WPP is prepared after all tiles have been encoded. However, using the dependent slice approach, the WPP approach allows the first row to be sent as soon as it is encoded. This is different from the initial WPP sub-stream assignment, where a “sub-stream” is defined for WPP as the concatenation of CU rows of slices that are decoded by the same decoder thread, i.e., the same core / processor. However, sub-streams for each row and each entropy slice are also possible before the entropy slice breaks the entropy encoding dependency and thus has low encoding efficiency, i.e., the WPP efficiency gain is lost.
[0123] In addition, the delay difference between the two approaches can be made actually low assuming a transmission as shown in FIG. 19. In particular, FIG. 19 shows WPP encoding with pipeline low-latency transmission.
[0124] Assuming that in FIG. 18 the encoding of two CUs after DS#1.1 in the WPP approach is not longer than the transmission of the first row SL#1, there is no difference between tiles and WPP in the case of low latency. However, the encoding efficiency of WP / DS is performance superior to the tile concept.
[0125] To increase the robustness for the WPP low-latency mode, FIG. 20 shows that the robustness improvement is achieved by using a regular slice (RS) as an anchor. In the image shown in FIG. 20, a dependent slice (DS) follows the (regular) slice (RS). Here, the (regular) slice acts as an anchor that breaks the dependency on the previous slice, and thus more robustness is provided at such insertion points of the (regular) slice. In principle, this is no different from inserting a (regular) slice anyway.
[0126] The concept of the dependent slice can also be implemented as follows.
[0127] Here, FIG. 21 shows a possible slice header syntax.
[0128] The slice header semantics are as follows:
[0129] A dependent_slice_flag equal to 1 indicates that the value of each slice header syntax element that does not exist is presumed to be equal to the value of the corresponding slice header syntax element in the previous slice that contains the coded tree block whose coded tree block address is SliceCtbAddrRS - 1. When it does not exist, the value of the dependent_slice_flag is presumed to be equal to 0. The value of the dependent_slice_flag is equal to 0 when SliceCtbAddrRS is equal to 0.
[0130] slice_address indicates the address at the slice granularity resolution where the slice starts. The length of the slice_address syntax element is (Ciel(Log2(PicWidthInCtbs*PicHeightInCtbs)) + SliceGranularity) bits.
[0131] The variable SliceCtbAddrRS, which indicates the coding tree block at which a slice starts in the coding tree block raster scan order, is derived as follows. SliceCtbAddrRS=(slice_address>>SliceGranularity)
[0132] The variable SliceCbAddrZS, which indicates the address of the first coding block in a slice at the minimum coding block granularity in the z-scan order, is derived as follows. SliceCbAddrZS=slice_address <<((log2_diff_max_min_coding_block_size-SliceGranularity)<<1)
[0133] Slice decoding starts with the largest possible coding unit, or in other terms, at the slice start coordinate, with a CTU.
[0134] first_slice_in_pic_flag indicates whether a slice is the first slice of an image. If first_slice_in_pic_flag is equal to 1, both the variables SliceCbAddrZS and SliceCtbAddrRS are set to 0, and furthermore, decoding starts with the first coding tree block in the image.
[0135] pic_parameter_set_id indicates the picture parameter set in use. The value of pic_parameter_set_id is in the range from 0 to 255.
[0136] num_entry_point_offsets indicates the number of entry_point_offset[i] syntax elements in the slice header. When tiles_or_entropy_coding_sync_idc is equal to 1, the value of num_entry_point_offsets is in the range from 0 to (num_tile_columns_minus1+1)*(num_tile_rows_minus1+1)-1. When tiles_or_entropy_coding_sync_idc is equal to 2, the value of num_entry_point_offsets is in the range from 0 to PicHeightInCtbs-1. When it does not exist, the value of num_entry_point_offsets is presumed to be equal to 0.
[0137] offset_len_minus1+1 indicates the length of the entry_point_offset[i] syntax element in bits.
[0138] entry_point_offset[i] indicates the i-th entry point offset in bytes and is further represented by offset_len_minus1+1 bits. The coded slice data after the slice header consists of num_entry_point_offsets+1 subsets having subset index values in the range from 0 to num_entry_point_offsets. Subset 0 consists of bytes 0 to entry_point_offset[0]-1 of the coded slice data, subset k in the range from 1 to num_entry_point_offsets-1 consists of bytes entry_point_offset[k-1] to entry_point_offset[k]+entry_point_offset[k-1]-1 of the coded slice data, and the last subset (having a subset index equal to num_entry_point_offsets) consists of the remaining bytes of the coded slice data.
[0139] When tiles_or_entropy_coding_sync_idc is equal to 1 and further when num_entry_point_offsets is greater than 0, each subset contains all the coded bits of exactly one tile, and further, the number of subsets (i.e., the value of num_entry_point_offsets + 1) is less than or equal to the number of tiles in the slice.
[0140] Note - When tiles_or_entropy_coding_sync_idc is equal to 1, each slice must contain a subset of one tile (when entry point signaling is not required) or an integer number of complete tiles.
[0141] When tiles_or_entropy_coding_sync_idc is equal to 2 and further when num_entry_point_offsets is greater than 0, each subset k having k in the range from 0 to num_entry_point_offsets - 1 contains all the coded bits of exactly one row of the coding tree block, and the last subset (having subset index equal to num_entry_point_offsets) contains all the coded bits of the remaining coding blocks included in the slice, and the remaining coding blocks consist of exactly one row of the coding tree block or a subset of one row of the coding tree block, and further, the number of subsets (i.e., the value of num_entry_point_offsets + 1) is equal to the number of rows of the coding tree block in the slice, and a subset of one row of the coding tree block in the slice is also counted.
[0142] Note that when tiles_or_entropy_coding_sync_idc is equal to 2, a slice may include a number of lines of coding tree blocks and a subset of the lines of coding tree blocks. For example, if a slice includes 2.5 lines of coding tree blocks, the number of subsets (i.e., the value of num_entry_point_offsets + 1) is equal to 3.
[0143] The corresponding picture parameter set RBSP syntax can be selected as shown in Figure 22.
[0144] The picture parameter set RBSP semantics can be as follows:
[0145] A dependent_slice_enabled_flag equal to 1 indicates the presence of the syntax element dependent_slice_flag in the slice header for the coded picture that references the picture parameter set. A dependent_slice_enabled_flag equal to 0 indicates the absence of the syntax element dependent_slice_flag in the slice header for the coded picture that references the picture parameter set. When tiles_or_entropy_coding_sync_idc is equal to 3, the value of the dependent_slice_enabled_flag is equal to 1.
[0146] A tiles_or_entropy_coding_sync_idc equal to 0 indicates that there is only one tile in each picture that references the picture parameter set, there is no specific synchronization process for the context variables called before decoding the first coding tree block of the lines of coding tree blocks in each picture that references the picture parameter set, and furthermore, the values of cabac_independent_flag and dependent_slice_flag for the coded picture that references the picture parameter set are not both equal to 1.
[0147] Note that when both the cabac_independent_flag and the depedent_slice_flag are equal to 1 for a slice, the slice is an entropy slice.
[0148] A tiles_or_entropy_coding_sync_idc equal to 1 indicates that there may be more than one tile in each image referring to the picture parameter set, there is no specific synchronization process for context variables called before decoding the first coded tree block of the rows of coded tree blocks in each image referring to the picture parameter set, and further, the values of the cabac_independent_flag and the dependent_slice_flag for the coded image referring to the picture parameter set are not both equal to 1.
[0149] A tiles_or_entropy_coding_sync_idc equal to 2 indicates that there is only one tile in each image referring to the picture parameter set, a specific synchronization process for context variables is called before decoding the first coded tree block of the rows of coded tree blocks in each image referring to the picture parameter set, and further, a specific storage process for context variables is called after decoding two coded tree blocks of the rows of coded tree blocks in each image referring to the picture parameter set, and further, the values of the cabac_independent_flag and the dependent_slice_flag for the coded image referring to the picture parameter set are not both equal to 1.
[0150] A tiles_or_entropy_coding_sync_idc equal to 3 specifies that there is only one tile in each image that refers to the picture parameter set, and there is no specific synchronization process for context variables that are called before decoding the first coding tree block of the rows of coding tree blocks in each image that refers to the picture parameter set. Further, the values of cabac_independent_flag and dependent_slice_flag for the coded pictures that refer to the picture parameter set may both be equal to 1.
[0151] When the dependent_slice_enabled_flag is equal to 0, the tiles_or_entropy_coding_sync_idc is not equal to 3.
[0152] It is a bitstream compliance requirement that the value of tiles_or_entropy_coding_sync_idc is for all picture parameter sets activated within the coded video sequence.
[0153] For each slice that refers to the picture parameter set, when the tiles_or_entropy_coding_sync_idc is equal to 2, and further, when the first coded block in the slice is not the first coded block in the first coding tree block of the rows of coding tree blocks, the last coded block in the slice belongs to the same row of coding tree blocks as the first coded block in the slice.
[0154] num_tile_columns_minus1 + 1 indicates the number of tile columns that partition the image. num_tile_rows_minus1 + 1 indicates the number of tile rows that partition the image.
[0155] When num_tile_columns_minus1 is equal to 0, num_tile_rows_minus1 is not equal to 0. A uniform_spacing_flag equal to 1 indicates that the column boundaries and, similarly, the row boundaries are uniformly distributed across the image. A uniform_spacing_flag equal to 0 indicates that the column boundaries and, similarly, the row boundaries are not uniformly distributed across the image but are explicitly indicated using the syntax elements column_width[i] and row_height[i].
[0156] column_width[i] indicates the width of the i-th tile column in units of coded tree blocks.
[0157] row_height[i] indicates the height of the i-th tile row in units of coded tree blocks.
[0158] The vector colWidth[i] indicates the width of the i-th tile column in units of CTBs having column i in the range from 0 to num_tile_columns_minus1.
[0159] The vector CtbAddrRStoTS[ctbAddrRS] indicates the mapping from the CTB address in raster scan order to the CTB address in tile scan order for an index ctbAddrRS in the range from 0 to (picHeightInCtbs * picWidthInCtbs) - 1.
[0160] The vector CtbAddrTStoRS[ctbAddrTS] indicates the mapping from the CTB address in tile scan order to the CTB address in raster scan order for an index ctbAddrTS in the range from 0 to (picHeightInCtbs * picWidthInCtbs) - 1.
[0161] The vector TileId[ctbAddrTS] indicates the mapping from the CTB address in tile scan order, which has ctbAddrTS in the range from 0 to (picHeightInCtbs * picWidthInCtbs) - 1, to the tile id.
[0162] The values of colWidth, CtbAddrRStoTS, CtbAddrTStoRS, and TileId are derived by calling the tile scanning mapping process with the CTB raster and PicHeightInCtbs and PicWidthInCtbs as inputs, and further, the output is assigned to colWidth, CtbAddrRStoTS, and TileId.
[0163] The value of ColumnWidthInLumaSamples[i], which indicates the width of the i-th tile column in units of luma samples, is set equal to colWidth[i], where colWidth[i] << Log2CtbSize.
[0164] The array MinCbAddrZS[x][y], which indicates the mapping from the location (x, y) in units of the minimum CB, where x is in the range from 0 to picWidthInMinCbs - 1 and y is in the range from 0 to picHeightInMinCbs - 1, to the minimum CB address in z scan order, is derived by calling the Z scanning order array initialization process with Log2MinCbSize, Log2CtbSize, PicHeightInCtbs, PicWidthInCtbs, and the vector CtbAddrRStoTS as inputs, and further, the output is assigned to MinCbAddrZS.
[0165] The loop_filter_across_tiles_enabled_flag equal to 1 indicates that the loop filter operation is executed across tile boundaries. The loop_filter_across_tiles_enabled_flag equal to 0 indicates that the loop filter operation is not executed across tile boundaries. The loop filter operation includes the deblocking filter, sample adaptive offset, and adaptive loop filter operations. When it does not exist, the value of the loop_filter_across_tiles_enabled_flag is assumed to be equal to 1.
[0166] The cabac_independent_flag equal to 1 indicates that the CABAC decoding of the coded blocks in a slice is independent of any state of the previously decoded slices. The cabac_independent_flag equal to 0 indicates that the CABAC decoding of the coded blocks in a slice is dependent on the state of the previously decoded slices. When it does not exist, the value of the cabac_independent_flag is assumed to be equal to 0.
[0167] The derivation process for the availability of the coded block with the minimum coded block address can be as follows:
[0168] The input to this process is - the minimum coded block address minCbAddrZS in z-scan order - the current minimum coded block address currMinCBAddrZS in z-scan order is.
[0169] The output of this process is the availability of the coded block with the coded block address cbAddrZS in z-scan order cbAvailable.
[0170] Note 1 - The meaning of availability is determined when this process is called. Note 2 - Regardless of its size, any coding block is associated with a minimum coding block address, which is the address of the coding block having the minimum coding block size in the z-scan order. - If one or more of the following conditions are true, cbAvailable is set to false. - minCbAddrZS is less than 0 - minCbAddrZS is greater than currMinCBAddrZS - The coding block having the minimum coding block address minCbAddrZS belongs to a different slice than the coding block having the current minimum coding block address currMinCBAddrZS, and furthermore, the dependent_slice_flag of the slice containing the coding block having the current minimum coding block address currMinCBAddrZS is equal to 0. - The coding block having the minimum coding block address minCbAddrZS is included in a different tile than the coding block having the current minimum coding block address currMinCBAddrZS. - Otherwise, cbAvailable is set to true.
[0171] The CABAC parsing process for slice data can be as follows:
[0172] This process is called when parsing a specific syntax element having the descriptor ae(v).
[0173] The input to this process is the value of the syntax element and the request for the values of the previously parsed syntax elements.
[0174] The output of this process is the value of the syntax element.
[0175] When starting to parse the slice data of a slice, the initialization process of the CABAC parsing process is called.
[0176] Figure 23 shows how spatial neighbor T is used to call the coding tree block availability derivation process related to the current coding tree block (information).
[0177] The minimum coding block address of the coding tree block containing the spatial neighbor block T (Figure 23), ctbMinCbAddrT, is derived as follows using the position (x0, y0) of the top - left luma sample of the current coding tree block: x = x0+2<<Log2CtbSize - 1 y = y0 - 1 ctbMinCbAddrT = MinCbAddrZS[x>>Log2MinCbSize][y>>Log2MinCbSize]
[0178] The variable availableFlagT is obtained by calling the coding block availability derivation process with ctbMinCbAddrT as input.
[0179] When starting the analysis of the coding tree as specified, the following ordered steps are applied.
[0180] The arithmetic decoding engine is initialized as follows.
[0181] If CtbAddrRS is equal to slice_address, dependent_slice_flag is equal to 1, and further, entropy_coding_reset_flag is equal to 0, the following applies. The synchronization process of the CABAC analysis process is called with TableStateIdxDS and TableMPSValDS as input. The decoding process for the binary decision before termination is called, followed by the initialization process for the arithmetic decoding engine.
[0182] Otherwise, if tiles_or_entropy_coding_sync_idc is equal to 2 and further CtbAddrRS % PicWidthInCtbs is equal to 0, the following applies. When availableFlagT is equal to 1, the synchronization process of the CABAC parsing process is called with TableStateIdxWPP and TableMPSValWPP as inputs. The decoding process for the binary decision before termination is called, followed by the process for the arithmetic decoding engine.
[0183] When cabac_independent_flag is equal to 0, further when dependent_slice_flag is equal to 1, or when tiles_or_entropy_coding_sync_idc is equal to 2, the memory process is applied as follows. When tiles_or_entropy_coding_sync_idc is equal to 2 and further CtbAddrRS % PicWidthInCtbs is equal to 2, the memory process of the CABAC parsing process is called with TableStateIdxWPP and TableMPSValWPP as outputs. When cabac_independent_flag is equal to 0, dependent_slice_flag is equal to 1, and further end_of_slice_flag is equal to 1, the memory process of the CABAC parsing process is called with TableStateIdxDS and TableMPSValDS as outputs.
[0184] The parsing of syntax elements proceeds as follows:
[0185] For each requested value of the syntax element, a binarization is derived.
[0186] The binarization for the syntax element and the sequence of parsed bins determines the decoding process flow.
[0187] For each bin of the binarization of a syntax element, which is indexed by the variable binIdx, a context index ctxIdx is derived.
[0188] For each ctxIdx, an arithmetic decoding process is invoked.
[0189] The sequence (b0..bbinIdx) resulting from the parsed bin is compared to the set of bin strings given by the binarization process after decoding of each bin. When the sequence matches a bin string in the given set, the corresponding value is assigned to the syntax element.
[0190] Requirements for the value of the syntax element are processed for the syntax element pcm-flag, and further, if the decoded value of pcm_flag is equal to 1, the decoding engine is initialized after decoding of any pcm_alignment_zero_bit, num_subsequent_pcm, all pcm_sample_luma and pcm_sample_chroma data.
[0191] As described above, the above explanation reveals a decoder as shown in FIG. 24. This decoder, generally denoted by reference numeral 5, reconstructs the image 10 from the data stream 12 in which the image 10 is encoded in units of slices 14 that are partitioned, and the decoder 5 is configured to decode the slices 14 from the data stream 12 according to the slice order 16. Of course, the decoder 5 is not limited to decoding the slices 14 sequentially. Rather, the decoder 5 can use wavefront parallel processing to decode the slices 14 on the condition that the image 10 partitioned into the slices 14 is suitable for wavefront parallel processing. Therefore, the decoder 5 can start decoding the slices 14 by considering the slice order 16 to enable wavefront processing, for example, as described above and further described below, and can decode the slices 14 in an alternating parallel manner, and may be a decoder.
[0192] Decoder 5 responds to the syntax element portion 18 within the current slice of slice 14 to decode the current slice according to at least one of two modes 20 and 22. According to the first mode of the at least two modes, namely mode 20, the current slice is decoded from the data stream 12 using context adaptive entropy decoding including context derivation that crosses slice boundaries, i.e., crosses the dotted line in FIG. 24, i.e., using information resulting from the encoding / decoding of other "slices previous in slice order 16". Further, decoding of the current slice from the data stream 12 using the first mode 20 includes continuous updating of the codec symbol probabilities and initialization of the symbol probabilities at the start of decoding of the current slice that is dependent on the saved state of the symbol probabilities of the previously decoded slices. Such dependency is described above in connection with, for example, "synchronization processes for codec variables". Finally, the first mode 20 also includes predictive decoding that crosses slice boundaries. Such predictive decoding that crosses slice boundaries can include, for example, intra prediction across the slice boundary, i.e., predictive sample values within the current slice based on the already reconstructed sample values of the slice previous in "slice order 16", or prediction of coding parameters that cross slice boundaries such as, for example, prediction of motion vectors, prediction modes, coding modes, etc.
[0193] According to the second mode 22, the decoder 5 uses context-adaptive entropy decoding, but limits the derivation of context so as not to cross slice boundaries, and decodes the current slice, i.e., the slice currently being decoded, from the data stream 12. For example, if a template of adjacent positions used to derive the context for a particular syntax element related to a block within the current slice extends to an adjacent slice, thereby crossing the slice boundary of the current slice, the corresponding attributes of each part of the adjacent slice, such as the value of the corresponding syntax element of this adjacent part of the adjacent slice, are set to default values to suppress the mutual dependency attributes between the current slice and the adjacent slice. While the continuous update of the context symbol probability can occur as it does in the first mode 20, the initialization of the symbol probability in the second mode 22 is independent of any previously decoded slice. Further, predictive decoding is performed with the predictive decoding restricted so as not to cross slice boundaries.
[0194] To facilitate the understanding of the description of FIG. 24 and the following description, refer to FIG. 25 which shows a possible implementation of the decoder 5 in a more structural sense compared to FIG. 24. As it is in FIG. 24, the decoder 5 is, for example, a predictive decoder that uses context-adaptive entropy decoding to decode the data stream to obtain, for example, prediction residuals and prediction parameters.
[0195] As shown in FIG. 25, the decoder 5 can include an entropy decoder 24, an inverse quantization and inverse transformation module 26, and a combiner 28 implemented, for example, as an adder and a predictor 28 as shown in FIG. 25. The entropy decoder 24, the module 26, and the adder 27 are connected in series between the input and output of the decoder 5 in the order in which they are mentioned. Further, the predictor 28 is connected between the output of the adder 28 and its further input to form a prediction loop together with the combiner 27. Therefore, the decoder 24 has its output further connected to the encoded parameter input of the predictor 28.
[0196] Although FIG. 25 gives the impression that the decoder decodes the current image serially, the decoder 5 can be implemented to decode the image 10 in parallel, for example. The decoder 5 can include, for example, a plurality of cores each operating according to elements 24-28 in FIG. 25. However, parallel processing is optional, and furthermore, the decoder 5 operating serially can also decode the data stream entering the input of the entropy decoder 24.
[0197] To efficiently achieve the above-described ability to decode the current image 10 serially or in parallel, the decoder 5 operates in units of encoding blocks 30 to decode the image 10. The encoding block 30 is, for example, a leaf block in which an encoding tree block or a maximum encoding block 32 is partitioned by a recursive multi-tree partition such as a quad-tree partition. Next, the code tree blocks 32 can be regularly arranged in columns and rows to form a regular partition of the image 10 into these code tree blocks 32. In FIG. 25, the code tree blocks 32 are shown by solid lines, while the encoding blocks 30 are shown by dotted lines. For illustrative purposes, only one code tree block 32 is shown to be further partitioned into encoding blocks 30, while the other code tree blocks 32 are shown not to be further partitioned to directly form encoding blocks. The data stream 12 can include a syntax portion signaling how the image 10 is partitioned into code blocks 30.
[0198] For each encoding block 30, the data stream 12 conveys syntax elements that clarify how the modules 24-28 recover the image content within that encoding block 30. For example, these syntax elements can include the following: 1) Optionally, partitioning data for further partitioning the encoded block 30 into prediction blocks. 2) Optionally, partitioning data for further partitioning the encoded block 30 into residual and / or transform blocks. 3) A prediction mode for signaling whether a prediction mode is used to derive a prediction signal for the encoded block 30, wherein the granularity at which this prediction mode is signaled can be dependent on the encoded block 30 and / or the prediction block. 4) Prediction parameters can be signaled for each encoded block or, if present, for each prediction block having certain prediction parameters to be sent, for example, depending on the prediction mode. Possible prediction modes can include, for example, intra prediction and / or inter prediction. 5) Other syntax elements may be present, such as filtering information for filtering the image 10 with the encoded block 30 to obtain, for example, a prediction signal and / or a reconstructed signal to be reproduced. 6) Finally, residual information, particularly in the form of transform coefficients, can be included in the data stream for the encoded block 30. In units of residual blocks, the residual data can be signaled. For each residual block, if present, spectral decomposition can be performed, for example, in units of the above-mentioned transform blocks.
[0199] The entropy decoder 24 serves to obtain the above-described syntax elements from the data stream. For this purpose, the entropy decoder 24 uses context-adaptive entropy decoding. That is, the entropy decoder 24 provides several contexts. To derive a specific syntax element from the data stream 12, the entropy decoder 24 selects a specific context among the possible contexts. The selection among the possible contexts is performed according to the attributes of the neighborhood of the part of the image 10 to which the current syntax element belongs. For each of the possible contexts, the entropy decoder 24 manages the symbol probability, that is, the probability estimation for each possible symbol of the symbol alphabet in which the entropy decoder 24 operates. "Managing" includes the above-described continuous update of the symbol probability of the context in order to apply the symbol probability associated with each context to the actual image context. By this measure, the symbol probability is applied to the actual probability statistics of the symbols.
[0200] Another environment in which the neighboring attributes affect the reconstruction of the current part of the image 10, such as the current coding block 30 for example, is the prediction decoding in the predictor 28. The prediction is not limited to only the predicted content within the current coding block 30, but can also include, for example, parameters contained in the data stream 12 for the current coding block 30 such as prediction parameters, prediction of partitioned data or transform coefficients. That is, the predictor 28 can predict the image content or such parameters from the above-described neighborhood in order to obtain the prediction residual obtained by the module 26 from the data stream 12 and the written signal to be combined therewith later. When predicting parameters, the predictor 28 can use the syntax elements contained in the data stream as prediction residuals to obtain the actual values of the prediction parameters. The predictor 28 uses the later prediction parameter values to obtain the above-described prediction signal to be combined with the prediction residual in the combiner 27.
[0201] The "neighborhood" described above mainly covers the upper left part of the current partial environment to which the syntax element currently undergoing entropy decoding or the syntax element currently being predicted belongs. In FIG. 25, such a neighborhood is exemplarily denoted by 34 for one coding block 30.
[0202] The coding / decoding order is defined within coding block 30: At the coarsest level, the code tree blocks 32 of image 10 are scanned in scan order 36, shown here as a raster scan leading row by row from top to bottom. Within each code tree block, coding block 30 is scanned in depth-first traversal order such that it is also substantially scanned in a raster scan leading row by row from top to bottom by code tree block 32 at each hierarchical level.
[0203] The coding order defined within coding block 30 is consistent with the definition of neighborhood 34 used to derive attributes in the neighborhood in order to select a context and / or perform a spatial prediction in that the neighborhood 34 mainly covers the part of image 10 that has already undergone decoding according to the coding order. Whenever a part of neighborhood 34 covers an unavailable part of image 10, default data is used instead, for example. For example, neighborhood template 34 can extend outside image 10. However, another possibility is that neighborhood 34 extends to adjacent slices.
[0204] Slice division, for example, the image 10 along the encoding / decoding order defined along the encoding block 30, i.e., each slice, is a continuous and non-interrupted sequence of the encoding blocks 30 along the encoding block order described above. In FIG. 25, the slice is indicated by the dashed-dotted line 14. The order defined in the slice 14 results from those configurations of the execution of the continuous encoding blocks 30 as described above. When the syntax element part 18 of a specific slice 14 indicates that it is decoded in the first mode, the entropy decoder 24 enables the context-adaptive entropy decoding to derive a context that crosses the slice boundary. That is, the spatial neighborhood 34 is used to select a context in the entropy decoding data for the current slice 14. In the case of FIG. 25, for example, the slice number 3 may be the currently decoded slice, and further, in the entropy decoding syntax element or some parts included therein with respect to the encoding block 30, the entropy decoder 24 can use the attributes resulting from the decoded parts in adjacent slices such as slice number 1. The predictor 28 operates in the same way: for a slice in the first mode 20, the predictor 28 uses a spatial prediction that crosses the slice boundary surrounding the current slice.
[0205] However, for a slice having a second mode 22 associated therewith, i.e., the syntax element part 18 indicates the second mode 22, the entropy decoder 24 and the predictor 28 limit the derivation of the entropy context and the predictive decoding in order to be subordinate to the attributes related to the parts existing only within the current slice. Obviously, the encoding efficiency is subject to this limitation. On the other hand, the slices of the second mode 22 can break the mutual dependency between the sequences of slices. Therefore, the slices of the second mode 22 can be scattered within the image 10 or within the video to which the image 10 belongs to enable resynchronization points. However, it is not necessary for each image 10 to have at least one slice in the second mode 22.
[0206] As already mentioned above, the first and second modes 20 and 22 also differ in their initialization of symbol probabilities. The slices encoded in the second mode 22 result in an entropy decoder 24 that re-initializes the probabilities independently of any previously decoded slices, i.e., slices decoded earlier in the order defined within the slice. The symbol probabilities are set to default values known on both the encoder and decoder sides, for example, or the initialization values are included within the slices encoded in the second mode 22.
[0207] That is, for the slices encoded / decoded in the second mode 22, the adaptation of the symbol probabilities always starts immediately from the start of these slices. Thus, the adaptation accuracy is not good for these slices at the start of these slices.
[0208] The situation is different for the slices encoded / decoded in the first mode 20. For later slices, the initialization of the symbol probabilities performed by the entropy decoder 24 depends on the saved state of the symbol probabilities of the previously decoded slices. Whenever a slice encoded / decoded in the first mode 20 has its start located, for example, outside the left side of the image 10, i.e., on the side where the raster scan 36 does not start row-by-row before proceeding to the lower side of the next row, the symbol probabilities resulting from the end of the entropy decoding of the immediately preceding slice are adopted. This is shown in FIG. 2 by the arrow 38 for slice number 4, for example. Slice number 4 has its start somewhere between the right and left sides of the image 10 and thus, when initializing the symbol probabilities, the entropy decoder 24 adopts the symbol probabilities obtained in the entropy decoding of the immediately preceding slice, i.e., slice number 3, up to its end, i.e., up to the end of the entropy decoding of slice 3, which includes the continuous update of the symbol probabilities during the entropy decoding of slice 3.
[0209] It has a second mode 22 associated therewith, but slices starting on the left side of the image 10, such as slice number 5, for example, do not apply the symbol probabilities as would be obtained after the entropy decoding of the previous slice number 4 has been completed, because this would prevent the decoder 5 from decoding the image 10 in parallel by using wavefront processing. Instead, as described above, the entropy decoder 24 applies the symbol probabilities as would be obtained after the entropy decoding of the second code tree block 32 in the encoding / decoding order 36 has been completed in the immediately previous code tree block row in the encoding / decoding order 36, as indicated by the arrow 40.
[0210] In FIG. 25, for example, the image 10 is illustratively partitioned into 3 rows of code tree blocks and 4 columns of encoding tree root blocks 32, and further, each row of code tree blocks is subdivided into 2 slices 14, such that the start of all second slices coincides with the first encoding unit in the encoding unit order of each code tree root block row. Thus, the entropy decoder 24 can use wavefront processing in the decoded image 10 by starting from the first or top code tree root block row and starting to decode these code tree root block rows alternately, second and then third, in parallel.
[0211] Of course, the partitioning of the block 32 in a recursive manner to further encoding blocks 30 is optional, and thus, in a more general sense, the block 32 can likewise be referred to as an "encoding block". That is, more generally speaking, the image 10 can be partitioned into encoding blocks 32 arranged in rows and columns and having a raster scan order 36 further defined with respect to each other, and further, the decoder 5 can be considered to associate each slice 14 with a consecutive subset of the encoding blocks 32 in the raster scan order 36 such that the subsets follow each other along the raster scan order 36 in slice order.
[0212] Also, as is apparent from the above description, the decoder 5, or more specifically, the entropy decoder 24 can be configured to store symbol probabilities obtained in context adaptive entropy decoding of any slice up to the second encoded block in the encoded block row according to the raster scan order 36. When initializing the symbol probabilities for context adaptive entropy decoding of the current slice having the associated first mode 20, the decoder 5, or more specifically, the entropy decoder 24 checks whether the first encoded block 32 of a consecutive subset of the encoded blocks 32 associated with the current slice is the first encoded block 32 in the encoded block row according to the raster scan order 36. If so, the symbol probabilities for context adaptive entropy decoding of the current slice are initialized according to the stored symbol probabilities obtained in context entropy decoding of the previously decoded slices up to the second encoded block in the encoded block row according to the raster scan order 36, as explained with respect to arrow 40. If not, the initialization of the symbol probabilities for context adaptive entropy decoding of the current slice is performed according to the symbol probabilities obtained in context adaptive entropy decoding of the previously decoded slices up to the last of the previously decoded slices, i.e., according to arrow 38. Also, in the case of initialization according to 38, it means the state saved at the end of the entropy decoding of the immediately preceding slice in the slice order 36, but in the case of initialization 40, it is that the previously decoded slices include the last of the second block in the row immediately preceding block 32 in the block order 36.
[0213] As indicated by the dotted line in FIG. 24, the decoder may be configured to respond to the syntax element portion 18 within the current slice of slice 14 to decode the current slice according to at least one of three modes. That is, the third mode 42 may be next to the others 20 and 22. The third mode 42 may differ from the second mode 22 in that prediction across slice boundaries is made possible, but the entropy coding / decoding is still restricted so as not to cross slice boundaries.
[0214] Above, two embodiments are shown with respect to the syntax element portion 18. The following table summarizes these two embodiments.
[0215] TIFF2025093939000008.tif60130
[0216] In one embodiment, the syntax element part 18 is individually formed by the dependent_slice_flag, while in other embodiments, the combination of the dependent_slice_flag and the no_cabac_reset_flag forms the syntax element part. Refer to the synchronization process for context variables as far as the initialization of symbol probabilities according to the saved state of symbol probabilities of previously decoded slices is concerned. In particular, when last_ctb_cabac_init_flag = 0 and tiles_or_entropy_coding_sync_idc = 2, the decoder saves the symbol probabilities obtained in the context-adaptive entropy decoding of the previously decoded slices up to the second coded block sequentially according to the raster scan order, and further, when initializing the symbol probabilities for the context-adaptive entropy decoding of the current slice according to the first mode, checks whether the first coded block of the consecutive subset of coded blocks associated with the current slice is the first coded block sequentially according to the raster scan order, and further, if so, initializes the symbol probabilities for the context-adaptive entropy decoding of the current slice according to the saved symbol probabilities obtained in the context-adaptive entropy decoding of the previously decoded slices up to the second coded block sequentially according to the raster scan order, and further, if not, may be configured to initialize the symbol probabilities for the context-adaptive entropy decoding of the current slice according to the symbol probabilities obtained in the context-adaptive entropy decoding of the previously decoded slices up to the last of the previously decoded slices.
[0217] Thus, in other words, according to the second embodiment for syntax, the decoder reconstructs the image 10 from the data stream 12 in which the image is encoded in units of slices 14 into which the image (10) is partitioned, the decoder is configured to decode the slices 14 from the data stream 12 according to the slice order 16, and further, the decoder responds to the syntax element part 18, i.e., the dependent_slice_flag in the current slice of the slice, to decode the current slice according to one of at least two modes 20, 22. According to the first mode 20 of the at least two modes, i.e., when the dependent_slice_flag = 1, the decoder performs context-adaptive entropy decoding 24 including derivation of context across slice boundaries, continuous update of the symbol probabilities of the context, and initialization 38, 40 of the symbol probabilities according to the saved state of the symbol probabilities of the previously decoded slices, and predictive decoding across slice boundaries to decode the current slice from the data stream 12, and further, according to the second mode 22 of the at least two modes, i.e., when the dependent_slice_flag = 0, the decoder restricts the derivation of context so as not to cross slice boundaries, performs context-adaptive entropy decoding with continuous update of the symbol probabilities of the context and initialization of the symbol probabilities independent of any previously decoded slices, and predictive decoding that restricts predictive decoding so as not to cross slice boundaries to decode the current slice from the data stream 12. The image 10 may be partitioned in encoded blocks 32 arranged in rows and columns and having a raster scan order 36 defined with respect to each other, and further, the decoder is configured to associate each slice 14 with a continuous subset of the encoded blocks 32 in the raster scan order 36 such that the subsets follow each other along the raster scan order 36 according to the slice order.The decoder, i.e., in response to tiles_or_entropy_coding_sync_idc = 2, stores symbol probabilities as obtained in context adaptive entropy decoding of a slice decoded up to the second coded block 32 in succession according to raster scan order 36, and further, when initializing symbol probabilities for context adaptive entropy decoding of the current slice according to the first mode, checks whether the first coded block of a consecutive subset of coded blocks 32 associated with the current slice is the first coded block 32 in succession according to raster scan order, and further, if so, initializes the symbol probabilities for context adaptive entropy decoding of the current slice according to the stored symbol probabilities as obtained in context adaptive entropy decoding of a slice decoded up to the second coded block in succession according to raster scan order 36, and further, if not, initializes the symbol probabilities for context adaptive entropy decoding of the current slice according to the symbol probabilities as obtained in context adaptive entropy decoding of the slice decoded up to the end of the previously decoded slice 38, and may be configured to do so. The decoder may be configured to respond to a syntax element part (18) within the current slice of slice 14 to decode the current slice according to at least one of three modes, i.e., one of the first mode 20 and the third mode 42 or in the second mode 22, and the decoder, according to the third mode 42, i.e., when dependent_slice_flag = 1 and tiles_or_entropy_coding_sync_idc = 3, restricts the derivation of context so as not to cross slice boundaries, and decodes the current slice from the data stream using context adaptive entropy decoding with continuous update of context symbol probabilities and initialization of symbol probabilities independent of any previously decoded slice and prediction decoding across slice boundaries, and one of the first and third modes is selected according to a syntax element, i.e., cabac_independent_flag.The decoder may further be configured to, that is, when tiles_or_entropy_coding_sync_idc = 0, 1, and 3 (``3'' when cabac_independentflag = 0), save the symbol probabilities as obtained in the context-adaptive entropy decoding of the previously decoded slice up to the end of the previously decoded slice, and further, when initializing the symbol probabilities for the context-adaptive entropy decoding of the current slice according to the first mode, initialize the symbol probabilities for the context-adaptive entropy decoding of the current slice according to the saved symbol probabilities. The decoder may be configured to, that is, when tiles_or_entropy_coding_sync_idc = 1, limit the predictive decoding within the tiles in which the image is subdivided in the first and second modes.
[0218] Of course, the encoder can set the above syntax so that the decoder can obtain the above advantages. The encoder may be parallel processing such as, for example, a multi-core encoder, etc., but this is not necessary. To encode the image 10 into the data stream 12 in units of slices 14, the encoder is configured to encode the slices 14 into the data stream 12 according to the slice order 16. The encoder determines the syntax element part 18 for the current slice of the slice so as to signal the current slice in which the syntax element part is encoded according to at least one of two modes 20, 22, and further encodes the syntax element part 18 into it. Further, when the current slice is encoded according to the first mode 20 of at least two modes, it includes context adaptation entropy encoding 24 including derivation of context across slice boundaries, continuous update of context symbol probabilities, and initialization 38, 40 of symbol probabilities according to the saved state of symbol probabilities of previously encoded slices, and prediction encoding across slice boundaries to encode the current slice into the data stream 12. Further, when the current slice is encoded according to the second mode 22 of at least two modes, the derivation of context is restricted so as not to cross the slice boundary, and context adaptation entropy encoding with continuous update of context symbol probabilities and initialization of symbol probabilities independent of any previously encoded slices, and prediction encoding that restricts prediction encoding so as not to cross the slice boundary are used to encode the current slice into the data stream 12. At the same time as the image 10 can be partitioned in an encoding block 32 having a raster scan order 36 arranged in rows and columns and further defined with respect to each other, the encoder can be configured to associate each slice 14 with a continuous subset of the encoding block 32 in the raster scan order 36 such that the subsets follow one another along the raster scan order 36 according to the slice order.The encoder stores symbol probabilities as obtained in context - adaptive entropy coding of slices encoded previously up to the second encoded block 32 in raster scan order 36, and further, when initializing symbol probabilities for context - adaptive entropy coding of the current slice according to the first mode, checks whether the first encoded block of a consecutive subset of the encoded block 32 associated with the current slice is the first encoded block 32 in consecutive order according to the raster scan order 36, and further, if so, initializes 40 the symbol probabilities for context - adaptive entropy coding of the current slice according to the stored symbol probabilities as obtained in context - adaptive entropy coding of slices encoded previously up to the second encoded block in consecutive order according to the raster scan order 36, and further, if not, initializes 38 the symbol probabilities for context - adaptive entropy coding of the current slice according to the symbol probabilities as obtained in context - adaptive entropy coding of the slice decoded previously up to the end of the slice encoded previously. The encoder can be configured to encode the syntax element part (18) into the current slice of the slice (14) such that the current slice is signaled to be encoded according to at least one of three modes, namely one of the first mode (20) and the third mode (42) or in the second mode (22), and the encoder is configured to encode the current slice into the data stream using context - adaptive entropy coding with continuous update of context symbol probabilities and initialization of symbol probabilities independent of any previously encoded slice without crossing slice boundaries and prediction coding crossing slice boundaries according to the third mode (42), and the encoder distinguishes whether it is one of the first and third modes, for example, using a syntax element, namely the cabac_independent_flag.The encoder determines general syntax elements such as, for example, the dependent_slices_present_flag, and further operates in one of at least two general operation modes according to the general syntax elements, that is, according to the first general operation mode, performs encoding of the syntax element part for each slice, and further, according to the second general operation mode, is configured to write the general syntax element into the data stream necessarily using a different one of at least two modes other than the first mode. The encoder can be configured to continuously and necessarily update the symbol probability continuously from the start to the end of the current slice according to the first and second modes. The encoder ay stores the probability for the symbol as obtained in the context adaptive entropy encoding of the previously encoded slice up to the end of the previously encoded slice, and further, when initializing the symbol probability for the context adaptive entropy encoding of the current slice according to the first mode, initializes the symbol probability for the context adaptive entropy encoding of the current slice according to the stored symbol probability. Further, the encoder can limit the predictive coding within the tile in which the image is subdivided in the first and second modes.
[0219] A possible structure of the encoder is shown in FIG. 26 for completeness. The predictor 70 operates almost the same as the predictor 28, that is, performs prediction, but by optimization, also determines encoding parameters including, for example, prediction parameters and modes. Modules 26 and 27 also occur in the decoder. The subtractor 72 determines the lossless prediction residual, which is lossily encoded in the transform and quantization module 74 by the use of quantization and, optionally, using a spectral decomposition transform. The entropy encoder 76 performs context adaptive entropy encoding.
[0220] In addition to the above specific syntax examples, different examples are outlined below showing the correspondence between the terms used below and the terms used above.
[0221] In particular, without elaborating on the above, a dependent slice is not only "dependent" in the sense that it is able to utilize knowledge known from outside its boundary, but also has an entropy context that adapts more rapidly, or achieves better spatial prediction for allowance across its boundary, as outlined above. Rather, to save the rate cost that should be spent to define a slice header by dividing an image into slices, a dependent slice adopts a part of the slice header syntax from the previous slice, i.e., this slice syntax header part is not retransmitted for the dependent slice. This is shown, for example, at 100 in FIG. 16 and at 102 in FIG. 21, and accordingly, the slice type is adopted from the previous slice, for example. By this measure, the re-division of an image into slices, such as independent slices and dependent slices, is less expensive in terms of costly bit consumption.
[0222] It results in the above-described dependency that leads to slightly different turns of phrase in the examples outlined below: A slice is defined as a unit part of an image where the slice header syntax can be set individually. Thus, a slice consists of one independent / regular / normal slice, currently called an independent slice segment using the above nomenclature, and zero, one, or two or more dependent slices, currently called dependent slice segments using the above nomenclature.
[0223] FIG. 27 shows an image partitioned into, for example, two slices, one formed by slice segments 141 to 143 and the other formed by only slice segment 144. Indices 1 to 4 indicate the slice order in the encoding order. FIGS. 28a and 28b show different examples in the case of re-partitioning image 10 into two tiles. In the case of FIG. 28a, one slice covers both tiles 501 and 502 where the index goes up in the encoding order and is formed by all five slice segments 14. Further, in the case of FIG. 28a, two slices each re-partition tile 501 and are formed by slice segments 141, 142, 143, and 144, and another slice covers tile 502 and is formed by slice segments 145-146.
[0224] The definitions can be as follows:
[0225] Dependent slice segment: A slice segment in which the values of some syntax elements of the slice segment header are inferred from the values for the previous independent slice segment in the decoding order, and which was previously called a dependent slice in the above-described embodiments.
[0226] Independent slice segment: A slice segment in which the values of the syntax elements of the slice segment header are not inferred from the values for the previous slice segment, and which was previously called a normal slice in the above-described embodiments.
[0227] Slice: An integer of coding tree units included in one independent slice segment and all subsequent dependent slice segments (if any) preceding the next independent slice segment (if any) within the same access unit / image.
[0228] Slice header: The slice segment header of an independent slice segment that is the current slice segment or the independent slice segment preceding the current dependent slice segment.
[0229] Slice segment: An integer of coding tree units that are consecutively ordered in a tile scan and are further included in a single NAL unit, where the partitioning of each image into slice segments is a partitioned partitioning.
[0230] Slice segment header: A portion of the encoded slice segment that contains data elements related to the first or all coding tree units represented in the slice segment.
[0231] The signaling of "modes" 20 and 22, namely "dependent slice segment" and "independent slice segment", can be as follows:
[0232] In some additional NAL units such as PPS, syntax elements can be used to signal whether the use of dependent slices is formed for a particular image of a sequence for a particular image:
[0233] A dependent_slice_segments_enabled_flag equal to 1 indicates the presence of the syntax element dependent_slice_segment_flag in the slice segment header. A dependent_slice_segments_enabled_flag equal to 0 indicates the absence of the syntax element dependent_slice_segment_flag in the slice segment header.
[0234] The dependent_slice_segments_enabled_flag is similar to the range of the dependent_slices_present_flag described above.
[0235] Similarly, the dependent_slice_flag can be called the dependent_slice_segment_flag in order to occupy a different nomenclature for the slice.
[0236] A dependent_slice_segment_flag equal to 1 indicates that the value of each slice segment header syntax element that does not exist in the header of the current slice segment is presumed to be equal to the value of the corresponding slice segment header syntax element in the slice header, i.e., the slice segment header of the previous independent slice segment.
[0237] At the same level, such as the picture level, the following syntax elements may be included:
[0238] An entropy_coding_sync_enabled_flag equal to 1 indicates that a specific synchronization process for context variables is called before decoding a coding tree unit that includes the first coding tree block of a row of coding tree blocks in each tile in each picture that refers to the PPS, and further, a specific storage process for context variables is called after decoding a coding tree unit that includes the second coding tree block of a row of coding tree blocks in each tile in each picture that refers to the PPS. An entropy_coding_sync_enabled_flag equal to 0 indicates that it is not necessary for a specific synchronization process for context variables to be called before decoding a coding tree unit that includes the first coding tree block of a row of coding tree blocks in each tile in each picture that refers to the PPS, and further, it is not necessary for a specific storage process for context variables to be called after decoding a coding tree unit that includes the second coding tree block of a row of coding tree blocks in each tile in each picture that refers to the PPS. It is a bitstream compliance requirement that the value of the entropy_coding_sync_enabled_flag be for all PPSs activated within the CVS. It is a bitstream compliance requirement that when the entropy_coding_sync_enabled_flag is equal to 1 and further when the first coding tree block in a slice is not the first coding tree block of the row of coding tree blocks in a tile, the last coding tree block in the slice belongs to the same row of coding tree blocks as the first coding tree block in the slice. It is a bitstream compliance requirement that when the entropy_coding_sync_enabled_flag is equal to 1 and further when the first coding tree block in a slice segment is not the first coding tree block of the row of coding tree blocks in a tile, the last coding tree block in the slice segment belongs to the same row of coding tree blocks as the first coding tree block in the slice segment.
[0239] As already described, the encoding / decoding order in CTB30, when more than one tile is present in the image, scans the first tile in raster fashion and then visits the next tile, resulting in a start row by row from top to bottom.
[0240] Decoder 5 and the encoder thus operate as follows in the entropy decoding (encoding) of slice segment 14 of the image:
[0241] A1) Whenever the currently decoded / encoded syntax element synEl is the first syntax element of a tile 50, slice segment 14 or row of CTBs, the initialization process of FIG. 29 is started. A2) Otherwise, the decoding of this syntax element occurs using the current entropy context. A3) If the current syntax element was the last syntax element in CTB30, then the entropy context storage process as shown in FIG. 30 is started. A4) That process advances to the next syntax element in A1).
[0242] In the initialization process, it is checked 200 whether synEI is the first syntax element of slice segment 14 or tile 50. If yes, the context is initialized independently of any previous slice segment in step 202. If no, it is checked 204 whether synEI is the first syntax element of a row of CTB30 and further whether the entropy_coding_sync_enabled_flag is equal to 1. If yes, it is checked 206 whether a second CTB30 is available in the previous line of CTB30 of the equal tile (see Figure 23). If yes, the context adoption according to 40 is executed in step 210 using the currently stored context probability for adoption of type 40. Otherwise, the context is initialized independently of any previous slice segment in step 202. If check 204 reveals no, it is further checked in step 212 whether synEl is the first syntax element in the first CTB of the dependent slice segment 14 and further whether the dependent_slice_segement_flag is equal to 1, and further, if yes, the context adoption according to 38 is executed in step 214 using the currently stored context probability for adoption of type 38. After any of steps 214, 212, 210 and 202, decoding / encoding actually starts.
[0243] A dependent slice segment having a dependent_slice_segement_flag equal to 1 helps to further reduce the encoding / decoding delay with little encoding efficiency penalty.
[0244] In the memory process of FIG. 30, it is checked in step 300 whether the encoded / decoded synEl is the last syntax element of the second CTB30 of the row of CTB30 and whether the entropy_coding_sync_enabled_flag is equal to 1. In the case of yes, the current entropy context is stored in step 302, i.e., the entropy coding probability of the context is stored in the storage specified for the adoption of 40 types. Similarly, in addition to step 300 or 302, it is checked in step 304 whether the encoded / decoded synEl is the last syntax element of slice segment 14 and whether the dependent_slice_segement_flag is equal to 1. In the case of yes, the current entropy context is stored in step 306, i.e., the entropy coding probability of the context is stored in the storage specified for the adoption of 38 types.
[0245] Note that any check asking whether the syntax element is the first synEl of the CTB row uses, for example, the syntax element slice_adress400 in the header of the slice segment, i.e., the start syntax element that reveals the start position of each slice segment along the decoding order.
[0246] When reconstructing the image 10 from the data stream 12 using the WPP process, the decoder can accurately utilize the subsequent start syntax portion 400 to search for the WPP sub-stream entry point. Since each slice segment includes a start syntax portion 400 that indicates the position of the start of decoding of each slice segment within the image 10, the decoder can use the start syntax portion 400 of the slice segment to identify the slice segment that starts on the left side of the image, thereby identifying the entry point of the WPP sub-stream in which the slice segments are grouped. Then, the decoder can continuously start decoding the WPP sub-streams in slice order to decode the WPP sub-streams alternately in parallel. The slice segments may be smaller than one image width, i.e., one row of CTB, so that their transmissions can be interleaved in the WPP sub-streams to further reduce the overall end-to-end transmission delay. The encoder provides each slice (14) having a start syntax portion (400) that indicates the position of the start of encoding of each slice within the image (10), and further groups the slices into WPP sub-streams such that the first slice in the slice order starts on the left side of the image for each WPP sub-stream. The encoder can even use the WPP process alone when encoding the image: the encoder continuously starts encoding the WPP sub-streams in slice order to encode the WPP sub-streams alternately in parallel.
[0247] Incidentally, a subsequent aspect of using the start syntax portion of the slice segment as a means for determining the position of the entry point of the WPP sub-stream can be used without the dependent slice concept.
[0248] It is achievable for all for parallel processing of the image 10 by setting the above variables as follows:
[0249] TIFF2025093939000009.tif99159
[0250] It is possible to tile partition and mix the WPP. In that case, one can consider the tiles as individual images: each one using the WPP consists of slices having one or more dependent slice segments, and the checks in steps 300 and 208 refer to a second CTB in the CTB rows described above in the same tile in the same way that steps 204 and A1 refer to the first CTB in the CTB30 rows of the current tile. In that case, the table described above can be extended:
[0251] TIFF2025093939000010.tif240161
[0252] As a brief note, later extensions are enabled in Embodiment 2. Embodiment 2 enables the following processing:
[0253] TIFF2025093939000011.tif103166
[0254] However, for the following extensions, the following table results:
[0255] In addition to the semantics of the image parameter set: If tiles_or_entropy_coding_sync_idc is equal to 4, each, however, the first row of the CTB is included in a different slice having a dependent slice flag set to 1. CTBs of different rows must not be in the same slice. There may be more than one slice per CTB row.
[0256] If tiles_or_entropy_coding_sync_idc is equal to 5, each CTB, however, the first tile must be in a different slice. CTBs of different tiles must not be in the same slice. There may be more than one slice per tile.
[0257] For further explanation, refer to FIG. 31.
[0258] That is, the above table can be extended:
[0259] TIFF2025093939000012.tif180129
[0260] Regarding the above embodiment, it should be noted that the decoder should be configured to read information from the current slice that reveals the re - division of the current slice into parallel sub - sections, for example, in the first and second modes, in response to tiles_or_entropy_coding_sync_idc = 1, 2, where the parallel sub - sections can be WPP sub - streams or tiles. In the first mode, it includes the initialization of symbol probabilities depending on the saved state of symbol probabilities of the preceding parallel sub - section, and in the second mode, it includes the initialization of symbol probabilities independent of the previously decoded slice and previously decoded parallel sub - sections. Context - adaptive entropy decoding is aborted at the end of the first parallel sub - section and newly restarted at the start of any subsequent parallel sub - section.
[0261] Thus, the above description reveals a method for low - latency encoding, decoding, encapsulation, and transmission of video data configured to be provided by the new HEVC coding standard, for example, configured in tiles, wavefront parallel processing (WPP) sub - streams, slices, or entropy slices.
[0262] In particular, it is defined how to transfer the parallel - encoded data in a conversational scenario in order to obtain the minimum latency in the encoding, decoding, and transmission processes. Therefore, a pipeline parallel encoding, transmission, and decoding approach is described to enable minimum - latency applications such as games, remote surgery, etc.
[0263] Furthermore, the above-described embodiments fill the gap of wavefront parallel processing (WPP) to enable it to be used in low-latency transmission scenarios. Therefore, a new encapsulation format for WPP substream 0 is shown with a dependent slice. This dependent slice can include entropy slice data, the WPP substream, a complete row of LCUs, just a fragment of the slice, and the conventional transmitted slice header is also applied to the included fragment data. The included data is signaled in the sub-slice header.
[0264] Finally, it should be noted that while the naming for the new slice could also be "subset / lite slice", the name "dependent slice" has been found to be more suitable.
[0265] Signaling is shown that describes the level of parallelization in encoding and transfer.
[0266] Although several aspects have been described in relation to an apparatus, it is clear that these aspects also represent descriptions of corresponding methods, and that a block or apparatus corresponds to a method step or a feature of a method step. Similarly, aspects described in relation to method steps also represent descriptions of corresponding blocks or items or features of a corresponding apparatus. Some or all of the method steps may be performed by (or using) a hardware device such as, for example, a microprocessor, a programmable computer, or an electronic circuit. In some embodiments, one or more of any of the most important method steps may be performed by such a device.
[0267] Depending on specific implementation requirements, embodiments of the present invention can be implemented in hardware or in software. The implementation can be carried out using a digital storage medium, such as a floppy (registered trademark) disk, DVD, Blu-ray (registered trademark), CD, ROM, PROM, EPROM, EEPROM or FLASH memory, which stores electronically readable control signals that cooperate (or can cooperate) with a programmable computer system so that each method is executed. Therefore, the digital storage medium may be computer-readable.
[0268] Some embodiments according to the present invention include a data carrier having electronically readable control signals that can cooperate with a programmable computer system so that one of the methods described herein is executed.
[0269] Generally, embodiments of the present invention can be implemented as a computer program product having program code, which, when the computer program product is executed on a computer, serves to execute one of those methods. The program code may be stored, for example, on a machine-readable carrier.
[0270] Other embodiments include a computer program for executing one of the methods described herein, stored on a machine-readable carrier.
[0271] Therefore, in other words, an embodiment of the method of the present invention is a computer program having program code for executing one of the methods described herein when the computer program is executed on a computer.
[0272] Accordingly, a further embodiment of the method of the present invention is a data carrier (or digital storage medium or computer-readable medium) containing a computer program for carrying out one of the methods described herein, recorded thereon. The data carrier, digital storage medium or recording medium is typically tangible and / or non-transitory.
[0273] Accordingly, a further embodiment of the method of the present invention is a data stream or a series of signals representing a computer program for carrying out one of the methods described herein. The data stream or series of signals may be configured to be transferred via a data communication connection, such as the Internet.
[0274] A further embodiment includes processing means, such as a computer or a programmable logic device, configured or adapted to carry out one of the methods described herein.
[0275] A further embodiment includes a computer having installed thereon a computer program for carrying out one of the methods described herein.
[0276] A further embodiment according to the present invention includes an apparatus or system configured to transfer (e.g., electronically or optically) a computer program for carrying out one of the methods described herein to a receiver. The receiver may be, for example, a computer, a mobile device, a memory device, etc. The apparatus or system may include, for example, a file server for transferring the computer program to the receiver.
[0277] In some embodiments, a programmable logic device (e.g., a field programmable gate array) may be used to perform some or all of the functions of the methods described herein. In some embodiments, a field programmable gate array may cooperate with a microprocessor to perform one of the methods described herein. Generally, the methods are preferably performed by any hardware device.
[0278] The above embodiments are merely illustrative for the principles of the present invention. Modifications and changes to the configurations and details described herein will be apparent to other persons skilled in the art. Therefore, the present invention is intended to be limited only by the scope of the impending claims and not by the specific details shown herein as descriptions and explanations of the embodiments.
[0279] Literature [1] Thomas Wiegand, Gary J. Sullivan, Gisle Bjontegaard, Ajay Luthra, "Overview of the H.264 / AVC Video Coding Standard", IEEE Trans. Circuits Syst. Video Technol., vol. 13, N7, July 2003. [2] JCT-VC, "High-Efficiency Video Coding (HEVC) text specification Working Draft 6", JCTVC-H1003, February 2012. [3] ISO / IEC 13818-1: MPEG-2 Systems specification.
Claims
1. 1. A decoder for reconstructing an image (10) from a data stream (12) in which the image (10) is coded in slices (14) into which the image (10) is partitioned, the decoder being configured to decode the slices (14) from the data stream (12) according to a slice order (16), the decoder being responsive to a syntax element portion (18) within a current slice of the slices to decode the current slice according to one of at least two modes (20, 22), and decoding the current slice from the data stream (12) using a context-adaptive entropy decoding (24) according to a first mode (20) of the at least two modes, the context-adaptive entropy decoding (24) including deriving a context across slice boundaries, continuously updating symbol probabilities of the context and initializing the symbol probabilities (38, 40) according to stored states of symbol probabilities of previously decoded slices, and predictive decoding across the slice boundaries; A decoder that decodes the current slice from the data stream (12) according to a second mode (22) of the at least two modes using context adaptive entropy decoding, which restricts the derivation of the context so as not to cross the slice boundaries, and has continuous updating of symbol probabilities of the context and initialization of the symbol probabilities independent of any previously decoded slices, and predictive decoding, which restricts predictive decoding so as not to cross the slice boundaries.
2. 2. The decoder of claim 1, wherein the image (10) is partitioned in coding blocks (32) arranged in rows and columns and having a raster scan order (36) defined relative to one another, and the decoder is further configured to associate each slice (14) with a consecutive subset of the coding blocks (32) in the raster scan order (36) such that the subsets follow one another along the raster scan order (36) according to the slice order.
3. The decoder stores symbol probabilities as obtained in a context adaptive entropy decoding of the previously decoded slice up to a second coding block (32) in succession according to the raster scan order (36), and further checks, when initializing the symbol probabilities for the context adaptive entropy decoding of the current slice in accordance with the first mode, whether a first coding block of the contiguous subset of coding blocks (32) associated with the current slice is a first coding block (32) in succession according to the raster scan order, and if so, ...).
3. The decoder of claim 2, configured to initialize (40) the symbol probabilities for the context-adaptive entropy decoding of the current slice in dependence on the stored symbol probabilities as obtained in the context-adaptive entropy decoding of the previously decoded slices successively up to a second coding block according to (36), and further configured to initialize (38) the symbol probabilities for the context-adaptive entropy decoding of the current slice in dependence on the symbol probabilities as obtained in the context-adaptive entropy decoding of the previously decoded slices up to the end of the previously decoded slices otherwise.
4. The decoder is configured to respond to the syntax element portion (18) within the current slice of the slices (14) to decode the current slice according to one of at least three modes, namely one of the first mode (20) and third mode (42) or a second mode (22), the decoder comprising: According to the third mode (42), the method is configured to restrict the derivation of the context so as not to cross the slice boundaries, and to decode the current slice from the data stream using context-adaptive entropy decoding with continuous updating of symbol probabilities of the context and initialization of the symbol probabilities independent of any previously decoded slices, and predictive decoding across the slice boundaries, A decoder as claimed in any preceding claim, wherein one of the first and third modes is selected in response to a syntax element.
5. 4. A decoder as claimed in claim 1, wherein the decoder is configured to respond to general syntax elements in the data stream to operate in one of at least two general operational modes, performing the response to the syntax element portions on a slice-by-slice basis in accordance with a first general operational mode, and further comprising a decoder according to a second general operational mode which necessarily uses a different one of the at least two modes other than the first mode.
6. 3. The decoder of claim 2, wherein the decoder is configured to necessarily and uninterruptedly continue successive updating of the symbol probabilities from the beginning to the end of the current slice according to the first and second modes.
7. 3. The decoder of claim 2, further configured to: store symbol probabilities as obtained in context adaptive entropy decoding of the previously decoded slice until the end of the previously decoded slice; and further configured to initialize the symbol probabilities for the context adaptive entropy decoding of the current slice according to the first mode, depending on the stored symbol probabilities.
8. A decoder as claimed in any preceding claim, wherein the decoder is arranged to constrain the predictive decoding in the first and second modes within tiles into which the image is subdivided.
9. 9. A decoder as claimed in claim 1, wherein the decoder is configured to read information from a current slice revealing a subdivision of the current slice into parallel subsections in the first and second modes, to stop the context adaptive entropy decoding at the end of a first parallel subsection, and to restart the context adaptive entropy decoding anew at the start of any subsequent parallel subsection, including in the first mode an initialization of the symbol probabilities according to a saved state of the symbol probabilities of a previous parallel subsection, and in the second mode an initialization of the symbol probabilities independent of any previously decoded slice and any previously decoded parallel subsection.
10. 10. The decoder of claim 1, wherein the decoder is configured to copy for the current slice, in accordance with the first mode (20) of the at least two modes, a portion of a slice header syntax from a previous slice decoded in the second mode.
11. The decoder is configured to reconstruct the image (10) from the data stream (12) using a WPP process, each slice (14) including a start syntax portion (400) indicating a location within the image (10) for starting decoding of the respective slice, and further comprising: identifying an entry point of a WPP sub-stream into which the slice is grouped by identifying a slice starting on the left side of the image using a start syntax portion of the slice; and A decoder according to any preceding claim, configured for staggered parallel decoding of the WPP sub-streams by starting the decoding of the WPP sub-streams consecutively according to the slice order.
12. 1. An encoder for encoding an image (10) into a data stream (12) in units of slices (14) into which the image (10) is partitioned, the encoder configured to encode the slices (14) into the data stream (12) according to a slice order (16), the encoder further comprising: determining a syntax element portion (18) for a current slice of said slices such that the syntax element portion indicates a current slice to be coded according to one of at least two modes (20, 22); and coding the syntax element portion (18) into the current slice; if the current slice is coded according to a first mode (20) of the at least two modes, coding the current slice into the data stream (12) using a context-adaptive entropy coding (24) including deriving a context across slice boundaries, continuously updating symbol probabilities of the context and initializing the symbol probabilities (38, 40) according to a stored state of symbol probabilities of a previously coded slice, and predictive coding across the slice boundaries; an encoder configured to, when the current slice is encoded according to a second mode (22) of the at least two modes, encode the current slice into the data stream (12) using context-adaptive entropy coding that restricts the derivation of the context so as not to cross the slice boundaries and has continuous updating of symbol probabilities of the context and initialization of the symbol probabilities independent of any previously encoded slices, and predictive coding that restricts predictive coding so as not to cross the slice boundaries.
13. 13. The encoder of claim 12, wherein the image (10) is partitioned in coding blocks (32) arranged in rows and columns and having a raster scan order (36) defined relative to one another, and the encoder is configured to associate each slice (14) with a consecutive subset of the coding blocks (32) in the raster scan order (36) such that the subsets follow one another along the raster scan order (36) according to the slice order.
14. 14. The encoder of claim 13, further configured to: store symbol probabilities as obtained in context-adaptive entropy coding of the previously coded slice up to a second coding block (32) successively according to the raster scan order (36); and further configured, when initializing the symbol probabilities for the context-adaptive entropy coding of the current slice according to the first mode, to check whether a first coding block of the consecutive subset of coding blocks (32) associated with the current slice is a first coding block (32) successively according to the raster scan order; and, if so, to initialize (40) the symbol probabilities for the context-adaptive entropy coding of the current slice in dependence on the stored symbol probabilities as obtained in context-adaptive entropy coding of the previously coded slice up to a second coding block (32) successively according to the raster scan order (36); and, if not, to initialize (38) the symbol probabilities for the context-adaptive entropy coding of the current slice in dependence on the symbol probabilities as obtained in context-adaptive entropy coding of the previously decoded slice up to the end of the previously coded slice.
15. The encoder is configured to encode the syntax element portion (18) into the current slice of the slices (14) such that the current slice is signaled for being encoded therein according to one of at least three modes, namely in one of the first mode (20) and the third mode (42) or in the second mode (22), the encoder being configured to: According to the third mode (42), the derivation of the context is restricted so as not to cross the slice boundaries, and the current slice is adapted to be coded into the data stream using context-adaptive entropy coding with continuous updating of symbol probabilities of the context and initialization of the symbol probabilities independent of any previously coded slice, and predictive coding across the slice boundaries, An encoder as claimed in any one of claims 12 to 14, wherein the encoder distinguishes between one of the first and third modes using a syntax element.
16. 16. An encoder according to claim 12, wherein the encoder is configured to determine a generic syntax element and to operate in one of at least two generic operation modes depending on the generic syntax element, namely to perform encoding of the syntax element portions slice by slice according to a first generic operation mode and to write the generic syntax element into the data stream necessarily using a different one of the at least two modes other than the first mode according to a second generic operation mode.
17. The encoder of claim 13 , wherein the encoder is configured to necessarily and uninterruptedly continue successive updating of the symbol probabilities from the beginning to the end of the current slice according to the first and second modes.
18. 14. The encoder of claim 13, wherein the encoder is configured to: store symbol probabilities as obtained in context adaptive entropy coding of the previously coded slice until the end of the previously coded slice; and further configured to initialize the symbol probabilities for the context adaptive entropy coding of the current slice according to the first mode in response to the stored symbol probabilities when initializing the symbol probabilities for the context adaptive entropy coding of the current slice according to the first mode.
19. An encoder as claimed in any of claims 12 to 18, wherein the encoder is configured to constrain the predictive coding in the first and second modes within tiles into which the image is subdivided.
20. 1. A decoder for reconstructing an image (10) from a data stream (12) in which the image (10) is coded in slices (14) into which the image (10) is partitioned using a WPP process, the decoder being configured to decode the slices (14) from the data stream (12) according to a slice order (16), each slice (14) including a start syntax portion (400) indicating a position within the image (10) for starting decoding of the respective slice, the decoder further comprising: Identifying an entry point of a WPP substream into which the slice is grouped by identifying a slice that starts on the left side of the image using a start syntax portion (400) of the slice; A decoder configured to sequentially commence the decoding of the WPP sub-streams according to the slice order to perform staggered parallel decoding of the WPP sub-streams.
21. 1. An encoder for encoding an image (10) into a data stream (12) in which the image (10) is encoded in slices (14) into which the image (10) is partitioned using a WPP process, the encoder being configured to encode the slices (14) into the data stream (12) according to a slice order (16), the encoder being configured to provide each slice (14) with a start syntax portion (400) indicating a position within the image (10) where encoding of the respective slice begins, the encoder further comprising: grouping the slices into WPP substreams such that for each WPP substream, the first slice in slice order starts on the left side of the image; and An encoder configured to begin the encoding of the WPP sub-streams sequentially according to the slice order to perform staggered parallel encoding of the WPP sub-streams.
22. 1. A method for reconstructing an image (10) from a data stream (12) in which the image (10) is coded in slices (14) into which the image (10) is partitioned, the method comprising the steps of decoding the slices (14) from the data stream (12) according to a slice order (16), the method being responsive to a syntax element portion (18) within a current slice of the slice for decoding the current slice according to one of at least two modes (20, 22), According to a first mode (20) of the at least two modes, the current slice is decoded from the data stream (12) using a context-adaptive entropy decoding (24) including deriving a context across slice boundaries, continuously updating symbol probabilities of the context and initializing the symbol probabilities (38, 40) according to a stored state of symbol probabilities of previously decoded slices, and a predictive decoding across the slice boundaries; According to a second mode (22) of the at least two modes, the current slice is decoded from the data stream (12) using context adaptive entropy decoding, which restricts the derivation of the context so as not to cross the slice boundaries, and has continuous updating of symbol probabilities of the context and initialization of the symbol probabilities independent of any previously decoded slices, and predictive decoding, which restricts predictive decoding so as not to cross the slice boundaries.
23. A method for encoding an image (10) into a data stream (12) in units of slices (14) into which the image (10) is partitioned, the method comprising the step of encoding the slices (14) into the data stream (12) according to a slice order (16), the method further comprising: determining a syntax element portion (18) for a current slice of the slices, such that the syntax element portion signals the current slice being coded according to one of at least two modes (20, 22), and further coding the syntax element portion (18) into the current slice; if the current slice is coded according to a first mode (20) of the at least two modes, coding the current slice into the data stream (12) using a context-adaptive entropy coding (24) including derivation of a context across slice boundaries, continuous updating of symbol probabilities of the context and initialization of the symbol probabilities (38, 40) according to a saved state of the symbol probabilities of a previously coded slice, and predictive coding across the slice boundaries; If the current slice is coded according to a second mode (22) of the at least two modes, the method includes the step of coding the current slice into the data stream (12) using context-adaptive entropy coding, which restricts the derivation of the context so as not to cross the slice boundaries, and has continuous updating of symbol probabilities of the context and initialization of the symbol probabilities independent of any previously coded slice, and predictive coding, which restricts predictive coding so as not to cross the slice boundaries.
24. 1. A method for reconstructing an image (10) from a data stream (12) in which the image (10) is encoded in slices (14) into which the image (10) is partitioned using a WPP process, the method comprising the steps of: decoding the slices (14) from the data stream (12) according to a slice order (16), each slice (14) including a start syntax portion (400) indicating a position within the image (10) for starting decoding of the respective slice, the method further comprising: identifying an entry point of a WPP sub-stream into which said slice is grouped by identifying a slice starting on the left side of said image using the start syntax part (400) of said slice; The method, further comprising: commencing the decoding of the WPP sub-streams sequentially according to the slice order to perform staggered parallel decoding of the WPP sub-streams.
25. A method for encoding an image (10) into a data stream (12) in which the image (10) is encoded in slices (14) into which the image (10) is partitioned using a WPP process, the method comprising the steps of: encoding the slices (14) into the data stream (12) according to a slice order (16); providing each slice (14) with a start syntax portion (400) indicating a position within the image (10) where encoding of the respective slice begins; grouping the slices into WPP substreams such that for each WPP substream the first slice in slice order starts on the left side of the image; and 3. The method of claim 2, further comprising: sequentially commencing the encoding of the WPP sub-streams according to the slice order to perform staggered parallel encoding of the WPP sub-streams.
26. A computer program having a program code for performing the method according to claim 22 or 25, when the computer program runs on a computer.
Citation Information
Patent Citations
Video encoding and decoding methods and apparatuses
JP2009177787A