Image coding that supports block merging and skip modes

Common signaling for integration and skip modes in image and video coders optimizes encoding efficiency by minimizing redundant data transmission and parameter granularity, addressing inefficiencies in existing coders.

JP2026053421APending Publication Date: 2026-03-25DOLBY VIDEO COMPRESSION LLC
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-12-11
Publication Date
2026-03-25

AI Technical Summary

Technical Problem

Existing image and video coders face inefficiencies in balancing the amount of auxiliary information required for prediction parameters and the freedom to subdivide images, leading to increased bitrates due to redundant residual data transmission.

Method used

Implementing common signaling for both integration activation and skip mode activation within the bitstream, reducing the need for separate signaling of these modes and optimizing encoding parameters across groups of blocks.

Benefits of technology

This approach reduces the bitrate by minimizing redundant signaling, enhancing encoding efficiency and reducing the granularity of parameter transmission.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026053421000001_ABST
    Figure 2026053421000001_ABST
Patent Text Reader

Abstract

The present invention provides an apparatus, method, and digital storage medium that offer an encoding concept having increased encoding efficiency. [Solution] A method to achieve increased encoding efficiency by using a common signaling in the bitstream 30 with respect to both the activation of integration and the activation of skip mode, wherein the encoder 10 signals that one of one or more possible states of the syntax elements in the bitstream is that, with respect to the current sample set of image 20, each sample set is to be integrated and has no predicted residual encoded and inserted into the bitstream.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to image and / or video coding, and particularly to coders that support block partitioning and skip modes.

Background Art

[0002] Many image and / or video coders process images in blocks. For example, predictive coders need to use too much auxiliary information for prediction parameters, resulting in either very accurately setting the prediction parameter set at high spatial resolution or coarsely setting the prediction parameters, which leads to an increase in the amount of bits required to encode the prediction residuals due to the low spatial resolution of the prediction parameters. To obtain a good compromise between these two extremes, block granularity is used. In short, the optimal setting for the prediction parameters lies somewhere between the two extremes.

[0003] To obtain an optimal solution for the above problems, several attempts have been made. For example, instead of using regular subdivision of an image into blocks regularly arranged in rows and columns, multi-tree subdivision tries to increase the freedom to subdivide an image into blocks with appropriate requirements for subdivision information. However, even multi-tree subdivision requires a significant amount of data signaling, and the freedom to subdivide an image is quite limited even when using such multi-tree subdivision.

[0004] To enable a better trade-off between the amount of auxiliary information required to signal image subdivision and the freedom to subdivide an image, block integration (merge) can be used with a reasonable amount of additional data required to signal the integration information to increase the number of possible image subdivisions. For integrated blocks, the coding parameters also need to be transmitted only once in the bitstream as a whole, as if the resulting integrated group of blocks were directly subdivided parts of the image.

[0005] To further increase the efficiency of encoding image content, skip mode has been introduced in some block-based image codecs, allowing the encoder to avoid sending residual data for a particular block to the decoder. In other words, skip mode can suppress the transmission of residual data for a particular block. The ability to suppress the transmission of residual data for a particular block results in a wider granularity for encoding encoding / prediction parameters, where an optimal trade-off between encoding quality and the overall bitrate spent can be expected. Naturally, this reduces the residual portion that lowers the rate required to encode the residual data, but on the other hand, increasing the spatial resolution of encoding encoding / prediction parameters results in an increase in the rate of auxiliary information. However, since skip mode is available, it may be advantageous to obtain a significant saving in encoding rate by simply increasing the granularity at which encoding / prediction parameters are transmitted, so that the residual portion is so small that even the separate transmission of the residual portion can be omitted.

[0006] However, due to the remaining redundancy newly introduced by the combination of block consolidation and skip mode usage, it is still necessary to achieve better encoding efficiency. [Overview of the project] [Problems that the invention aims to solve]

[0007] Thus, the object of the present invention is to provide an encoding concept having increased encoding efficiency. This object is achieved by the independent claims of the application. [Means for solving the problem]

[0008] The underlying idea of ​​the present invention is that further increases in coding efficiency can be achieved when common signaling is used within the bitstream for both integration activation and skip mode activation. That is, one of the possible states of one or more syntax elements within the bitstream may signal that, with respect to the current sample set of the image, each sample set is to be integrated and the predicted residuals are not encoded and inserted into the bitstream. In other words, a common flag may signal whether the coding parameters associated with the current sample set will be set according to the integration candidates or extracted from the bitstream, and whether the current sample set of the image will be reconstructed based only on the predicted signal corresponding to the coding parameters associated with the current sample set, without any residual data, or whether it will be reconstructed by refining the predicted signal corresponding to the coding parameters associated with the current sample set with residual data within the bitstream.

[0009] The inventors of the present invention have found that the introduction of common signaling for the activation of the integration mode and the activation of the skip mode saves bitrate, such that the additional overhead of signaling the activation of the integration mode and / or the activation of the skip mode separately from each other can be reduced, or only needs to be incurred if the integration mode and the skip mode are not activated simultaneously.

[0010] Advantageous embodiments of the present invention are the subject of the accompanying dependent claims.

[0011] Preferred embodiments of the present invention are described below in further detail with respect to the figures. [Brief explanation of the drawing]

[0012] [Figure 1]Figure 1 shows a block diagram of an apparatus for encoding according to an embodiment. [Figure 2] Figure 2 shows a block diagram of an apparatus for encoding according to a more detailed embodiment. [Figure 3] Figure 3 shows a block diagram of the device for decoding according to the embodiment. [Figure 4] Figure 4 shows a block diagram of an apparatus for decoding according to a more detailed embodiment. [Figure 5] Figure 5 shows a block diagram of a possible internal structure of the encoder in Figure 1 or Figure 2. [Figure 6] Figure 6 shows a block diagram of a possible internal structure of the decoder shown in Figure 3 or Figure 4. [Figure 7a] Figure 7a schematically shows the possible subdivision of an image into tree root blocks, coding units (blocks), and prediction units (partitions). [Figure 7b] Figure 7b shows the subdivided tree of the tree root block shown in Figure 7a, down to the partition level, according to the illustrated example. [Figure 8] Figure 8 shows an embodiment for a set of possible supported split patterns according to the embodiment. [Figure 9] Figure 9 shows possible partitioning patterns that efficiently result from combining block merging and block partitioning when using block partitioning according to Figure 8. [Figure 10] Figure 10 schematically shows candidate blocks for skip / direct mode according to the embodiment. [Figure 11] Figure 11 shows the syntax portion of the syntax according to the embodiment. [Figure 12] Figure 12 shows the syntax portion of the syntax according to the embodiment. [Figure 13a] Figure 13a shows the syntax portion of the syntax according to the embodiment. [Figure 13b] Figure 13b shows the syntax portion of the syntax according to the embodiment. [Figure 14]FIG. 14 schematically shows the definition of adjacent partitions for the partition according to the embodiment.

Best Mode for Carrying Out the Invention

[0013] Regarding the following description, whenever the same reference numerals are used in relation to different figures, the description associated with each element presented with respect to one of these figures may be transferred from one figure to another figure as long as it does not conflict with the remaining description of this other figure. Similarly, note that it applies to other figures.

[0014] FIG. 1 shows an apparatus 10 for encoding an image 20 into a bitstream 30. Of course, the image 20 can be part of a video if the encoder 10 is a video encoder.

[0015] Although not explicitly shown in FIG. 1, the image 20 is shown as an array of samples. The sample array of the image 20 is divided into sample sets 40. And they can be any set of samples, such as a sample set that covers one contiguous area of the image 20 without overlap. For ease of understanding, the sample set 40 is shown as a block 40 and will be referred to as a block hereinafter, but the following description should not be considered limited to a particular type of sample set 40. According to a specific embodiment, the sample set 40 is a rectangular and / or square block.

[0016] For example, image 20 can be subdivided into a regular arrangement of blocks 40, such that the blocks 40 are arranged in rows and columns, as shown in Figure 1 as an example. However, any other subdivision of image 20 into blocks 40 is possible. In particular, the subdivision of image 20 into blocks 40 can be fixed, i.e., known to the decoder by default, or signaled within the bitstream 30 to the decoder. In particular, the blocks 40 of image 20 can vary in size. For example, multi-tree subdivision, such as quad-tree subdivision, can be applied to image 20, or to a regular pre-subdivision of image 20 into a regularly arranged tree root block, in this case to obtain blocks 40 that form leaf blocks of the multi-tree subdivision of the tree root block.

[0017] In any case, the encoder 10 is configured to encode a flag into the bitstream 30 that commonly signals whether, with respect to the current sample set 40, the encoding parameters associated with the current sample set 40 will be set according to the integrated candidate or extracted from the bitstream 30, and whether the current sample set of image 20 will be reconstructed based solely on the predicted signal corresponding to the encoding parameters associated with the current sample set, without any residual data, or whether it will be reconstructed by refining the predicted signal corresponding to the encoding parameters associated with the current sample set 40 using residual data in the bitstream 30. For example, the encoder 10 is configured to encode a flag in the bitstream 30 that commonly signals that, in the first state with respect to the current sample set 40, the encoding parameters associated with the current sample set 40 will be set according to the integration candidate rather than being extracted from the bitstream 30, and that the current sample set of image 20 will be reconstructed based solely on the predicted signal corresponding to the encoding parameters associated with the current sample set, without any residual data; and in the other state, the encoding parameters associated with the current sample set 40 will be extracted from the bitstream 30, or that the current sample set of image 20 will be reconstructed by refining the predicted signal corresponding to the encoding parameters associated with the current sample set 40 with residual data in the bitstream 30. This means that: the encoder 10 supports the integration of blocks 40. Integration is conditional; that is, not all blocks 40 will undergo integration. For some blocks 40, it is currently advantageous to merge them with the merger candidates, for example, from the perspective of rate distortion optimization, but for others, the opposite is true.To determine whether a particular block 40 should be subject to integration, the encoder 10 determines a set or list of integration candidates and, for each of these integration candidates, checks whether integrating that candidate with the current block 40 forms the most favorable encoding option, for example, in terms of rate distortion optimization. The encoder 10 is configured to determine a set or list of integration candidates for the current block 40 based on previously encoded portions of the bitstream 30. For example, the encoder 10 extracts at least a portion of the set or list of integration candidates by employing encoding parameters associated with spatially and / or temporally adjacent blocks 40 that were previously encoded according to the encoding order applied by the encoder 10. Temporal adjacency represents, for example, a block of image encoded before the video to which image 20 belongs, and that temporally adjacent block is spatially located so as to spatially overlap the current block 40 of the current image 20. Thus, with respect to this portion of the set or list of integration candidates, there is a one-to-one relationship between each integration candidate and its spatially and / or temporally adjacent block. Each integration candidate has encoding parameters associated with it. If block 40 is to be merged into one of the merger candidates, the encoder 10 sets the coding parameters of block 40 according to the merger candidate. For example, the encoder 10 can set the coding parameters of block 40 to be equal to each merger candidate. That is, the encoder 10 can replicate the coding parameters of block 40 from each merger candidate. Thus, with respect to this above portion of the set or list of merger candidates, the coding parameters of the merger candidates are adopted directly from spatially and / or temporally adjacent blocks, or the coding parameters of each merger candidate are obtained from the coded data of such spatially and / or temporally adjacent blocks by adopting them, i.e., by setting an equal merger candidate. However, on the other hand, region changes are taken into account, for example, by scaling the adopted coding parameters according to region changes.For example, at least some of the encoding parameters that follow the integration can include motion parameters. However, the motion parameters may relate to different reference image indices. More precisely, the motion parameters that will be adopted may relate to a specific time interval between the current image and the reference image, and when integrating the current block with each integration candidate having each motion parameter, the encoder 10 can be configured to scale the motion parameter of each integration candidate in order to apply that time interval to the time interval selected for the current block.

[0018] In any case, the integration candidates described so far have in common that all of them have the encoding parameters associated with them, and there is a one-to-one relationship between these integration candidates and the blocks adjacent to them. Therefore, integrating any of the above integration candidates with block 40 can be considered as the integration of these blocks into one or more groups of block 40 so that, except for scaling adaptation etc., the encoding parameters do not change across the entire image 20 within these groups of block 40. In fact, the integration regarding any of the above integration candidates reduces the granularity with which the encoding parameters change across image 20. Moreover, the integration regarding any of the above integration candidates results in additional degrees of freedom when subdividing image 20 into block 40 and groups of block 40, respectively. Thus, in this regard, the integration of block 40 into this kind of group of blocks allows the encoder 10 to consider encoding image 20 using encoding parameters that vary across the entire image 20 in units of these groups of block 40.

[0019] In addition to the integration candidates described above, the encoder 10 can also add integration candidates to a set / list of integration candidates that are the result of a combination of the encoding parameters of two or more adjacent blocks, such as their arithmetic mean, geometric mean, or the median of the encoding parameters of adjacent blocks.

[0020] Thus, in effect, the encoder 10 reduces the granularity at which the encoding parameters are explicitly transmitted within the bitstream 30 compared to the granularity determined by the subdivision of the image 20 into blocks 40. Some of these blocks 40 form a group of blocks using the exact same encoding parameters with the integration options outlined above. Some blocks are linked together by integration but use different encoding parameters that are correlated with each other by the functions of each scaling adaptation and / or combination. Some blocks 40 do not undergo integration, and therefore the encoder 10 directly encodes the encoding parameters into the bitstream 30.

[0021] The encoder 10 uses the coding parameters of the blocks 40 thus defined to determine the prediction signal for the image 20. The encoder 10 performs this block-by-block determination of the prediction signal in such a way that the prediction signal is determined by the coding parameters associated with each block 40.

[0022] Another decision made by the encoder 10 is whether the residual portion, i.e., the difference between the predicted signal and the original image content in each local part of the current block 40, will be transmitted in the bitstream 30. That is, the encoder 10 decides with respect to block 40 whether a skip mode is applied to each block. If a skip mode is applied, the encoder 10 simply encodes the image 20 in the current portion 40 in the form of a predicted signal obtained from or corresponding to the encoding parameters associated with each block 40; if a skip mode is not selected, the encoder 10 encodes the image 20 into the bitstream 30 of block 40 using both the predicted signal and the residual data.

[0023] To conserve bitrate for signaling decisions regarding merge and skip modes, the encoder 10 signals both decisions in common using a single flag for block 40. More precisely, common signaling can be achieved such that the activation of both merge and skip modes is commonly indicated by a flag in each block 40 of the bitstream 30 taking a first possible flag state, while another flag state of that flag simply indicates to the decoder that neither merge nor skip mode will be activated. For example, the encoder 10 may decide with respect to a particular block 40 to activate merge but deactivate skip mode. In that case, the encoder 10 uses another flag state to signal the deactivation of at least one of merge and skip modes within the bitstream 30, while subsequently signaling the activation of merge within the bitstream 30 using another flag, for example. Therefore, the encoder 10 only needs to send this additional flag in the case of block 40 where the merge and skip modes are not activated simultaneously. In embodiments further described later, the first flag is called mrg_cbf or skip_flag, while the auxiliary merge indicator flag is called mrg or merge_flag. The inventors of the present application have found that this common use of a single signaling state to commonly signal the activation of the merge and skip modes reduces the overall bitrate of the bitstream 30.

[0024] Regarding the signaling states described above, it should be noted that this type of signaling state can be determined by the state of one bit of the bitstream 30. However, the encoder 10 may be configured to entropically encode the bitstream 30, and therefore the correspondence between the flag and the signaling state of the bitstream 30 can become more complex. In that case, the state can correspond to one bit of the bitstream 30 in the entropy decoding region. Furthermore, the signaling state can correspond to one of two states of the flag to which the codeword is assigned according to the variable-length coding scheme. In the case of arithmetic coding, the signaling state that commonly signals the activation of the integration and skip modes can correspond to one of the symbols of the symbol alphabet that forms the basis of the arithmetic coding scheme.

[0025] As outlined above, the encoder 10 signals the simultaneous activation of the merge and skip modes using a flag in the bitstream 30. This flag may be transmitted within a syntax element having two or more possible states, as outlined in more detail below. This syntax element may also signal other coding options, for example. The details will be further explained below. However, in that case, one of the possible states of the syntax element signals the simultaneous activation. That is, whenever the syntax element just mentioned in block 40 takes this default possible state, the encoder 10 signals the activation of both the merge and skip modes. The decoder does not need to signal further for the activation of the merge and the activation of the skip mode, respectively.

[0026] With regard to the description outlined above, it should be noted that the division of image 20 into block 40 may not represent the finest resolution at which the encoding parameters are determined for image 20. Rather, the encoder 10 may accompany each block 40 with further division information to subblocks 50 and 60, i.e., sample subsets, to signal within the bitstream 30 one of the supported division patterns for dividing block 40. In this case, simultaneous merge / skip decisions are performed by the encoder 10 on a block 40 basis, but the encoding parameters associated with, for example, separate auxiliary merge and / or skip mode decisions are determined with respect to image 20 on a sub-block 40 basis, i.e., on a sub-block 50 and 60 basis, in the block 40 shown as an example in Figure 1. Of course, a non-divided mode can represent one of the supported division patterns, thereby resulting in the encoder 10 simply determining one set of encoding parameters for block 40. Regardless of the number of subblocks 50 and 60 in each splitting pattern, the merge decision can be applied to all subblocks, i.e., one or more of its subblocks. That is, if merge is activated with respect to block 40, this activation can be effective for all subblocks. According to embodiments further outlined below, the above common state that commonly signals the activation of merge and skip modes may additionally and simultaneously signal non-splitting patterns among the supported splitting patterns for block 40, such that when a flag or syntax element takes this state, no further transmission of splitting information for the current block is required. Of course, alternatively, any other splitting patterns among the supported splitting patterns can also be indicated simultaneously in addition to the activation of merge and skip modes.

[0027] According to some embodiments of the present invention, the encoder 10 avoids the bit efficiency penalty resulting from the common use of block partitioning of block 40 and the consolidation of subblocks 50 and 60. More precisely, the encoder 10 can determine, for example, whether further partitioning of block 40 is better in terms of rate distortion optimization, and which of the supported partitioning patterns must be used for the current block 40 in order to apply a particular encoding parameter to the granularity set or defined within the current block 40 of the image 20. As outlined in more detail below, the encoding parameter can represent, for example, predictive parameters such as inter-predictive parameters. Such inter-predictive parameters may include, for example, a reference image index, a motion vector, and so on. Supported partitioning patterns may include, for example, a non-partitioning mode, i.e., an option in which the current block 40 is not further subdivided; a horizontally partitioning mode, i.e., an option in which the current block 40 is subdivided along a horizontally extending line into an upper or topmost and lower or bottommost section; and a vertically partitioning mode, i.e., an option in which the current block 40 is vertically subdivided along a vertically extending line into a left and right section. In addition, supported partitioning patterns may also include an option in which the current block 40 is further regularly subdivided into four more blocks, each considered to be a quarter of the current block 40. Furthermore, the partitioning may relate to all blocks 40 of image 20, or only to an appropriate subset thereof, such as those having a specific encoding mode associated with it, such as interpredictive mode. Similarly, it should be noted that integration may, by itself, only be available with respect to a specific block, for example, one encoded in interpredictive mode. According to embodiments further outlined below, the commonly interpreted state described above also simultaneously signals that each block is in interpredictive mode rather than intrapredictive mode.Therefore, with respect to block 40, one state of the flag described above may signal that this block is an interpredictive coding block in which both merge and skip modes are activated and will not be further divided. However, as an auxiliary decision when the flag is in other states, each partition or sample subset 50 and 60 may be individually and simultaneously determined by further flags within bitstream 30 to signal whether merge is applied to each partition 50 and 60. Furthermore, different subsets of supported division modes may be available with respect to block 40, which are determined, for example, by the block size, and the level of subdivision of block 40, if it is a multi-tree subdivision leaf block, either combined or individually.

[0028] That is, the subdivision of image 20 into blocks to obtain block 40 in particular can be fixed or signaled within the bitstream. Similarly, the subdivision pattern currently used to further subdivide block 40 can be signaled within the bitstream 30 in the form of subdivision information. Thus, this subdivision information can be considered as a kind of extension of the subdivision of image 20 into blocks 40. On the one hand, further relevance of the original granularity of the subdivision of image 20 into blocks 40 can still be maintained. For example, the encoder 10 can be configured to signal within the bitstream 30 the encoding mode used for each part of image 20 or block 40 at the granularity defined by block 40, while the encoder 10 can be configured to vary the encoding parameters of each encoding mode within each block 40 at an increased (fine) granularity defined by each subdivision pattern selected for each block 40. For example, the encoding modes transmitted at the granularity of block 40 can be distinguished among intra-prediction modes such as time inter-prediction modes, inter-view prediction modes, and inter-prediction modes. The type of encoding parameters associated with one or more subblocks (partitions) resulting from the division of each block 40 depends on the encoding mode assigned to each block 40. For example, with respect to intra-encoded blocks 40, the encoding parameters may include spatial orientation regarding which image content from the portion decoded before image 20 is used to fill each block 40. In the case of inter-encoded blocks 40, the encoding parameters may include, in particular, motion vectors for motion compensation prediction.

[0029] Figure 1 shows, as an example, the current block 40 as subdivided into two subblocks 50 and block 60. In particular, the vertical partitioning mode is shown as an example. The smaller blocks 50 and block 60 can also be called subblocks 50 and subblock 60, or partitions 50 and partition 60, or prediction units 50 and prediction units 60. In particular, when one of the signaled supported partitioning patterns identifies the subdivision of the current block 40 into two or more further blocks 50 and block 60, the encoder 10 can be configured to exclude, for all further blocks of subblocks 50 and subblock 60 except the first subblock in the coding order, coding parameter candidates from the set of coding parameter candidates for each subblock, which have coding parameters that are the same as those associated with any of the subblocks that, when consolidated into each subblock, would become one of the supported partitioning patterns. More precisely, for each of the supported partitioning patterns, the coding order is determined among the resulting one or more partitions 50 and partition 60. In Figure 1, the coding order is indicated by arrow 70, which, for example, determines that the left partition 50 is coded before the right partition 60. In the case of horizontal split mode, it may be determined that the upper partition is coded before the lower partition. In any case, with respect to the second partition 60 in coding order 70, the encoder 10 is configured to exclude coding parameter candidates from the set of coding parameter candidates for each second partition 60 that have the same coding parameters as the first partition 50, in order to avoid the result of this merger, i.e., both partitions 50 and 60 would have the same coding parameters as associated with it, which could indeed occur equally by selecting a non-split mode for block 40 at a low coding rate.

[0030] More precisely, the encoder 10 may be configured to use block merging in an efficient manner along with block partitioning. As far as block merging is concerned, the encoder 10 may determine each set of coding parameter candidates for each partition 50 and partition 60. The encoder may be configured to determine a set of coding parameter candidates for each partition 50 and partition 60 based on the coding parameters associated with a previously decoded block. In particular, at least some of the coding parameter candidates in the set of coding parameter candidates may be equal to the coding parameters of a previously decoded partition, i.e., they may be adopted from the coding parameters of a previously decoded partition. In addition, or instead, at least some of the coding parameter candidates may be obtained from coding parameter candidates associated with two or more previously coded partitions by a suitable combination of medians, means, etc. However, since the encoder 10 is configured to determine a reduced set of coding parameter candidates and, if one or more of these coding parameter candidates remain after elimination, to perform a selection of one of the remaining uneliminated coding parameter candidates for each non-first partition 60 in order to set the coding parameter associated with each partition, depending on one of the uneliminated or selected coding parameter candidates, the encoder 10 is configured to perform elimination in such a way that coding parameter candidates that would efficiently lead to the recombination of partition 50 and partition 60 are eliminated. In other words, the efficient partitioning situation is thus effectively avoided as it would result in a collection of syntaxes that would be more complex to encode than if this partitioning were directly signaled using only the partitioning information.

[0031] Furthermore, as the set of coding parameter candidates becomes smaller, the amount of auxiliary information required to encode the integration information into the bitstream 30 can decrease due to the smaller number of elements in these candidate sets. In particular, since the decoder can determine and then reduce the set of coding parameter candidates in the same way as the encoder in Figure 1, the encoder 10 in Figure 1 can, for example, use fewer bits to insert syntax elements into the bitstream 30 and identify which of the unremoved coding parameters will be used for integration. Naturally, the introduction of syntax elements into the bitstream 30 can be completely suppressed if the number of unremoved coding parameter candidates for each partition is simply one. In any case, by integration, i.e., by setting the coding parameters associated with each partition depending on one of the remaining or selected unremoved coding parameter candidates, the encoder 10 can suppress the insertion of entirely new coding parameters for each partition into the bitstream 30, thereby reducing auxiliary information as well. According to some embodiments of the present invention, the encoder 10 may be configured to signal refinement information within the bitstream 30 for refining one of the remaining or selected candidate encoding parameters for each partition.

[0032] The above possibility of reducing the list of merger candidates allows the encoder 10 to be configured to determine merger candidates that will be eliminated by comparing their encoding parameters with those of the partitions, and the resulting mergers will produce another supported partition pattern. For example, if the encoding parameters of the left partition 50 form one element of the set of encoding parameter candidates for the right partition 60, this method of processing encoding parameter candidates will efficiently eliminate at least one encoding parameter candidate as shown in Figure 1. However, further encoding parameter candidates may also be eliminated if they are equal to the encoding parameters of the left partition 50. However, according to another embodiment of the present invention, the encoder 10 may be configured to determine a set of candidate blocks for each partition from the second onward in the encoding order by removing that or those candidate blocks from this set of candidate blocks that, when merged into each partition, would result in one of the supported partition patterns. In a sense, this means the following: The encoder 10 can be configured to determine a combined candidate for each of partitions 50 or 60 (i.e., the first and the next in coding order) such that each element of the candidate set has exactly one partition in either the currently coded block 40 or the previously coded block 40, which is associated with the candidate in that the candidate adopts each coding parameter of its associated partition. For example, each element of the candidate set may be equal to one of these coding parameters of a previously coded partition, i.e., it may be adopted from among them, or at least it may be obtained from the coding parameters of just one of these previously coded partitions, such as by additional scaling or refinement with additionally transmitted refinement information.However, the encoder 10 may also be configured to add further elements or candidates to this type of candidate set, i.e., to add coding parameter candidates obtained from a combination of coding parameters of one or more previously coded partitions, or, by modification, from the coding parameters of one previously coded partition, such as by taking only the coding parameters of one movement parameter list. For “combined” elements, there is no one-to-one relationship between the coding parameters of each candidate element and each partition. According to the first modification of the description in Figure 1, the encoder 10 can be configured to remove all candidates from the entire candidate set, whose coding parameters are equal to the coding parameters of partition 50. According to the latter modification of the description in Figure 1, the encoder 10 can be configured to remove only the elements of the candidate set that are associated with partition 50. If both views are agreed upon, the encoder 10 can be configured to remove candidates from a portion of the candidate set that shows a one-to-one relationship with several (e.g., adjacent) previously coded partitions without extending the removal (and search for candidates with equal coding parameters) to the rest of the candidate set with the coding parameters obtained by the combination. However, if one combination naturally leads to a redundant representation, this can be resolved by removing the redundant encoding parameter from the list, or by performing a redundancy check on the combined candidates as well.

[0033] Before describing an embodiment of the decoder that fits the above embodiment in Figure 1, a more detailed embodiment of the encoding apparatus, i.e., the encoder, shown in Figure 1, is outlined in more detail below with respect to Figure 2. Figure 2 shows an encoder that includes a subdivider 72 configured to subdivide the image 20 into blocks 40, a merger 74 configured to consolidate the blocks 40 into a group of one or more sample sets as outlined above, an encoder or encoding stage 76 configured to encode the image 20 using encoding parameters that vary across the entire image 20 in units of groups of sample sets, and a stream generator 78. The encoder 76 is configured to encode the image 20 by predicting the image 20 and encoding the predicted residuals for a given block. That is, the encoder 76 encodes the predicted residuals for some, but not all, of the blocks 40, as described above. Rather, some of them activate a skip mode. The stream generator 78 is configured to insert prediction residuals and coding parameters into the bitstream 30, along with one or more syntax elements, for at least one subset of blocks 40, signaling whether each block 40 is merged into one of the groups with another block or whether each block uses a skip mode. As described above, the subdivision information that forms the basis of the subdivision of the subdivider 72 can be encoded into the bitstream 30 with respect to the image 20 by the stream generator 78. This is shown by the dashed line in Figure 2. The merger decision by the merger 74 and the skip mode decision by the encoder 76 are, as outlined above, currently encoded into the bitstream 30 by the stream generator 78 so as to signal that one of the possible states of one or more syntax elements of block 40 is now to be merged into one of the groups of blocks with another block of the image 20, and that the bitstream 30 does not have encoded and inserted prediction residuals. The stream generator 78 may use, for example, entropy coding to perform its insertion.The subdivider 72 may be responsible for subdividing the image 20 into block 40, as well as for any further division into partitions 50 and 60, respectively. The merger 74 may be responsible for the integration decision outlined above, while the encoder 76 may, for example, determine the skip mode for block 40. Naturally, all of these decisions affect the rate / distortion measure in the combination, and therefore the device 10 may be configured to try out several decision options to determine which option is preferred.

[0034] After describing an encoder according to an embodiment of the present invention with respect to Figures 1 and 2, an apparatus for decoding according to the embodiment, namely a decoder 80, will be described with respect to Figure 3. The decoder 80 in Figure 3 is configured to decode a bitstream 30 having an encoded image 20, as described above. In particular, the decoder 80 is configured to respond in common to the aforementioned flags in the bitstream 30 with respect to the current sample set or block 40, a first determination of whether the encoding parameters currently associated with block 40 will be set by a combined candidate or extracted from the bitstream 30, and a second determination of whether the current block 40 of image 20 will be reconstructed without residual data, based only on a predicted signal corresponding to the encoding parameters currently associated with block 40, or by refining the predicted signal corresponding to the encoding parameters currently associated with block 40 with residual data within the bitstream 30.

[0035] In other words, the function of the decoder largely coincides with that of the encoder described with respect to Figures 1 and 2. For example, decoder 80 may be configured to perform subdivision of image 40 into blocks 40. This subdivision may be known to decoder 80 by default, or decoder 80 may be configured to extract each subdivision information from bitstream 30. Whenever block 40 is merged, decoder 80 may be configured to obtain the encoding parameters associated with that block 40 by setting its encoding parameters by merge candidates. To determine merge candidates, decoder 80 may perform the decision outlined on a set or list of merge candidates in exactly the same way that the encoder did. This may even include a reduction of the pre-set / list of merge candidates to avoid the redundancy outlined above between block division and block merge, according to some embodiments of the present application. Whenever merge is activated, selection from the determined set or list of merge candidates may be performed by decoder 80 by extracting each merge index from bitstream 30. The integration index points to the integration candidate to be used from the (reduced) set or list of integration candidates determined as described above. Furthermore, as also described above, the decoder 80 can also be configured to multiply block 40 by a partition according to one of the supported partition patterns. Naturally, one of these partition patterns may include an unpartitioned mode in which block 40 is not further partitioned. In the case of well-described flags that take a commonly defined state indicating the activation of integration and skip modes for a particular block 40, the decoder 80 can be configured to reconstruct the current block 40 based only on the prediction signal rather than in combination with any residual signal. In other words, the decoder 80 in that case simply reconstructs the image 20 within the current block 40 by using the prediction signal extracted from the encoding parameters of the current block, suppressing the extraction of residual data for the current block 40.As already explained above, the decoder 80 can also interpret the common state of the flag as a signaling for the current block 40 that this block is an interpredicted block and / or a block that has not been further divided. That is, if the flag of the current block 40 in the bitstream 30 signals that the coding parameters associated with the current block 40 will be set using integration, the decoder 80 may be configured to obtain the coding parameters associated with the current block 40 by setting these coding parameters according to the integration candidates, and to reconstruct the current block 40 of image 20 based solely on the prediction signal corresponding to the coding parameters of the current block 40 without any residual data. However, if the flag in question signals that block 40 is not subject to integration or that skip mode is not used, the decoder 80 may respond to another flag in the bitstream 30, and as a result, the decoder 80 may rely on this other flag to obtain the coding parameters associated with the current block by setting it for each integration candidate, obtain residual data about the current block from the bitstream 30, and reconstruct the current block 40 of image 20 based on the predicted signal and residual data, or extract the coding parameters associated with the current block 40 from the bitstream 30, obtain residual data about the current block 40 from the bitstream 30, and reconstruct the current block 40 of image 20 based on the predicted signal and residual data. As outlined above, the decoder 80 may be configured to predict the presence of another flag in the bitstream 30 only if the first flag does not take a commonly signaling state that simultaneously signals the activation of integration and skip mode. Only at that time does the decoder 80 extract another flag from the bitstream to check whether integration occurs without skip mode.Of course, another method is for the decoder 80 to use this third flag, which signals whether the skip mode is active or inactive, to now wait for another third flag in the bitstream 30 for block 40 when the second flag signals the inactivity of integration.

[0036] Similar to Figure 2, Figure 4 shows a possible embodiment of the apparatus for decoding Figure 3. Thus, Figure 4 shows the apparatus for decoding, namely the decoder 80, which includes a subdivider 82 configured to subdivide the image 20 encoded in the bitstream 30 into blocks 40, a merger 84 configured to combine the blocks 40 into groups of one or more blocks, a decoder 86 configured to decode or reconstruct the image 20 using encoding parameters that vary across the entire image 20 in units of groups of sample sets, and an extractor 88. The decoder 86 is also configured to decode the image 20 by predicting the image 20 with respect to a default block 40, i.e., the one with the skip mode switched off, decoding the prediction residual for the default block 40, and combining the prediction resulting from predicting the image 20 with the prediction residual. The extractor 88 is configured to extract predicted residuals and coding parameters from a bitstream 30 that signals whether each block 40 will be merged with another block 40 into one group, along with one or more syntax elements for each of at least a subset of blocks 40. Here, the merger 84 is configured to perform its merger in response to one or more syntax elements, one of which of the possible states of the one or more syntax elements signals that each block 40 is to be merged with another block 40 into one group of blocks, and that it does not have any coded and inserted predicted residuals in the bitstream 30.

[0037] Thus, comparing Figure 4 with Figure 2, subdivider 82 acts like subdivider 72, reversing the subdivision caused by subdivider 72. Subdivider 82 either knows about the subdivision of image 20 by default or extracts subdivision information from bitstream 30 via extractor 88. Similarly, merger 84 forms the merger of block 40 and is activated with respect to block 40 and block partitions within bitstream 30 via the signaling outlined above. Decoder 86 performs the generation of a predicted signal for image 20 using the encoded parameters within bitstream 30. In the case of merger, decoder 86 either replicates the encoded parameters of the current block 40 or current block partition from an adjacent block / partition, or otherwise sets its encoded parameters according to the merger candidate.

[0038] As outlined above, the extractor 88 is configured to interpret one of the possible states of a flag or syntax element relating to the current block as a signal that simultaneously signals the activation of the integration and skip modes. Simultaneously, the extractor 88 can also interpret the state to signal one of the default supported partition patterns for the current block 40. For example, a default partition pattern could be a non-partition mode in which the block 40 remains unpartitioned, thus forming a single partition itself. Therefore, the extractor 88 expects that if each flag or syntax element does not take on a state that simultaneously signals, the bitstream 30 will contain partition information that simply signals the partitioning of the block 40. As outlined in more detail below, the partition information may be transmitted within the bitstream 30 via a syntax element that currently controls the encoding mode of the block 40, i.e., partitioning the block 40 into inter-encoded and intra-encoded parts in parallel. In that case, the common signaling state of the first flag / syntax element can also be interpreted as signaling for the interpredictive coding mode. For each partition resulting from the signaled partition information, the extractor 88 may extract another merge flag from the bitstream if the first flag / syntax element for block 40 does not take a common signaling state that simultaneously signals the activation of merge and skip modes. In that case, the skip mode may necessarily be interpreted by the extractor 88 to be switched off, and merge may be individually activated by the bitstream 30 with respect to the partitions, but the residual signal is now extracted from the bitstream 30 for block 40.

[0039] Thus, the decoder 80 in Figure 3 or Figure 4 is configured to decode the bitstream 30. As described above, the bitstream 30 can signal one of the supported partitioning patterns for the current block 40 of the image 20. If one of the signaled supported partitioning patterns identifies the subdivision of the current block 40 into two or more partitions 50 and partition 60, the decoder 80 may be configured to exclude, for all partitions in the coding order 70 except the first partition 50, i.e., for partition 60 in the examples shown in Figures 1 and 3, from the set of coding parameter candidates for each partition, those that, when combined with each partition, have coding parameters that are the same as or equal to the coding parameters associated with one of the partitions, i.e., one of the supported partitioning patterns that was not signaled in the bitstream 30.

[0040] For example, the decoder 80 can be configured to set the coding parameters associated with each partition 60 depending on one of the unremoved coding parameter candidates, provided that the number of unremoved coding parameter candidates is not zero. For example, the decoder 80 sets the coding parameters of each partition 60 to be equal to one of the unremoved coding parameter candidates, with or without additional refinement and / or with or without scaling by the temporal distance to which the coding parameters are associated. For example, the coding parameter candidates to be integrated from the unremoved candidates may have a separate reference image index associated with it, distinct from the reference image index explicitly signaled within the bitstream 30 for partition 60. In this case, the coding parameters of the coding parameter candidates can each define a motion vector associated with the reference image index, and the decoder 80 can be configured to scale the motion vector of the last selected unremoved coding parameter candidate according to the ratio between the two reference image indices. Thus, according to the above modification, the coding parameters subject to integration include motion parameters, while the reference image indices are separated from them. However, as described above, according to another embodiment, the reference image index may also be part of the encoding parameters that follow the integration.

[0041] The fact that merge behavior may be limited to inter-predicted block 40 applies equally to the encoders in Figures 1 and 2 and the decoders in Figures 3 and 4. Therefore, the decoder 80 and encoder 10 can be configured to currently support intra and inter-prediction modes with respect to block 40, and to perform merge only when block 40 is currently encoded in inter-prediction mode. Thus, only the encoding / prediction parameters of this type of inter-predicted previously encoded partition can be used to determine / construct the candidate list.

[0042] As already mentioned above, the coding parameters may also be prediction parameters, and the decoder 80 can be configured to use the prediction parameters of partitions 50 and 60 to obtain prediction signals for each partition. Naturally, the encoder 10 also performs prediction signal extraction in a similar manner. However, the encoder 10 also sets the prediction parameters along with all other syntax within the bitstream 30 to obtain some optimizations in the sense of appropriate optimization.

[0043] Furthermore, as already explained above, the encoder can be configured to insert an index into the (unexcluded) coding parameter candidates only if the number of (unexcluded) coding parameter candidates for each partition is greater than 1. Thus, the decoder 80 can be configured to simply expect that, if the number of (unexcluded) coding parameter candidates is greater than 1, the bitstream 30 will contain a syntax element that specifies which of the (unexcluded) coding parameter candidates will be used for the merger, for example, depending on the number of (unexcluded) coding parameter candidates for partition 60. However, if the candidate set totals less than 2, as described above, this can usually be excluded by limiting the reduction of the candidate set to those candidates obtained by adopting or extracting the coding parameters of just one previously coded partition, thereby expanding the list / set of candidates using combined coding parameters, i.e., parameters obtained by combinations of coding parameters of one or more or two or more previously coded partitions. Similarly, the reverse is also possible, i.e., excluding all coding parameter candidates that have the same values ​​as those of partitions that are usually other supported split patterns.

[0044] With regard to the decision, the decoder 80 operates as the encoder 10 does. That is, the decoder 80 can be configured to determine a set of merge candidates for a partition or a partition in block 40 based on the encoding parameters associated with a previously decoded partition. That is, the encoding order is determined not only within partitions 50 and 60 of each block 40, but also within block 40 of the image 20 itself. All partitions that were encoded before partition 60 thus serve as a criterion for determining a set of merge candidates for any subsequent partition, such as partition 60 in the case of Figure 3. As also described above, the encoder and decoder can restrict the determination of a set of merge candidates to partitions that are in a particular spatial and / or temporal adjacency. For example, the decoder 80 can be configured to determine a set of merge candidates based on the encoding parameters associated with a previously decoded partition adjacent to the current partition, and such partitions can be located outside and inside the current block 40. Of course, the determination of merge candidates can also be performed for the first partition in the encoding order. Simply removing partitions may not be done.

[0045] In accordance with the description in Figure 1, the decoder 80 can be configured to determine a set of coding parameter candidates for each non-first partition 60 from the first set of previously decoded partitions, excluding those coded in intra-predictive mode.

[0046] Furthermore, if the encoder introduces subdivision information into the bitstream in order to subdivide the image 20 into blocks 40, the decoder 80 can be configured to revert the subdivision of the image 20 into these types of encoded blocks 40 according to the subdivision information in the bitstream 30.

[0047] With respect to Figures 1 to 4, it should be noted that the residual signal for block 40 can now be transmitted via the bitstream 30 with a granularity that may differ from the granularity defined by the partition with respect to the encoding parameters. For example, for blocks where the skip mode is not activated, the encoder 10 in Figure 1 can be configured to subdivide block 40 into one or more transform blocks in a manner parallel to or independent of the partitioning into partitions 50 and 60. The encoder can signal each transform block subdivision for block 40 by further subdivision information. The decoder 80 can then be configured to reverse this further subdivision of block 40 into one or more transform blocks by further subdivision information of the bitstream, and derive the residual signal for block 40 now from the bitstream in units of these transform blocks. The significance of transform block subdivision may be that transforms such as DCT in the encoder and corresponding inverse transforms such as IDCT in the decoder are performed individually within each transform block of block 40. To reconstruct the image 20 within block 40, the encoder 10 combines the predicted and residual signals obtained by applying coding parameters in each partition 50 and partition 60, for example, by adding them together. However, it should be noted that residual coding does not involve transforms or inverse transforms, and the predicted residuals are instead coded in the spatial domain, for example.

[0048] Before describing further possible details of the following further embodiments, the possible internal structures of the encoders and decoders in Figures 1 to 4 are described with respect to Figures 5 and 6. However, the mergers and subdividers are not shown in these figures in order to focus on the nature of hybrid coding. Figure 5 shows, as an example, how encoder 10 may be constructed internally. As shown in the figure, encoder 10 may include a subtractor 108, a transducer 100, and a bitstream generator 102, which can perform entropy coding as shown in Figure 5. Elements 108, 100, and 102 are connected in series between an input 112 that receives an image 20 and an output 114 that outputs the aforementioned bitstream 30. In particular, the subtractor 108 has its non-inverting input connected to input 112, the transducer 100 is connected between the output of the subtractor 108 and the first input of the bitstream generator 102, and the bitstream generator 102 has an output connected to output 114. The encoder 10 in Figure 5 further includes an inverse converter 104 and an adder 110, which are connected in series to the output of the converter 100 in the order described. The encoder 10 further includes a predictor 106 connected between the output of the adder 110 and a further input of the adder 110 and the inverting input of the subtractor 108.

[0049] The elements in Figure 5 interact as follows: The predictor 106 is applied to the inverting input of the subtractor 108 to predict a portion of the image 20 using the prediction result, i.e., the prediction signal. The output of the subtractor 108 then shows the difference between the prediction signal and each portion of the image 20, i.e., the residual signal. The residual signal follows the transform coding of the converter 100. That is, the converter 100 can perform a transform such as DCT and subsequent quantization of the transformed residual signal, i.e., the transform coefficients, in order to obtain the transform coefficient levels. The inverse converter 104 reconstructs the final residual signal output by the converter 100 in order to obtain a reconstructed residual signal corresponding to the residual signal input to the converter 100, excluding information loss for the quantization of the converter 100. The sum of the reconstructed residual signal and the prediction signal as the output of the predictor 106 results in a reconstruction of each portion of the image 20, which is sent from the output of the adder 110 to the input of the predictor 106. The predictor 106 operates in various modes as described above, such as intra-prediction mode and inter-prediction mode. The prediction mode and corresponding coding or prediction parameters applied by the predictor 106 to obtain the predicted signal are sent by the predictor 106 to the entropy encoder 102 for insertion into the bitstream.

[0050] Possible embodiments of the internal structure of the decoder 80 in Figures 3 and 4, corresponding to the possibilities shown in Figure 5 with respect to the encoder, are shown in Figure 6. As shown in the figure, the decoder 80 may include a bitstream extractor 150, an inverse converter 152, and an adder 154, which can be implemented as an entropy decoder as shown in Figure 6, connected between the input 158 ​​and output 160 of the decoder in the order described. Furthermore, the decoder in Figure 6 includes a predictor 156 connected between the output of the adder 154 and its further input. The entropy decoder 150 is connected to the parameter input of the predictor 156.

[0051] To briefly explain the function of the decoder in Figure 6, the entropy decoder 150 is there to extract all the information contained in the bitstream 30. The entropy coding scheme used may be variable-length coding or arithmetic coding. In this way, the entropy decoder 150 reverts the bitstream transformation coefficient levels that represent the residual signal and sends it to the inverse decoder 152. Furthermore, the entropy decoder 150 acts as the extractor 88 described above, reverting all coding modes and associated coding parameters from the bitstream and sending it to the predictor 156. In addition, partitioned and integrated information is extracted from the bitstream by the extractor 150. The inverse-transformed, i.e., reconstructed residual signal and the predicted signal obtained from the predictor 156 are combined, for example by being added by the adder 154, and then the reconstructed signal thus reverted is output at output 160 and sent to the predictor 156.

[0052] As is evident from comparing Figures 5 and 6, elements 152, 154, and 156 functionally correspond to elements 104, 110, and 106 in Figure 5.

[0053] In the above description of Figures 1 to 6, several different possibilities were presented with respect to the possible subdivision of image 20 and the corresponding granularity when varying some of the parameters related to encoding image 20. This type of possibility is also described with respect to Figures 7a and 7b. Figure 7a shows a portion of image 20. According to the embodiment of Figure 7a, the encoder and decoder are first configured to subdivide image 20 into tree root blocks 200. Such tree root blocks are shown in Figure 7a. The subdivision of image 20 into tree root blocks is done regularly in rows and columns, as shown by the dotted lines. The size of the tree root blocks 200 can be selected by the encoder and signaled to the decoder by the bitstream 30. Alternatively, the size of these tree root blocks 200 can be fixed by default. The tree root blocks 200 are subdivided using quad-tree subdivision to produce the distinguished blocks 40 described above, which may be called encoded blocks or encoded units. These encoded blocks or encoded units are drawn by the thin solid lines in Figure 7a. This causes the encoder to add subdivision information to each tree root block 200 and insert that subdivision information into the bitstream. This subdivision information indicates how the tree root block 200 will be subdivided into blocks 40. At the granularity of these blocks 40, and in units thereof, the prediction mode changes within image 20. As described above, each block 40, or each block having a specific prediction mode such as interprediction mode, is accompanied by subdivision information about which supported subdivision pattern will be used for each block 40. However, in this regard, it should be noted that the aforementioned flag / syntax elements can simultaneously signal one of the supported subdivision modes for each block 40, so that when a common signaling state is taken, the explicit transmission of separate subdivision information for each block 40 can be suppressed on the encoder side and therefore cannot be expected on the decoder side.In the case shown in Figure 7a, for many coding blocks 40, a non-partitioned mode is selected so that coding block 40 coincides with the spatially corresponding partition. In other words, coding block 40 is simultaneously a partition having each set of prediction parameters associated with it. The type of prediction parameters then depends on the mode associated with each coding block 40. However, another coding block is shown to be further subdivided, for example. For example, coding block 40 in the top right corner of tree root block 200 is shown to be subdivided into four partitions, while coding block in the bottom right corner of tree root block 200 is shown exemplarily to be vertically subdivided into two partitions. The subdivision for partitioning is indicated by dotted lines. Figure 7a also shows the coding order within the partitions thus defined. As shown in the figure, a depth-first scan order is used. Across the tree root block boundaries, the coding order can proceed in a scan order in which rows of tree root block 200 are scanned row by row from top to bottom of image 20. This method makes it possible to maximize the possibility that a particular partition has encoded partitions before it is adjacent to its upper and left boundaries. Each block 40, or each block having a particular prediction mode such as interprediction mode, may have a merge switch indicator in the bitstream indicating whether merge is working for the corresponding partition within it. It should be noted that the division of a block into partition / prediction units can be limited to a maximum of two partitions, with the exception that this rule is only made with respect to the smallest possible block size of block 40. This avoids redundancy between the subdivision information for subdividing image 20 into block 40 and the division information for subdividing block 40 into partitions when using quad-tree subdivision to obtain block 40. Alternatively, only a division into one or two partitions, with or without asymmetrical ones, may be allowed.

[0054] Figure 7b shows a subdivision tree. The solid lines represent the subdivision of the tree root block 200, while the dotted lines represent the division of the leaf blocks of the quad tree subdivision, which is the encoded block 40. In other words, the division of the encoded block represents a kind of extension of the quad subdivision.

[0055] As already mentioned above, each coding block 40 may be subdivided in parallel with the transform blocks such that the transform blocks may represent different subdivisions of each coding block 40. For each of these transform blocks not shown in Figures 7a and 7b, the transform for transforming the residual signal of the coding block may be performed separately.

[0056] Further embodiments of the present invention are described below. While the above embodiments focused on the relationship between block integration and block splitting, the following description also includes aspects of the present invention relating to other encoding principles known in current codecs, such as skip / direct mode. Nevertheless, the following description is not to be considered merely a description of other embodiments, i.e., embodiments separated from those described above. Rather, the following description also clarifies the details of possible embodiments relating to the above embodiments. Therefore, the following description uses the reference numerals already shown in the figures above, and as a result, each possible embodiment described below also defines possible variations of the above embodiments. Most of these variations can be individually adapted to the above embodiments.

[0057] In other words, embodiments of the present invention describe a method for reducing the rate of auxiliary information in image and video coding applications with respect to a set of samples by signaling a combination of integration and the absence of residual data. In other words, the rate of auxiliary information in image and video coding applications is reduced by combining syntax elements that indicate the use of an integration scheme and syntax elements that indicate the absence of residual data.

[0058] Furthermore, before describing these variations and further details, an overview of image and video codecs is provided.

[0059] In image and video coding applications, a sample array associated with an image is typically divided into specific sets of samples (or sample sets) that may represent arbitrarily shaped regions, triangles, or rectangular or square blocks containing other shapes, or other sets of samples. The subdivision of a sample array may be fixed by syntax, or the subdivision may be signaled (at least partially) within the bitstream. To keep the auxiliary information rate for signaling subdivision information low, syntax typically allows only a limited number of choices, resulting in simple subdivisions such as subdividing a block into smaller blocks. Commonly used subdivision schemes include dividing a square block into four smaller square blocks, or into two rectangular blocks of the same size, or into two rectangular blocks of different sizes. Here, the subdivision actually used is signaled within the bitstream. A sample set is associated with specific coding parameters that can identify predictive information or residual coding modes, etc. In video coding applications, subdivision is often done for motion representation. All samples in a block (within a split pattern) relate to the same set of motion parameters, which may include parameters that identify the type of prediction (e.g., List 0, List 1, or bidirectional prediction; and / or translation or affine prediction or prediction using a different motion model), parameters that identify the reference image used, parameters that identify the motion relative to the reference image, which are typically sent to the predictor as a difference (e.g., displacement vector, affine motion parameter vector, or motion parameter vector for other motion models), the precision of the motion parameters (e.g., precision of half-sample or quarter-sample), parameters that identify the weighting of the reference sample signal (e.g., for illumination compensation), or parameters that identify the interpolation filter used to obtain the motion-compensated prediction signal for the current block. For each sample set, it is assumed that individual coding parameters are sent (e.g., to identify the prediction and / or residual coding).To obtain improved coding efficiency, the present invention provides a method and specific embodiments for integrating two or more sample sets into a so-called group of sample sets. All sample sets in this type of group share the same coding parameters, which can then be transmitted together with one of the sample sets in the group. In this way, the coding parameters do not need to be transmitted individually for each sample set in the group of sample sets; instead, the coding parameters are transmitted only once for all groups of sample sets.

[0060] As a result, the rate of auxiliary information for transmitting coding parameters is reduced, and the overall coding efficiency is improved. Alternatively, an additional refinement for one or more coding parameters can be transmitted for one or more sample sets in a group of sample sets. That refinement can be applied to all sample sets in the group, or only to the sample sets to which it is transmitted.

[0061] Some embodiments of the present invention combine the division and integration processes of a block into various subblocks 50, 60 (as described above). Typically, an image or video encoding system supports various division patterns for a block 40. For example, a square block may be left undivided or divided into four square blocks of the same size, two rectangular blocks of the same size (divided vertically or horizontally), or rectangular blocks of different sizes (vertically or horizontally). Typical partition patterns described are shown in Figure 8. In addition to the above description, division may even involve one or more levels of division. For example, a square subblock may also be further divided using the same division pattern, at the option of. The problem that arises when this type of division process is combined with an integration process that allows a block (square or rectangular) to be integrated with, for example, one of its adjacent blocks is that the resulting division can be obtained by different combinations of division patterns and integration signals. Thus, the same information can be transmitted in a bitstream using different codewords, which obviously does not reach the optimal state with respect to encoding efficiency. As a simple example, consider a square block that is not further subdivided (as shown in the upper left corner of Figure 8). This subdivision can be directly signaled by sending a syntax element indicating that this block 40 is not subdivided. However, the same pattern can also be signaled by sending a syntax element that specifies that this block is subdivided into, for example, two vertically (or horizontally) arranged rectangular blocks 50, 60. Then, integration information can be sent that specifies that the second of these rectangular blocks is merged with the first rectangular block, resulting in exactly the same subdivision as when signaling that the block is not further subdivided. The same can be achieved by first specifying that the block is subdivided into four square subblocks, and then sending integration information that efficiently merges all these four blocks. This concept obviously does not reach the optimal state (because we have different codewords to signal the same thing).

[0062] Some embodiments of the present invention increase coding efficiency due to a combination of the concept of reducing the auxiliary information rate and thus providing different partitioning patterns for blocks and the concept of integration. Looking at the example partitioning patterns in Figure 8, the "simulation" of blocks that were not further partitioned by either of the partitioning patterns using two rectangular blocks can be avoided by prohibiting (i.e., excluding from the bitstream syntax specification) cases where the rectangular block is integrated with the first rectangular block. Looking at the problem more deeply, it is also possible to "simulate" patterns that were not subdivided by integrating the second rectangle with any other adjacent (i.e., a rectangular block other than the first) relating to the same parameters as the first rectangular block (e.g., information for identifying predictions). When these integration parameters result in a pattern that can also be obtained by signaling one of the supported partitioning patterns as a result, redundancy can be avoided by conditioning the transmission of integration information in a way that the transmission of specific integration parameters is excluded from the bitstream syntax. For example, if the current splitting pattern identifies subdivision into two rectangular blocks, as shown in Figures 1 and 3, before sending integration information for the second block, i.e., block 60 in Figures 1 and 3, it is possible to check which of the possible integration candidates has the same parameters (e.g., parameters for identifying the predicted signal) as the first rectangular block, i.e., block 50 in Figures 1 and 3. Then, all candidates with the same motion parameters (including the first rectangular block itself) are excluded from the set of integration candidates. The codeword or flag sent to signal the integration information is fitted to the resulting set of candidates. If the set of candidates is empty due to the parameter check, the integration information cannot be sent. If the set of candidates consists of just one entry, it only signals whether that block is to be integrated, and that candidate does not need to be signaled as it can be obtained on the decoder side, etc.In the example above, the same concept is also used for a division pattern that divides a square block into four smaller square blocks. Here, the transmission of the integration flag is adapted in a way that neither the division pattern that does not specify the subdivision nor the two division patterns that specify subdivision into two rectangular blocks of the same size can achieve by the combination of integration flags. While the example above with a specific division pattern illustrates the most common concept, it is clear that the same concept can be used for other sets of division patterns (by avoiding the specification of a specific division pattern through the combination of other division patterns and corresponding integration information).

[0063] Another aspect that needs to be considered is that the integrated concept is in a sense analogous to the skip or direct modes found in video coding design. In skip / direct modes, motion parameters are essentially not sent for the current block but are inferred from spatial and / or temporal adjacencies. In a particular efficient concept of skip / direct modes, a list of motion parameter candidates (reference frame index, displacement vector, etc.) is generated from spatial and / or temporal adjacencies, and an index to this list is sent that specifies which of the candidate parameters will be selected. For bidirectional predicted blocks (or multiple hypothetical frames), another candidate may be signaled for each reference list. Possible candidates may include the block above the current block, the block to the left of the current block, the block to the upper left of the current block, the block to the upper right of the current block, a median predictor of various of these candidates, and a block (or other already coded block, or a combination derived from already coded blocks) located in the same position in one or more previous reference frames.

[0064] Combining skip / direct with the unified concept means that a block can be encoded using either skip / direct or unified mode. While the skip / direct and unified concepts are similar, there are differences between the two concepts, which are described in more detail in Section 1. The main difference between skip and direct is that skip mode further signals that no residual signal is transmitted. When the unified concept is used, a flag is typically transmitted that signals whether the block contains a non-zero conversion coefficient level.

[0065] To obtain improved coding efficiency, the embodiments described above and below combine signaling whether a sample set uses the coding parameters of another sample set and signaling whether residual signals are not transmitted for blocking. The combined flags indicate that a sample set uses the coding parameters of another sample set and that residual data is not transmitted. In this case, it is required that only one flag is transmitted, rather than two.

[0066] As described above, some embodiments of the present invention also provide the encoder with greater degrees of freedom to generate the bitstream, as the integration approach significantly increases the number of possibilities for selecting a partition for the image sample array without resulting in bitstream redundancy. Encoding efficiency can be improved because the encoder can choose from more options, for example, to minimize a particular rate distortion measure. As an example, some of the additional patterns that can be shown by a combination of subdivision and integration (e.g., the patterns in Figure 9) can be additionally tested (using the corresponding block sizes for motion estimation and mode determination), by pure partitioning (Figure 8), and the best pattern provided by partitioning and integration (Figure 9) can be selected based on a particular rate distortion measure. In addition, for each block, it can be tested whether integration using any of the already encoded candidate sets results in a reduction of a particular rate distortion measure, and the corresponding integration flag is set during the encoding process. In summary, there are several possibilities for operating the encoder. In a simple approach, the encoder can first determine the largest subdivision of the sample array (as the highest-level encoding scheme). This allows us to check, for each sample set, whether integration with other sample sets or other groups of sample sets reduces a particular rate-distortion cost measure. Here, the predictive parameters associated with the integrated group of sample sets can be re-estimated (for example, by performing a new motion search). Alternatively, the predictive parameters already determined for the current sample set and candidate sample sets (or groups of sample sets) for integration can be evaluated for the considered group of sample sets. In a broader approach, a particular rate-distortion cost measure can be evaluated for additional candidate groups of sample sets.As an exception, when testing various possible partitioning patterns (see, e.g., Figure 8), some or all of the patterns that can be represented by combinations of partitioning and merging (see, e.g., Figure 9) may also be tested. That is, for all patterns, specific motion estimation and mode determination processes are performed, and the pattern that yields the smallest rate distortion measure is selected. This process may also be combined with the low-complexity process described above. As a result, for the resulting blocks, it is additionally tested whether merging with respect to already encoded blocks (e.g., outside the patterns in Figures 8 and 9) results in a reduction of the rate distortion measure.

[0067] In the following, several possible detailed embodiments for the embodiments outlined above will be described, for example, with respect to the encoders in Figures 1, 2, and 5 and the decoders in Figures 3, 4, and 6. As already mentioned above, it is usable in image and video coding. As stated above, an image or a particular set of sample arrays for an image can be decomposed into blocks, which are associated with specific coding parameters. An image typically consists of multiple sample arrays. In addition, an image may be associated with additional auxiliary sample arrays, which can, for example, identify transparency information or depth maps. The sample arrays of an image (including auxiliary sample arrays) can be classified into one or more so-called plane groups, where each plane group consists of one or more sample arrays. The plane groups of an image can be coded independently, or, if the image is associated with multiple plane groups, by prediction from other plane groups of the same image. Each plane group is typically decomposed into blocks. Blocks (or corresponding blocks of sample arrays) are predicted by inter-image prediction or intra-image prediction. Blocks can have different sizes and may be square or rectangular. The division of an image into blocks can be fixed by syntax, or it can be signaled (at least partially) within the bitstream. Often, a syntax element is sent that signals the subdivision for blocks of a given size. This type of syntax element can specify whether and how the image is subdivided into smaller blocks and is associated with coding parameters, for example, for prediction purposes. An example of possible subdivision patterns is shown in Figure 8. For all samples in a block (or corresponding block in a sample array), the decoding of the associated coding parameters is specified in a specific way. In the example, all samples in a block are predicted using the same set of prediction parameters, such as a reference index (identifying a reference image in a set of already encoded images), a motion parameter (identifying a measure of the block's movement between the reference image and the current image), parameters for identifying an interpolation filter, and an intra-prediction mode.Motion parameters can be represented by displacement vectors having horizontal and vertical components, or by higher-order motion parameters such as affine motion parameters consisting of six components. Multiple sets of specific prediction parameters (e.g., reference index and motion parameters) can also be associated with a single block. In this case, for each of these specific sets of prediction parameters, one intermediate prediction signal is generated for the block (or corresponding block in the sample array), and the final prediction signal is constructed by a combination that includes superimposing the intermediate prediction signals. Corresponding weighting parameters and, optionally, a constant offset (added to the weighted sum) can be fixed with respect to the image, or to the reference image, or to a set of reference images, or they can be included in the set of prediction parameters for the corresponding block. The difference (also called the residual signal) between the original block (or corresponding block in the sample array) and their prediction signals is typically transformed and quantized. Often, a two-dimensional transformation is applied to the residual signal (or the corresponding sample array for the residual block). For transform coding, the block (or corresponding block in the sample array) for which a specific set of prediction parameters was used can be further divided before the transformation is applied. A transformation block can be equal to or smaller than the block used for prediction. A transformation block can also contain two or more blocks used for prediction. Different transformation blocks can have different sizes, and a transformation block can represent a square or rectangular block. In the above examples relating to Figures 1-7, it was noted that the leaf node of the first subdivision, i.e., the coding block 40, can be further divided in parallel into partitions that define the granularity of the coding parameters, on the one hand, and into transformation blocks to which the two-dimensional transformation is individually applied, on the other hand. After the transformation, the resulting transformation coefficients are quantized, and so-called transformation coefficient levels are obtained. The transformation coefficient levels, as well as the prediction parameters and, if any, the subdivision information, are entropy coded. In particular, the coding parameters for the transformation blocks are called residual parameters.Similar to prediction parameters, residual parameters, and any subdivision information, can be entropy encoded. In the latest video coding standards, such as H.264, a flag called the coded block flag (CBF) can signal that all transformation coefficient levels are 0, and therefore the residual parameters are not encoded. According to the present invention, this signaling is combined with integrated activation signaling.

[0068] In modern image and video encoding standards, the possibility of subdividing an image (or plane set) into syntax-supplied blocks is very limited. Typically, only whether (and possibly how) a block of a predetermined size can be subdivided into smaller blocks is specified. For example, the maximum block size in H.264 is 16x16. A 16x16 block is also called a macroblock, and each image is divided into macroblocks in the first step. For each 16x16 macroblock, it may be signaled whether it will be encoded as a 16x16 block, or as two 16x8 blocks, or as two 8x16 blocks, or as four 8x8 blocks. If a 16x16 block is subdivided into four 8x8 blocks, each of these 8x8 blocks can be encoded as one 8x8 block, or as two 8x4 blocks, or as two 4x8 blocks, or as four 4x4 blocks. A small set of possible divisions into blocks in the latest image and video coding standards has the advantage that the auxiliary information rate for signaling the division information can be kept small, but it has the disadvantage that the bitrate required to transmit the predictive parameters for the blocks can become significant, as described below. The auxiliary information rate for transmitting the predictive information usually represents a significant amount of the total bitrate for the blocks. And when this auxiliary information is reduced, coding efficiency can be increased, which can be achieved, for example, by using a larger block size. It is also possible to increase the set of supported division patterns compared to H.264. For example, the division patterns shown in Figure 8 can be supplied for square blocks of all sizes (or selected sizes). The real image or picture of a video sequence consists of arbitrarily shaped objects having certain properties. For example, this type of object or part of an object is characterized by a unique texture or unique motion.Typically, the same set of prediction parameters can be used for this type of object or part of an object. However, object boundaries usually do not coincide with possible block boundaries for large prediction blocks (e.g., 16x16 macroblocks in H.264). The encoder typically determines a subdivision (within a limited set of possibilities) that results in the minimum rate-distortion cost measure. For an arbitrarily shaped object, this results in a large number of small blocks. This statement also holds true when more subdivision patterns than those mentioned above are supplied. It should be noted that the number of subdivision patterns should not be too large, for a lot of auxiliary information and / or encoder / decoder computation is required to transmit and process the signals of these patterns. Thus, an arbitrarily shaped object can often result in a large number of small blocks due to subdivision. And since each of these small blocks is associated with a set of prediction parameters that need to be transmitted, the auxiliary information rate can be a significant part of the overall bit rate. However, since some of the small blocks still represent a region or part of the same object, the prediction parameters for the large number of resulting blocks are the same or very similar. Intuitively, coding efficiency can be increased when the syntax is extended in a way that not only allows for the subdivision of a block, but also allows for the merging of two or more blocks obtained after subdivision. As a result, a group of blocks are obtained that are coded by the same predictive parameters. The predictive parameters for this kind of group of blocks need to be coded only once. In the above example in Figures 1 to 7, for example, if merging occurs, the coding parameters for the current block 40 are not transmitted. That is, the encoder does not transmit the coding parameters associated with the current block, and the decoder does not expect the bitstream 30 to contain the coding parameters for the current block 40. Rather, according to its particular embodiment, only refinement information may be transmitted for the merged current block 40.The determination, reduction, and consolidation of candidate sets are performed for the other coding blocks 40 of the image 20. The coding blocks somehow form groups of coding blocks along the coding chain. There, the coding parameters for these groups are sent only once in the bitstream.

[0069] If the bitrate saved by reducing the number of coding prediction parameters is greater than the additional bitrate spent on coding the integrated information, the described integration results in increased coding efficiency. It should be further noted that the described syntax extension (for integration) provides the encoder with an additional degree of freedom in selecting the division of an image or plane group into blocks without introducing redundancy. The encoder is not limited to first performing the subdivision and then checking whether some of the resulting blocks have the same set of prediction parameters. In one simple variation, the encoder could first decide on the subdivision as a modern coding technique. Then, it could check on a block-by-block basis whether the integration with one of its adjacent blocks (or associated already determined groups of blocks) reduces the rate-distortion cost measure. In this, the prediction parameters associated with the new group of blocks can be re-estimated (e.g., by performing a new motion search), or the prediction parameters already determined for the current block and adjacent blocks or groups of blocks can be evaluated for the new group of blocks. The encoder could also directly check the patterns (or subsets thereof) supplied by combinations of division and integration. In other words, motion estimation and mode determination can be performed in the resulting shape, as already mentioned above. The integrated information can be signaled on a block basis. In practice, this integration can also be interpreted as the result obtained from estimating the predictive parameters for the current block, where the estimated predictive parameters are set to be equal to the predictive parameters of one of the adjacent blocks.

[0070] Regarding modes other than skip, an additional flag such as CBF is required to signal that the residual signal is not transmitted. There are two variations of the skip / direct mode in the modern H.264 technical video coding standard, temporal direct mode and spatial direct mode, which are selected at the image level. Both direct modes are only applicable to B-pictures. In temporal direct mode, the reference index for reference image list 0 is set to equal 0, and the reference index for reference image list 1, as well as the motion vectors for both reference lists, are obtained based on the motion data of a macroblock located at the same position in the first reference image of reference image list 1. Temporal direct mode uses motion vectors from temporally collocated blocks and scales the motion vectors according to the time distance between the current block and its collocated block. In spatial direct mode, the reference index and motion vectors for both reference image lists are essentially estimated based on motion data in spatial adjacency. The reference index is selected as the minimum value of the corresponding reference index in the spatial adjacency, and each motion vector component is set to equal the median value of the corresponding motion vector components in the spatial adjacency. Skip mode can only be used to encode 16x16 macroblocks of H.264 (P and B pictures), while direct mode can be used to encode 16x16 macroblocks or 8x8 sub-macroblocks. In contrast to direct mode, when merge is applied to the current block, all prediction parameters can be duplicated from the block into which the current block is merged. Merging can also be applied to any block size, resulting in the more flexible partitioning pattern described above, where all samples of a single pattern are predicted using the same prediction parameters.

[0071] The underlying concept of the embodiments outlined above or below is to reduce the bitrate required to transmit the CBF flag by combining the integration and the CBF flag. If the sample set uses integration and no residual data is transmitted, one flag is transmitted signaling both.

[0072] To reduce the rate of auxiliary information in image and video coding applications, a particular set of samples (which may represent rectangular or square blocks, or regions of arbitrary shape, or other sets of samples) is typically associated with a particular set of coding parameters. For each of these sample sets, the coding parameters are included in the bitstream. The coding parameters may indicate prediction parameters, which specify how the corresponding set of samples is predicted using already coded samples. The division of an image's sample array into sample sets may be fixed by syntax or signaled by corresponding subdivision information in the bitstream. Multiple subdivision patterns may be allowed for a single block. The coding parameters for a sample set are sent in a default order given by syntax. With respect to the current set of samples, it may be signaled that it will be merged into one or more other sample sets (for example, for prediction purposes) in a group of sample sets. A possible set of values ​​for the corresponding merger information may be applied to the adopted subdivision pattern in a way that a particular subdivision pattern cannot be shown by the combination of other subdivision patterns and the corresponding merger data. The coding parameters for a group of sample sets need to be transmitted only once. In addition to prediction parameters, residual parameters (e.g., transformation aid information, quantization aid information, and transformation coefficient levels) may be transmitted. If the current sample sets are to be merged, aid information indicating the merge process is transmitted. This aid information is further referred to as merge information. The embodiments described above and below illustrate a concept in which signaling of merge information is combined with signaling of coded block flags (which identify whether residual data exists for a block).

[0073] In a special embodiment, the integration information includes a combined, so-called mrg_cbf flag, which is equal to 1 if the current sample set is integrated and residual data is not transmitted. In this case, no further coding parameters and residual parameters are transmitted. If the combined mrg_cbf flag is equal to 0, another flag is encoded indicating whether integration is applied. Furthermore, a flag is encoded indicating that residual parameters are not transmitted. In CABAC and context-adaptive VLC, the context for deriving probabilities (and VLC table switching) for syntax elements associated with the integration information can be selected as a function of already transmitted syntax elements and / or decoding parameters (e.g., the combined mrg_cbf flag).

[0074] In a preferred embodiment, the integrated information, including the combined mrg_cbf flag, is encoded before the encoded parameters (e.g., prediction information and segmentation information).

[0075] In a preferred embodiment, the integrated information containing the composite mrg_cbf flag is encoded after a subset of encoding parameters (e.g., prediction information and segmentation information). The integrated information can be encoded for all sample sets resulting from the segmentation information.

[0076] In embodiments described further below with respect to Figures 11 to 13, mrg_cbf is called skip_flag. Typically, mrg_cbf can be called merge_skip to indicate that it is another version of skip associated with block merge.

[0077] The following preferred embodiments are described with respect to a set of samples representing rectangular and square blocks, but can be directly extended to regions or other sets of samples of any shape. The preferred embodiments show a combination of syntax elements associated with the integration scheme and syntax elements indicating the absence of residual data. The residual data may include residual auxiliary information and transformation coefficient levels. Again, with respect to the preferred embodiments, the absence of residual data is identified by an encoded block flag (CBF), which can similarly be represented by other means or flags. A CBF equal to 0 relates to the case where no residual data is transmitted.

[0078] 1. Combination of Integration Flag and CBF Flag Below, the flag that activates auxiliary mergers is called mrg, but in relation to Figures 11-13 later, it will be called merge_flag. Similarly, the merged index is currently called mrg_idx, but later merge_idx will be used.

[0079] The possible combinations of integration flags and CBF flags using a single syntax element are described in this section. The descriptions of these possible combinations outlined below are shown in Figures 1 to 6 and can be converted to any of the above descriptions.

[0080] In a preferred embodiment, up to three syntax elements are transmitted to identify the integration information and the CBF.

[0081] The first syntax element (hereinafter referred to as mrg_cbf) identifies whether the current set of samples is to be merged with other sample sets and whether all corresponding CBFs are equal to 0. The mrg_cbf syntax element can only be encoded if the extracted set of candidate sample sets is not empty (after possible elimination of candidates resulting in splits that may be signaled by different splitting patterns without merge). However, it can be guaranteed by default that the list of merge candidates never disappears, i.e., there are at least one or at least two available merge candidates. In a preferred embodiment of the present invention, if the extracted set of candidate sample sets is not empty, the mrg_cbf syntax element is encoded as follows:

[0082] - If the blocks are currently merged and the CBF is equal to 0 for all components (e.g., luminance and the two chroma components), the mrg_cbf syntax element is set to 1 and encoded. - Otherwise, the mrg_cbf syntax element is set to equal to 0 and encoded.

[0083] The values ​​0 and 1 for the mrg_cbf syntax element can also be toggled.

[0084] A second syntax element, separately called mrg, determines whether the current set of samples is to be merged with another set of samples. If the mrg_cbf syntax element is equal to 1, the mrg syntax element is not encoded and is instead assumed to be equal to 1. If the mrg_cbf syntax element does not exist (because the extracted set of candidate samples is empty), the mrg syntax element also does not exist and is assumed to be equal to 0. However, it can be guaranteed by default that the list of merge candidates never disappears, i.e., there are at least one or at least two available merge candidates.

[0085] A third syntax element, also called mrg_idx, is encoded only if the mrg syntax element is equal to (or presumed to be equal to) 1, and specifies which of the candidate sample sets is adopted for integration. In a preferred embodiment, if the extracted set of candidate sample sets contains one or more candidate sample sets, the mrg_idx syntax element is encoded only. In another preferred embodiment, if at least two sample sets of the extracted set of candidate sample sets are associated with different encoding parameters, the mrg_idx syntax element is encoded only.

[0086] It should be noted that the unified candidate list can even be fixed to separate parsing and reconstruction, improving parsing processing capabilities and making it more robust with respect to information loss. More precisely, this separation can be ensured by using a fixed assignment of list entries and codewords. This does not require fixing the length of the list. However, simultaneously fixing the length of the list by adding additional candidates allows for compensation of the coding efficiency loss of the fixed (longer) codewords. Thus, as previously stated, if the list of candidates contains one or more candidates, the unified index syntax element can only be transmitted. However, this would require extracting the list before parsing the unified index, and would prevent performing these two operations simultaneously. To enable increased parsing throughput and make the parsing process more robust with respect to transmission errors, it is possible to eliminate this dependency by using a fixed codeword for each index value and a fixed number of candidates. If this number cannot be reached by candidate selection, it is possible to extract auxiliary candidates to complete the list. These additional candidates may include so-called combined candidates, constructed from the motion parameters of possibly different candidates already in the list, and zero motion vectors.

[0087] In a preferred embodiment, after a subset of prediction parameters (or, more generally, specific encoded parameters associated with a sample set) is transmitted, integrated information for the sample set is encoded. The subset of prediction parameters may consist of one or more reference image indices, or one or more components of a motion parameter vector or reference image index, and one or more components such as a motion parameter vector.

[0088] In a preferred embodiment, the mrg_cbf syntax element of the integrated information is encoded only with respect to a reduced set of partitioning modes. The possible sets of partitioning modes are shown in Figure 8. In a preferred embodiment, this reduced set of partitioning modes is limited to one, corresponding to the first partitioning mode (top left of the list in Figure 8). For example, mrg_cbf is encoded only if the block is not further partitioned. As a further example, mrg_cfb can be encoded only for square blocks.

[0089] In another preferred embodiment, the mrg_cbf syntax element of the integration information is encoded only for one block of the partition, where this partition is one of the possible partition modes shown in Figure 8, for example, a partition mode using the four blocks in the lower left. In a preferred embodiment, if there are one or more blocks to be combined in one of these partition modes, the integration information for the first integrated block (in decoding order) includes the mrg_cbf syntax element for the entire partition. For all other blocks of the same partition mode that are subsequently decoded, the integration information includes only the mrg syntax element that identifies whether the current set of samples is to be integrated with other sample sets. Information on whether residual data exists is inferred from the mrg_cbf syntax element encoded in the first block.

[0090] In a further preferred embodiment of the present invention, integration information for a set of samples is encoded before the prediction parameters (or, more generally, specific coding parameters associated with the set of samples). The integration information, including the mrg_cbf, mrg, and mrg_idx syntax elements, is encoded in the manner shown in the first preferred embodiment above. If the integration information signals that the current set of samples will not be integrated with any other set of samples, and that the CBF is equal to 1 for at least one of the components, then the prediction or coding parameters and residual parameters are simply transmitted. In a preferred embodiment, if the mrg_cbf syntax element indicates that the current block is integrated and the CBF for all components is equal to zero, then no further signaling is required after the integration information for this current block.

[0091] In other preferred embodiments of the present invention, the syntax elements mrg_cbf, mrg, and mrg_idx are combined and encoded as one or two syntax elements. In a preferred embodiment, mrg_cbf and mrg are combined into a single syntax element which identifies one of the following cases: (a) the block is merged and does not contain residual data, (b) the block is merged and contains (or may contain) residual data, or (c) the block is not merged. In another preferred embodiment, the syntax elements mrg and mrg_idx are combined into a single syntax element which, if N is the number of merge candidates, identifies one of the following cases: the block is not merged, the block is merged into candidate 1, the block is merged into candidate 2, ..., the block is merged into candidate N. In a further preferred embodiment of the present invention, the syntax elements mrg_cfb, mrg, and mrg_idx are combined into a single syntax element, which (using N, the number of candidates) identifies one of the following cases: the block is not combined, the block is combined into candidate 1 and does not contain residual data, the block is combined into candidate 2 and does not contain residual data, ..., the block is combined into candidate N and does not contain residual data, the block is combined into candidate 1 and contains (or may contain) residual data, the block is combined into candidate 2 and contains (or may contain) residual data, ..., the block is combined into candidate N and contains (or may contain) residual data. The combined syntax element can be transmitted by variable-length coding, or by arithmetic coding, or by binary arithmetic coding using any particular binarization scheme.

[0092] 2. Combinations of Integration Flag and CBF Flag, and Skip / Direct Mode Skip / direct modes may be supported for all or only a particular block size and / or block shape. In extensions of skip / direct modes, such as those specified in the modern technical video coding standard H.264, a set of candidate blocks is used for skip / direct modes. The difference between skip and direct is whether residual parameters are sent. Parameters for skip and direct (e.g., predictions) can be estimated from one of the corresponding candidates. A candidate index that signals which candidate is used to estimate the coding parameters is coded. When multiple predictions are combined to form a final prediction signal for a current block (such as the bidirectional predictive blocks used in H.264B-frames), each prediction may relate to a different candidate. Thus, for each prediction, a candidate index can be coded.

[0093] In a preferred embodiment of the present invention, the candidate list for skip / direct may include different candidate blocks than the candidate list for integrated mode. One example is shown in Figure 10. The candidate list may include the following blocks (the current block is indicated by Xi): ● Motion vector (0,0) ● Median (between Left, Above, and Corner points) ●Left block (Li) ● Above block (Ai) ●Corner blocks (in order: Above Right (Ci1), Below Left (Ci2), Above Left (Ci3)) ●Collocated blocks of images that are different but have already been encoded.

[0094] The following symbols are used to describe the embodiments described below. ●set_mvp_ori is a set of candidates used for skip / direct mode. This set consists of {Median, Left, Above, Corner, Collocated}. The Median is the median of the ordered set of Left, Above, and Corner, and the Collocated is given by the nearest reference frame (or the first reference image of one of the reference image lists), and the corresponding motion vector is scaled according to the time distance. For example, if there are no Left, Above, or Corner blocks, a Motion Vector with both components equal to 0 may be additionally inserted into the list of candidates. ●set_mvp_comb is a subset of set_mvp_ori.

[0095] In a preferred embodiment, both skip / direct mode and block integration mode are supported. Skip / direct mode uses the original set of candidates, set_mvp_ori. Integration information associated with block integration mode may include combined mrg_cbf syntax elements.

[0096] In other embodiments, skip / direct mode and block merge mode are supported, but skip / direct mode uses a modified set of candidates set_mvp_comb. This modified set of candidates may be a specific subset of the original set set_mvp_ori. In a preferred embodiment, the modified set of candidates consists of corner blocks and collocated blocks. In other embodiments, the modified set of candidates consists only of collocated blocks. Further subsets are possible.

[0097] In another embodiment, integrated information containing the mrg_cbf syntax element is encoded before the parameter associated with the skip mode.

[0098] In another embodiment, the parameter associated with the skip mode is encoded before the integrated information containing the mrg_cbf syntax element.

[0099] According to other embodiments, the direct mode may not be activated (non-existent), and the block integration has an extended set of candidates having a skip mode which is replaced with mrg_cbf.

[0100] In a preferred embodiment, the candidate list for block integration may include different candidate blocks. One example is shown in Figure 10. The candidate list may include the following blocks (the current block is indicated by Xi): ● Motion Vector (0,0) ●Left block (Li) ● Above block (Ai) ●Collocated blocks of images that are different but have already been encoded. ●Corner blocks (in order: Above Right (Ci1), Below Left (Ci2), Above Left (Ci3)) ● Combined bidirectional predictive candidates ● Unscaled bidirectional predictive candidates

[0101] It must be stated that the candidate locations for block consolidation can be the same as the MVP list for interpretation in order to avoid memory access.

[0102] Furthermore, lists can be “fixed” in the manner outlined above to separate parsing and reconstruction in order to improve parsing processing capabilities and make them more robust with respect to information loss.

[0103] 3. Encoding of CBF In a preferred embodiment, if the mrg_cfb syntax element is equal to 0 (which signals that the block will not be merged, or that it contains non-zero residual data), a flag is sent that signals whether all components of the residual data (e.g., luminance and the two chroma components) are zero. If mrg_cfb is equal to 1, this flag is not sent. In a particular configuration, if mrg_cfb is equal to 0, this flag is not sent, and the syntax element mrg indicates that the block will be merged.

[0104] In another preferred embodiment, if the mrg_cfb syntax element is equal to 0 (which signals that the block will not merge or that it contains non-zero residual data), several syntax elements are sent for each component, signaling whether the residual data for that component is zero.

[0105] Different context models can be used for mrg_cbf.

[0106] Thus, the above embodiment is particularly, A subdivider configured to subdivide the sample set of samples, A merger configured to integrate sample sets into one or more independent sets of sample sets, An encoder configured to encode an image using encoding parameters that vary across the entire image in units of independent sets of sample sets, wherein the encoder is configured to encode an image by predicting the image and encoding the predicted residuals for a given set of samples, The present invention describes an image encoding apparatus that includes a stream generator configured to insert predicted residuals and coding parameters into a bitstream, in addition to one or more syntax elements for each of at least a subset of sample sets, which signal whether each sample set is to be combined with other sample sets into one of a set of independent sets.

[0107] Furthermore, a device for decoding a bitstream containing an encoded image is described, which is... A subdivider configured to subdivide the image into a sample set, A merger configured to integrate each sample set into one or more independent sets of sample sets, A decoder configured to decode an image using coding parameters that vary across the entire image in units of independent sets of sample sets, wherein the decoder is configured to decode an image by predicting an image with respect to a default set of sample sets, decoding the prediction residuals for the default set of sample sets, and combining the predictions resulting from predicting the image with respect to the prediction residuals, An extractor configured to extract prediction residuals and coding parameters from a bitstream, along with one or more syntax elements for each of at least subsets of the sample sets, which signal whether each of the sample sets will be merged with other sample sets into one of independent sets, wherein the merger is configured to perform the merger in response to the syntax elements.

[0108] One of the possible states of one or more syntax elements is that each sample set is merged with another sample set into one independent set, signaling that the bitstream does not have encoded and inserted prediction residuals.

[0109] The extractor may also be configured to extract subdivision information from the bitstream, and the subdivider may be configured to subdivide the image into a sample set in response to the subdivision information.

[0110] The extractor and merger, for example, move sequentially through the sample set according to the sample set scanning order, and with respect to the current sample set, Extract the first binary syntax element (mrg_cbf) from the bitstream, If the first binary syntax element takes the first binary state, the current sample set is merged into one of the independent sets by estimating that the coding parameters for the current sample set are equal to the coding parameters associated with this independent set, skipping the extraction of predicted residuals for the current sample set, and proceeding to the next sample set in the sample set scan order. If the first binary syntax element takes on the second binary state, the second syntax element (mrg, mrg_idx) is extracted from the bitstream. Depending on the second syntax element, the system may be configured to extract at least one further syntax element relating to the predicted residuals for the current sample set, and to integrate the current sample set by estimating that the coding parameters for the current sample set are equal to the coding parameters associated with one of the independent sets, or to perform the extraction of coding parameters for the current sample set.

[0111] At least one syntax element for each subset of the sample set can also signal which of the set of default candidate sample sets adjacent to each sample set will be used to merge each sample set into one of the independent sets, if each sample set is to be merged with another sample set.

[0112] If one or more syntax elements do not signal that each sample set will be merged into one of two independent sets along with another sample set, the extractor will not... The bitstream may be configured to extract one or more additional syntax elements (skip / direct mode) that signal whether at least some of the encoding parameters for each sample set will be predicted, or whether they will be predicted from any of the further sets of predefined candidate sample sets adjacent to each sample set.

[0113] In that case, the default candidate sample set and any further default candidate sample set are independent of or may intersect with respect to a small number of default candidate sample sets within each of the default candidate sample set and any further default candidate sample set.

[0114] The extractor can also be configured to extract subdivision information from the bitstream, and the subdivider is configured to hierarchically subdivide the image into sample sets in response to the subdivision information, and the extractor moves sequentially through the child sample sets of the parent sample set, which consist of the sample sets into which the image is subdivided, and with respect to the current child sample set, The first binary syntax element (mrg_cbf) is extracted from the bitstream, If the first binary syntax element takes on the first binary state, the current child sample set is integrated into one of the independent sets by estimating that the coding parameter for the current child sample set is equal to the coding parameter associated with this independent set, skipping the extraction of the predicted residual for the current child sample set and proceeding to the next child sample set. If the first binary syntax element takes on the second binary state, the second syntax element (mrg, mrg_idx) is extracted from the bitstream, and then, Depending on the second syntax element, extract at least one further syntax element relating to the predicted residual for the current child sample set, proceed to the next child sample set, and integrate the current child sample set into one of the independent sets by estimating that the coding parameters for the current child sample set are equal to the coding parameters associated with this independent set, or perform the extraction of coding parameters for the current child sample set. For the next child sample set, if the first binary syntax element of the current child sample set takes a first binary state, the extraction of the first binary syntax element is skipped and the second syntax element is extracted instead, and so on. If the first binary syntax element of the current child sample set takes a second binary state, the first binary syntax element is extracted.

[0115] For example, let's assume that a parent sample set (CU) is split into two child sample sets (PU). For the first PU, if the first binary syntax element (merge_cbf) has a first binary state, then 1) the first PU uses merge, and the first and second PUs (the entire CU) have no residual data in the bitstream, and 2) for the second PU, the second binary syntax elements (merge_flag, merge_idx) are signaled. However, if the first binary syntax element for the first PU has a second binary state, then 1) for the first PU, the second binary syntax elements (merge_flag, merge_idx) are signaled, and similarly, residual data is present in the bitstream, while 2) for the second PU, the first binary syntax element (merge_cbf) is signaled. Thus, if merge_cbf is in a second binary state with respect to all previous child sample sets, merge_cbf may signal at the PU level, i.e., with respect to consecutive child samples. If merge_cbf is in a first binary state for consecutive child sample sets, it does not mean that all child sample sets following this child sample set have residual data in the bitstream. For example, with respect to a CU divided into four PUs, merge_cbf may be in a first binary state with respect to the second PU, meaning that, for example, the third and fourth PUs in encoding order do not have residual data in the bitstream, but the first PU does or could have it.

[0116] The first and second binary syntax elements can be encoded using context-adaptive variable-length coding or context-adaptive (binary) arithmetic coding, and the context for encoding the syntax elements is obtained based on the values ​​for these syntax elements in an already encoded block.

[0117] As described in other preferred embodiments, the syntax element merge_idx may only be sent if the list of candidates contains one or more candidates. This requires obtaining the list before parsing the merged index, preventing these two processes from being performed in parallel. To enable increased parsing throughput and make the parsing process more robust with respect to transmission errors, this dependency can be eliminated by using a fixed codeword and a fixed number of candidates per index value. If this number is not reached by candidate selection, it is possible to obtain auxiliary candidates to complete the list. These additional candidates may include so-called combined candidates, constructed from the motion parameters and zero motion vectors of possibly different candidates already in the list.

[0118] In another preferred embodiment, the syntax for signaling any of the blocks in the candidate set is applied simultaneously in the encoder and decoder. For example, given three choices of blocks for integration, those three choices are considered for entropy coding solely by their syntax. The probability of all other choices is assumed to be 0, and the entropy codec is fitted simultaneously in the encoder and decoder.

[0119] The predictive parameters estimated as a result of the integration process may represent the entire set of predictive parameters associated with the block, or they may represent a subset of these predictive parameters (e.g., predictive parameters for one hypothesis of the block in which multi-hypotheses prediction is used).

[0120] In a preferred embodiment, syntax elements associated with the integrated information are entropy-encoded using context modeling.

[0121] One way to translate the embodiments outlined above to a specific syntax is described below with respect to the following figures. In particular, Figures 11 to 13 show different parts of the syntax that utilize the embodiments outlined above. Specifically, according to the embodiments outlined below, image 20 is first divided into coding tree blocks whose image content is encoded using the syntax coding_tree shown in Figure 11. As shown there, for example, for entropy_coding_mode_flag=1 for context-adaptive binary arithmetic coding or other specific entropy coding modes, the quad-tree subdivision of the current coding tree block is signaled within the syntax part coding_tree by a flag called split_coding_unit_flag in code 400. As shown in Figure 11, according to the embodiments described below, the tree root block is subdivided so as to be signaled by split_coding_unit_flag in depth-first scan order as shown in Figure 7a. Whenever a leaf node is reached, it indicates a coding unit that is immediately encoded using the syntax function coding_unit. This can be seen from Figure 11 when we look at the conditional clause 402, which checks whether the current split_coding_unit_flag is set. If YES, the function coding_tree is called recursively, resulting in further transmission / extraction of split_coding_unit_flag in the encoder and decoder, respectively. Otherwise, i.e., split_coding_unit_flag=0, the current subblock of the tree root block 200 in Figure 7a is a leaf block, and the function coding_unit in Figure 10 is called at 404 to encode this coding unit.

[0122] In the embodiments described above, the above options are used to determine which integration is only available for images where the inter-prediction mode is available. That is, intra-encoded slices / images do not use integration in any case. This can be seen from Figure 12, and the flag skip_flag is sent in 406 only in the case of a slice type different from the intra-image slice type, i.e., the current slice to which the current encoding unit belongs allows the partition to be inter-encoded. According to this embodiment, integration is only about the prediction parameters associated with inter-prediction. According to this embodiment, skip_flag is signaled for all encoding units 40, and if skip_flag is equal to 1, this flag value is 1) The partitioning mode for the current coding unit is non-partitioning mode, where it is not partitioned and only the partitions of that coding unit represent itself. 2) The current coding unit / partition is to be intercoded, i.e., assigned to an intercoding mode. 3) The current encoding unit / partition will be affected by the merger. 4) The current encoding unit / partition is subordinate to skip mode, i.e., it signals the decoder to activate skip mode in parallel.

[0123] Therefore, if skip_flag is set, the function prediction_unit is called at 408, indicating that the current coding unit is the prediction unit. However, this is not the only possibility of switching the integration option. Rather, if skip_flag associated with all coding units is not set at 406, the prediction type of the coding unit of the non-intra image slice is signaled at 410 by the syntax element pred_type, and accordingly, the function prediction_unit is called for any partition of the current coding unit, for example at 412, if the current coding unit is not further partitioned. In Figure 12, only four different partitioning options are shown to be available, but other partitioning options shown in Figure 8 may also be available. Another possibility is that the partitioning option PART_NxN is not available, but the others are. The relationship between the partitioning options shown in Figure 8 and the names of the partitioning modes used in Figure 12 is shown in Figure 8 by the respective subscripts below the individual partitioning options. Note that the prediction type syntax element pred_type signals not only the prediction mode, i.e., intra-encoded or inter-encoded, but also the partition in the case of inter-encoded mode. The case of inter-encoded mode is described further. The function prediction_unit is called for each partition, such as partitions 50 and 60 in the aforementioned encoding order. The function prediction_unit begins by checking skip_flag at 414. If skip_flag is set, merge_idx is necessarily followed at 416. The check in step 414 is about checking whether a skip_flag associated with the entire encoding unit, as signaled at 406, has been set. If not, merge_flag is signaled again at 418, and if the latter is set, merge_idx, indicating the merge candidate for the current partition, is followed at 420.Furthermore, merge_flag is signaled at 418 for the current partition only if the current prediction mode of the current coding unit is inter-prediction mode (see 422). That is, if skip_flag is not set, the prediction mode is signaled at 410 via pred_type. Here, for each prediction unit, if pred_type signals that the inter-coding mode is active (see 422), the merge-specific flag, i.e., merge_flag, is sent individually for each subsequent partition when the merge is activated for each partition by the merge index merge_idx.

[0124] As can be seen from Figure 13, according to this embodiment, the transmission of prediction parameters used for the 424 current prediction units is performed only when the merge is not used for the current prediction units, i.e., the merge is not activated by either the skip_flag or the merge_flag for each partition.

[0125] As already shown above, skip_flag=1 signals that the residual data will not be sent. This is derived from the fact that the transmission of residual data in Figure 12, 426 for the current coding unit only occurs if skip_flag is equal to 0, as can be seen from the residual data transmission in the other options of conditional clause 428, which checks the state of skip_flag immediately after the transmission.

[0126] Up to this point, the embodiments in Figures 11-13 have been described only under the assumption that entropy_coding_mode_flag is equal to 1. However, the embodiments in Figures 11-13 include one embodiment of the embodiments outlined above in the case of entropy_coding_mode_flag=0, where other entropy coding modes are used to entropy code syntax elements, such as variable-length coding, or more precisely, context-adaptive variable-length coding. In particular, the possibility of simultaneously signaling the activation of integration and skip modes follows the modification outlined above, where the common signaling state is only one of two or more states for each syntax element. This will be described in more detail here. However, it is emphasized that the possibility of switching between both entropy coding modes is arbitrary, and therefore another embodiment can be easily obtained from Figures 11-13 by simply enabling one of the two entropy coding modes.

[0127] For example, see Figure 11. If entropy_coding_mode_flag is equal to 0 and the slice_type syntax element signals that the current tree root block belongs to an inter-coding slice, i.e., if an inter-coding mode is available, the syntax element cu_split_pred_part_mode is transmitted at 430, and this syntax element, as indicated by its name, signals information about further subdivision of the current coding unit, activation or deactivation of skip mode, and activation or deactivation of merge and predict mode, along with each splitting information. See Table 1.

[0128] [Table 1]

[0129] Table 1 identifies the meaning of the possible states of the syntax element cu_split_pred_part_mode when the current coding unit has a size that is not the smallest in the quad-tree subdivision within the current tree root block. The possible states are listed in the leftmost column of Table 1. Since Table 1 relates to the case where the current coding unit does not have the smallest size, there is a state for cu_split_pred_part_mode, i.e., state 0. And it signals that the current coding unit needs to be subdivided into four further coding units, not the actual coding unit, and that they will be traversed in depth-first traversal order as outlined by calling the function coding_tree 432 further. That is, cu_split_pred_part_mode=0 signals that the current quad-tree subdivision unit of the current tree root block will be subdivided into four even smaller units, i.e., split_coding_unit_flag=1. However, if cu_split_pred_part_mode takes any other possible state, split_coding_unit_flag=0, and the current unit forms a leaf block of the current tree root block, i.e., a coding unit. In this case, one of the remaining possible states of cu_split_pred_part_mode is the common signaling state described above, which simultaneously signals that the current coding unit is under the influence of merge and that the skip mode is activated, as indicated by skip_flag being equal to 1 in the third column of Table 1. At the same time, it signals that no further partitions of the current coding unit will occur, i.e., PART_2Nx2N is selected as the splitting mode. There is also a possible state of cu_split_pred_part_mode where the skip mode is not activated and it signals the activation of merge. This is possible state 2, which corresponds to the non-splitting mode being active, i.e., PART_2Nx2N, where skip_flag=0 and merge_flag=1.In other words, in that case, merge_flag is signaled beforehand, rather than within the prediction_unit syntax. For the remaining possible states of cu_split_pred_part_mode, the interpretation modes for other split modes are signaled for these split modes that divide the current coding unit into one or more partitions.

[0130] [Table 2]

[0131] Table 2 shows the significance or semantics of the possible states of cu_split_pred_part_mode when the current coding unit has the smallest possible size, according to the quad-tree subdivision of the current tree root block. In this case, all possible states of cu_split_pred_part_mode do not correspond to further subdivision by split_coding_unit_flag=0. However, possible state 0 signals that its skip_flag=1, i.e., it signals that integration is activated and skip mode is activated simultaneously. Furthermore, it signals that non-splitting occurs, i.e., split mode PART_2Nx2N. Possible state 1 corresponds to possible state 2 in Table 1, which corresponds to possible state 2 in Table 2, which corresponds to possible state 3 in Table 1.

[0132] The above description of the embodiments shown in Figures 11 to 13 has already explained most of the functions and meanings, but some further information is shown below.

[0133] A skip_flag[x0][y0] equal to 1 indicates that, for the current coding unit (see Figure 40), no syntax elements other than the motion vector predictor index (merge_idx) will be parsed after skip_flag[x0][y0] when decoding a P or B slice. A skip_flag[x0][y0] equal to 0 indicates that the coding unit will not be skipped. The array index x0,y0 identifies the position (x0,y0) of the top-left luminance sample of the coding unit considered, relative to the top-left luminance sample of the image (see Figure 20).

[0134] When skip_flag[x0][y0] is not present, it is inferred to be equal to 0.

[0135] As described above, if skip_flag[x0][y0] is equal to 1, - PredMode is presumed to be equal to MODE_SKIP - PartMode is estimated to be equal to PART_2Nx2N

[0136] cu_split_pred_part_mode[x0][y0] identifies the split_coding_unit_flag, and when the coding unit is not split, according to skip_flag[x0][y0], merge_flag[x0][y0], PredMode, and the PartMode of the coding unit. Array indices x0 and y0 identify the position (x0,y0) of the top-left luminance sample of the coding unit in relation to the top-left luminance sample of the image.

[0137] merge_flag[x0][y0] identifies whether the inter-prediction parameters for the current prediction unit (50 and 60 in the figure, i.e., partitions within the coded unit 40) are estimated from adjacent inter-predicted partitions. The array index x0,y0 identifies the position (x0,y0) of the top-left luminance sample of the considered prediction block in relation to the top-left luminance sample of the image.

[0138] merge_idx[x0][y0] identifies the merge candidate index of the merge candidate list. Here, x0,y0 identifies the position (x0,y0) of the top-left luminance sample of the considered prediction block in relation to the top-left luminance sample of the image.

[0139] Although not specifically shown in the above description of Figures 11-13, in this embodiment, the merge candidates or list of merge candidates are not only determined, for example, using only the coding or prediction parameters of spatially adjacent prediction units / partitions, but rather, the list of candidates is also formed by using the prediction parameters of temporally adjacent partitions of previously coded images. Furthermore, combinations of prediction parameters of spatially and / or temporally adjacent prediction units / partitions are used and included in the list of merge candidates. Of course, only a subset thereof may be used. In particular, Figure 14 shows one possibility for determining spatial adjacency, i.e., spatially adjacent partitions or prediction units. Figure 14 shows, as an example, a prediction unit or partition 60, and pixels B0-B2 and A0-A1, positioned directly adjacent to the boundary 500 of partition 60. That is, B2 is diagonally adjacent to the top-left pixel of partition 60, B1 is vertically adjacent to the top-right pixel of partition 60, B0 is diagonally positioned to the top-right pixel of partition 60, A1 is positioned to the left of the bottom-left pixel of partition 60, and A0 is diagonally positioned to the bottom-left pixel of partition 60. Partitions containing pixels B0-B2 and at least one of A0 and A1 form spatial adjacencies, and their prediction parameters form a combined candidate.

[0140] The following function may be used to perform the aforementioned removal of those candidates that would lead to other available split modes.

[0141] In particular, the coding / prediction parameters arising from a candidate N, i.e., pixel N=(B0,B1,B2,A0,A1), i.e., a prediction unit / partition covering the position (xN,yN), are excluded from the candidate list if any of the following conditions are true (see Figure 8 for the partition mode PartMode and the corresponding partition index PartIdx which shows each partition within the coding unit).

[0142] - The current prediction unit's PartMode is PART_2NxN, PartIdx is equal to 1, and the prediction units covering the luminance positions (xN, yN-1) (PartIdx=0) and (xP, yP-1) (Cand.N) have the same motion parameters. mvLX[xP,yP-1]==mvLX[xN,yN] refIdxLX[xP,yP-1]==refIdxLX[xN,yN] predFlagLX[xP,yP-1]==predFlagLX[xN,yN]

[0143] - The current prediction unit's PartMode is PART_Nx2N, PartIdx is equal to 1, and the prediction units covering the luminance position (xP-1, yP) (PartIdx=0) and the luminance position (xN, yN) (Cand.N) have the same motion parameters. mvLX[xP-1,yP]==mvLX[xN,yN] refIdxLX[xP-1,yP]==refIdxLX[xN,yN] predFlagLX[xP-1,yP]==predFlagLX[xN,yN]

[0144] - The current prediction unit's PartMode is PART_NxN, PartIdx is equal to 3, and the prediction units covering the luminance positions (xP-1,yP)(PartIdx=2) and (xP-1,yP-1)(PartIdx=0) have the same motion parameters. mvLX[xP-1,yP]==mvLX[xP-1,yP-1] refIdxLX[xP-1,yP]==refIdxLX[xP-1,yP-1] predFlagLX[xP-1,yP]==predFlagLX[xP-1,yP-1] Furthermore, the prediction units covering the luminance position (xP, yP-1) (PartIdx=1) and the luminance position (xN, yN) (Cand.N) have the same motion parameters. mvLX[xP,yP-1]==mvLX[xN,yN] refIdxLX[xP,yP-1]==refIdxLX[xN,yN] predFlagLX[xP,yP-1]==predFlagLX[xN,yN]

[0145] - The current prediction unit's PartMode is PART_NxN, PartIdx is equal to 3, and the prediction units covering the luminance positions (xP,yP-1)(PartIdx=1) and (xP-1,yP-1)(PartIdx=0) have the same motion parameters. mvLX[xP,yP-1]==mvLX[xP-1,yP-1] refIdxLX[xP,yP-1]==refIdxLX[xP-1,yP-1] predFlagLX[xP,yP-1]==predFlagLX[xP-1,yP-1] Furthermore, the prediction units covering the luminance positions (xP-1, yP) (PartIdx=2) and (xN, yN) (Cand.N) have the same motion parameters. mvLX[xP-1,yP]==mvLX[xN,yN] refIdxLX[xP-1,yP]==refIdxLX[xN,yN]

[0146] In this regard, note that the position or location (xP, yP) indicates the largest pixel of the current partition / prediction unit. That is, all coding parameter candidates obtained by directly adopting each coding parameter of adjacent prediction units, i.e., prediction unit N, according to the first item, are checked. However, other additional coding parameter candidates can be checked in the same way as to whether they are equal to the coding parameters of each occurring prediction unit, which would result in other partition patterns supported by the syntax. According to the embodiment just described, the identity of coding parameters includes checking the identity of the parameters, i.e., the motion vector, i.e., mvLX, the reference index, i.e., refIxLX, and the prediction flag predFlagLX, which indicates that the parameters, i.e., the motion vector and reference index, are used in interprediction, associated with the motion vector, i.e., mvLX, the reference index, i.e., refIxLX, and the reference list X, where X is 0 or 1.

[0147] It should be noted that the aforementioned possibility of removing candidate coding parameters for adjacent prediction units / partitions also applies when supporting the asymmetric partitioning modes shown in the right half of Figure 8. In that case, mode PART_2NxN may represent a mode where all partitions are subdivided horizontally, and PART_Nx2N may correspond to a mode where all partitions are subdivided vertically. Furthermore, mode PART_NxN can be excluded from the supported partitioning modes or patterns, in which case only the first two exclusion checks must be performed.

[0148] Regarding the embodiments shown in Figures 11 to 14, it should be noted that intra-predicted partitions may be excluded from the list of candidates, meaning that their encoding parameters may not be included in the list of candidates.

[0149] Furthermore, note that three contexts can be used for skip_flag, merge_flag, and merge_idx, respectively.

[0150] While several embodiments have been described in relation to the apparatus, it is clear that these embodiments also provide descriptions of the corresponding methods. Here, a block or device corresponds to a method step or a function of a method step. Similarly, embodiments described in relation to a method step also provide descriptions of the corresponding blocks, items, or features of the corresponding apparatus. Some or all of the method steps can be performed by (or using) hardware devices such as, for example, a microprocessor, a programmable computer, or an electronic circuit. In some embodiments, one or more of the most important method steps can be performed by such devices.

[0151] Depending on the specific implementation requirements, embodiments of the present invention may be implemented in hardware or in software. The embodiments may be implemented using a digital storage medium, such as a floppy disk, DVD, Blu-ray, CD, ROM, PROM, EPROM, EEPROM, or FLASH memory, having electronically readable control signals stored thereon, which cooperates (or can cooperate) with a programmable computer system, so that each method can be performed. Thus, the digital storage medium may be computer-readable.

[0152] Some embodiments of the present invention include a data carrier having electronically readable control signals that can cooperate with a programmable computer system so that one of the methods described in the present specification is performed.

[0153] Typically, embodiments of the present invention can be executed as a computer program product having program code, such that when the computer program product is running on a computer, the program code executes one of the methods. The program code can be stored, for example, in a machine-readable carrier.

[0154] Other embodiments include a computer program stored in a machine-readable carrier for performing one of the methods described herein.

[0155] Therefore, in other words, an embodiment of the method of the present invention is a computer program having program code for executing one of the methods described in the present specification when the computer program is running on a computer.

[0156] Accordingly, a further embodiment of the method of the present invention is a data carrier (or digital storage medium or computer-readable medium) on which a computer program is recorded for performing one of the methods described in the present specification. The data carrier, digital storage medium or recorded medium is generally tangible and / or non-transient.

[0157] Accordingly, a further embodiment of the method of the present invention is a data stream or sequence of signals representing a computer program for performing one of the methods described in the present specification. The data stream or sequence of signals may be configured to be transmitted, for example, over a data communication connection, such as the Internet.

[0158] Further embodiments include processing means, such as a computer or a programmable logic device, configured or adapted to perform one of the methods described in this specification.

[0159] Further embodiments include a computer on which a computer program for performing one of the methods described in this specification is installed.

[0160] Further embodiments of the present invention include an apparatus or system configured to transfer (e.g., electronically or optically) a computer program for performing one of the methods described in the present specification to a receiver. The receiver may be, for example, a computer, a mobile device, a storage device, etc. The apparatus or system may include, for example, a file server for transferring the computer program to the receiver.

[0161] In some embodiments, a programmable logic device (e.g., a field-programmable gate array) can be used to perform some or all of the functions of the methods described herein. In some embodiments, a field-programmable gate array can cooperate with a microprocessor to perform one of the methods described herein. Typically, the methods are preferably performed by any hardware device.

[0162] The embodiments described above are merely illustrative for the principle of the present invention. Modifications and changes to the apparatus and details described in this specification will be obvious to those skilled in the art. Therefore, it is intended that the invention is limited only by the immediate claims and not by the specific details shown in the description and explanation of the embodiments in this specification.

Claims

1. An apparatus for decoding a bitstream in which an image (20) is encoded, A subdivider (82) configured to subdivide the aforementioned image (20) into a sample set of sample (40), A merger (84) configured to merge the aforementioned sample set (40) into a group of one or more sample sets, A decoder (86) configured to decode the image (20) using encoding parameters transmitted in the bitstream in units of groups of the sample set, wherein the decoder (86) is configured to predict the image (20) with respect to a default sample set, decode the prediction residuals for the default sample set, and decode the image by combining the prediction residuals with the predictions resulting from predicting the image (20), An extractor (88) configured to extract from the bitstream (30) the prediction residuals and coding parameters, and one or more syntax elements for at least each subset of the sample set (40), indicating whether each of the sample set (40) will be merged with another sample set (40) into one of the groups, wherein the merger (84) is configured to perform a merge in response to the one or more syntax elements, and Equipped with, One of the possible states of the one or more syntax elements indicates that each of the sample sets (40) is merged with another sample set (40) into one of the groups, and that each of the sample sets (40) has zero predicted residuals. The extractor is also configured to extract subdivision information from the bitstream using entropy decoding, and the subdivider is configured to subdivide the image into a sample set in response to the subdivision information. An apparatus characterized by the following features.

2. The extractor and the merger are configured to sequentially advance through the sample sets according to the sample set scanning order, and with respect to the current sample set, The first binary syntax element is extracted from the bitstream, If the first binary syntax element takes on a first binary state, the current sample set is merged into the group by estimating that the coding parameter for the current sample set is equal to the coding parameter associated with one of the groups, the extraction of the predicted residual for the current sample set is skipped, and the process proceeds to the next sample set in the sample set scan order. If the first binary syntax element takes on a second binary state, the second syntax element is extracted from the bitstream, and Depending on the second syntax element, the system is configured to merge the current sample set into a group by estimating that the coding parameter of the current sample set is equal to the coding parameter associated with one of the groups, or to perform the extraction of the coding parameter of the current sample set and extract at least one further syntax element relating to the predicted residual for the current sample set. The apparatus according to claim 1, characterized in that

3. The apparatus according to claim 1, wherein at least one syntax element for each subset of the sample set signals which of the set of predefined candidate sample sets adjacent to each sample set will be merged with if each sample set is merged with another sample set into any one of the groups.

4. The extractor is also configured to extract subdivision information from the bitstream, the subdivider is configured to hierarchically subdivide the image into sample sets in response to the subdivision information, the extractor is configured to sequentially proceed through the child sample sets of the parent sample set that the sample set into which the image is subdivided contains, and with respect to the current child sample set, The first binary syntax element is extracted from the bitstream, If the first binary syntax element takes on a first binary state, the current child sample set is merged into this group by estimating that the coding parameter of the current child sample set is equal to the coding parameter associated with one of the groups, the extraction of the predicted residual for the current child sample set is skipped, and the process moves on to the next child sample set. If the first binary syntax element takes on a second binary state, the second syntax element is extracted from the bitstream. Depending on the second syntax element, the current child sample set is merged into this group by estimating that the coding parameter of the current child sample set is equal to the coding parameter associated with one of the groups, or the current child sample set is extracted along with at least one further syntax element relating to the predicted residual for the current child sample set, With respect to the next child sample set, if the first binary syntax element of the current child sample set takes the first binary state, the extraction of the first binary syntax element is skipped and the extraction of the second syntax element is performed instead; if the first binary syntax element of the current child sample set takes the second binary state, the first binary syntax element is extracted. The apparatus according to claim 1, characterized in that

5. The apparatus according to claim 4, characterized in that the first binary syntax element and the second binary syntax element are encoded using context-adaptive variable-length coding or context-adaptive binary arithmetic coding, and the context for encoding the syntax elements is derived based on the values ​​for these syntax elements in an already encoded block.

6. The apparatus according to claim 1, wherein the bitstream further comprises a depth map to which the image is associated.

7. The apparatus according to claim 1, characterized in that the sample array is one of a set of sample arrays relating to different faces of the image, which are encoded independently of each other.

8. A device for encoding images, A subdivider (72) configured to subdivide the aforementioned image into a sample set of samples, A merger (74) configured to merge each of the aforementioned sample sets into a group of one or more sample sets, An encoder (76) configured to encode the image using encoding parameters transmitted in a bitstream in units of groups of the sample set, wherein the encoder (76) is configured to encode the image by predicting the image and encoding the predicted residual for a predetermined sample set, A stream generator (78) configured to insert the prediction residuals and coding parameters, as well as one or more syntax elements for at least each subset of the sample set, which indicate whether each sample set is merged with another sample set into one of the groups, into the bitstream. Includes, One of the possible states of the one or more syntax elements indicates that each sample set is merged with another sample set into one of the groups, and that each sample set has zero predictive residuals. The stream generator is also configured to insert subdivision information into the bitstream using entropy coding, and the subdivision information indicates how to subdivide the image into the sample set. An apparatus characterized by the following features.

9. The apparatus according to claim 8, wherein the bitstream further comprises a depth map to which the image is associated.

10. The apparatus according to claim 8, characterized in that the array of information samples is one of a set of sample arrays relating to different faces of the image, which are encoded independently of each other.

11. A method for decoding a bitstream in which an image (20) is encoded, wherein the method is: The steps include subdividing the aforementioned image (20) into a sample set (40) of samples, The steps include merging each of the aforementioned sample sets (40) into a group of one or more sample sets, A step of decoding the image (20) using encoding parameters transmitted in the bitstream in units of groups of the sample sets, wherein the decoder (86) is configured to decode the image by predicting the image (20) with respect to a default sample set, decoding the prediction residuals for the default sample set, and combining the prediction residuals with the prediction resulting from predicting the image (20), Steps include extracting from the bitstream (30) the predicted residuals and the coding parameters, as well as one or more syntax elements for at least each subset of the sample set (40), which indicate whether each sample set (40) is merged with another sample set (40) into one of the groups, wherein the merger (84) is configured to perform the merge in response to the one or more syntax elements. Includes, One of the possible states of the one or more syntax elements signals that each sample set (40) is merged with another sample set (40) into one of the groups, and that each sample set (40) has zero predicted residuals. Subdivision information is extracted from the bitstream using entropy decoding, and the image is subdivided into the sample set in response to the subdivision information. A method characterized by the following features.

12. The method according to claim 11, wherein the bitstream further comprises a depth map to which the image is associated.

13. The method according to claim 11, characterized in that the array of information samples is one of a set of sample arrays relating to different faces of the image, which are encoded independently of each other.

14. A method for encoding an image, The steps include subdividing the image into a sample set of samples, The steps include merging each of the aforementioned sample sets into a group of one or more sample sets, A step of encoding the image using encoding parameters transmitted in a bitstream in units of groups of the sample set, wherein the encoder (76) is configured to encode the image by predicting the image and encoding the predicted residuals for a predetermined sample set. The steps include inserting the predicted residuals and the coding parameters, as well as one or more syntax elements for at least each subset of the sample set, which indicate whether each sample set is merged with another sample set into one of the groups, into a bitstream; Includes, One of the possible states of the one or more syntax elements indicates that each sample set is merged with another sample set into one of the groups, and that each sample set has zero predictive residuals. The subdivision information is inserted into the bitstream using entropy coding, and the subdivision information indicates how the image is subdivided into the sample set. A method characterized by the following features.

15. The method according to claim 14, wherein the bitstream further comprises a depth map to which the image is associated.

16. The method according to claim 14, characterized in that the array of information samples is one of a set of sample arrays relating to different faces of the image, which are encoded independently of each other.

17. Subdividing an image into a sample set of samples, Merging each of the aforementioned sample sets into a group of one or more sample sets, The image is encoded using encoding parameters transmitted in the bitstream in units of groups of the sample set, wherein the encoder (76) is configured to encode the image by predicting the image and encoding the predicted residual for a predetermined sample set. Inserting the prediction residuals and coding parameters, as well as one or more syntax elements for at least each subset of the sample set, into the bitstream, each syntax element indicating whether the respective sample set is merged with another sample set into one of the groups. A digital storage medium storing the bitstream in which the image is encoded, One of the possible states of the one or more syntax elements indicates that each sample set is merged with another sample set into one of the groups, and that each sample set has zero predictive residuals. The subdivision information is inserted into the bitstream using entropy coding, and the subdivision information indicates how the image is subdivided into the sample set. A digital storage medium characterized by the following features.

18. A digital storage medium storing a bitstream in which an image is encoded, wherein the bitstream is Predicted residuals and coding parameters, wherein the coding parameters are transmitted in the bitstream in units of groups of sample sets of the image, which are obtained by subdividing the image into sample sets and merging the sample sets into groups of sample sets, and the predicted residuals and coding parameters, At least one syntax element for each subset of the sample set, indicating whether each of the sample sets is merged with another sample set into one of the groups, and Includes, One of the possible states of the one or more syntax elements indicates that each sample set is merged with another sample set into one of the groups, and that each sample set has zero predictive residuals. The subdivision information is inserted into the bitstream using entropy coding, and the subdivision information indicates how the image is subdivided into the sample set. A digital storage medium characterized by the following features.

19. The digital storage medium according to claim 18, further comprising a depth map to which the image is associated.

20. The digital storage medium according to claim 18 or 19, characterized in that the array of information samples is one of a set of sample arrays relating to different faces of the image, which are encoded independently of each other.

21. A method for decoding a bitstream, Subdividing an image into a sample set of samples, Merging each of the aforementioned sample sets into one or more groups of sample sets, The image is encoded using encoding parameters transmitted in the bitstream in units of groups of the sample set, wherein the encoder (76) is configured to encode the image by predicting the image and encoding the predicted residual for a predetermined sample set. Inserting the prediction residuals and coding parameters, as well as one or more syntax elements for at least each subset of the sample set, into the bitstream, which indicate whether each sample set is merged with another sample set into one of the groups. The steps include receiving and decoding a bitstream encoded by One of the possible states of the one or more syntax elements indicates that each sample set is merged with another sample set into one of the groups, and that each sample set has zero predictive residuals. The subdivision information is inserted into the bitstream using entropy coding, and the subdivision information indicates how the image is subdivided into the sample set. A method characterized by the following features.

22. A method for decoding a bitstream, Predicted residuals and coding parameters, wherein the coding parameters are transmitted in the bitstream in units of groups of sample sets of an image, which are obtained by subdividing the image into sample sets and merging the sample sets into groups of sample sets, and the predicted residuals and coding parameters, At least one syntax element for each subset of the sample set, one or more syntax elements indicating whether each sample set is merged with another sample set into one of the groups, The step includes receiving a bitstream containing and decoding it, One of the possible states of the one or more syntax elements indicates that each sample set is merged with another sample set into one of the groups, and that each sample set has zero predictive residuals. The subdivision information is inserted into the bitstream using entropy coding, and the subdivision information indicates how the image is subdivided into the sample set. A method characterized by the following features.

23. A method for storing an image, Predicted residuals and coding parameters, wherein the coding parameters are transmitted in a bitstream in units of groups of sample sets resulting from subdividing the image into sample sets and merging the sample sets into groups of sample sets, and the predicted residuals and coding parameters, At least one syntax element for each subset of the sample set, one or more syntax elements indicating whether each sample set is merged with another sample set into one of the groups, and The step includes storing a bitstream containing the bitstream into a digital storage medium, One of the possible states of the one or more syntax elements indicates that each sample set is merged with another sample set into one of the groups, and that each sample set has zero predictive residuals. The subdivision information is inserted into the bitstream using entropy coding, and the subdivision information indicates how the image is subdivided into the sample set. A method characterized by the following features.

24. A method for transmitting an image, Predicted residuals and coding parameters, wherein the coding parameters are transmitted in a bitstream in units of groups of sample sets resulting from subdividing the image into sample sets and merging the sample sets into groups of sample sets, and the predicted residuals and coding parameters, At least one syntax element for each subset of the sample set, one or more syntax elements indicating whether each sample set is merged with another sample set into one of the groups, and The step includes transmitting a bitstream containing the bitstream via transmission, One of the possible states of the one or more syntax elements indicates that each sample set is merged with another sample set into one of the groups, and that each sample set has zero predictive residuals. The subdivision information is inserted into the bitstream using entropy coding, and the subdivision information indicates how the image is subdivided into the sample set. A method characterized by the following features.

25. The method according to any one of claims 21 to 24, wherein the bitstream further comprises a depth map to which the image is associated.

26. The method according to any one of claims 21 to 24, characterized in that the array of information samples is one of a set of sample arrays relating to different faces of the image, which are coded independently of each other.

Citation Information

Patent Citations

  • Method and apparatus for context dependent merging for skip-direct modes for video encoding and decoding

    CN101682769A

  • Separating and merging device, method for coding signal and computer program product

    CN1339922A

  • Method and apparatus for supporting the syntax structure of multipass video for slice data

    JP2010529811A

  • Image processing apparatus, image processing circuit, and image processing method

    US20080279463A1

  • Method and an apparatus for processing a video signal

    US20100086051A1