Picture coding supporting block partitioning and block merging

By removing redundant coding parameters in block integration, the solution improves encoding efficiency in image and video codecs, reducing auxiliary information and enhancing segmentation freedom.

JP2025111518AActive Publication Date: 2025-07-30DOLBY VIDEO COMPRESSION LLC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2025064843
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2010-10-08
Filing Date
2025-04-10
Publication Date
2025-07-30
Estimated Expiration
2031-10-10

AI Technical Summary

Technical Problem

Existing image and video codecs face inefficiencies due to redundancy in block merging and segmentation, leading to increased auxiliary information requirements and limited freedom in segmenting images.

Method used

The solution involves avoiding redundancy by removing coding parameter candidates that are identical across integrated blocks, thereby reducing signaling overhead and maintaining efficient block integration and splitting.

Benefits of technology

This approach enhances encoding efficiency by minimizing auxiliary information and increasing the number of achievable split patterns, while maintaining the benefits of block integration, thus optimizing the trade-off between auxiliary information and segmentation freedom.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025111518000001_ABST
    Figure 2025111518000001_ABST
Patent Text Reader

Abstract

To provide a method of avoiding redundancy in block partitioning and block merging of picture and / or video coding.SOLUTION: In an encoder configured to encode a picture 20 into a bitstream 30, if the signaled one of supported partitioning patterns specifics a subdivision of a block 40 into two or more blocks 50, 60, a removal of certain coding parameter candidates for all further blocks, except the first further block of the further blocks in a coding order, is performed. In particular, those coding parameter candidates are removed from a set of coding parameter candidates for the respective further blocks, the coding parameters which are the same as coding parameters associated with any of the further blocks which, when being merged with the respective further blocks, would result in one of the supported partitioning pattern. By this measure, redundancy between partitioning coding and merging coding can be avoided.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to image and / or video coding, and more particularly, to a codec that supports block splitting and block integration.

Background Art

[0002] Many image and / or video codecs process images in blocks. For example, predictive codecs use block granularity to find a good compromise between setting the prediction parameter set very accurately at high spatial resolution, which would consume too much auxiliary information for the prediction parameters, and setting the prediction parameters very coarsely, which would lead to an increase in the amount of bits required to encode the prediction residuals due to the low spatial resolution of the prediction parameters. In short, the optimal setting for the prediction parameters lies somewhere between the two extremes.

[0003] To obtain an optimal solution to the above problems, several attempts have been made. For example, instead of using a regular subdivision of the image into blocks regularly arranged in rows and columns, a subdivision that performs a multi-tree split tries to increase the freedom to subdivide the image into blocks with appropriate requirements for the subdivision information. However, even multi-tree subdivision requires a significant amount of data signaling, and the freedom to subdivide the image is quite limited even when using such multi-tree subdivision.

[0004] To enable a better trade-off between the amount of auxiliary information required to signal image segmentation and the degree of freedom in segmenting an image, block merging can be used to increase the number of possible image segmentations with a reasonable amount of additional data required to signal the merging information. For the merged blocks, the encoding parameters likewise need to be transmitted only once in total in the bitstream, as if the resulting merged group of blocks were a directly segmented part of the image.

[0005] However, there is still redundancy newly created by the combination of block merging and block segmentation, so there is a need to achieve better encoding efficiency.

SUMMARY OF THE INVENTION

PROBLEMS TO BE SOLVED BY THE INVENTION

[0006] Thus, an object of the present invention is to provide an encoding concept with further encoding efficiency. This object is achieved by the independent claims according to the present application.

MEANS FOR SOLVING THE PROBLEMS

[0007] The underlying idea of the present invention is that, for the current block of an image that signals one of the supported split patterns for a bitstream, if the reversal of the split by block integration is avoided, a further increase in coding efficiency can be achieved. In particular, when one of the signaled supported split patterns defines the subdivision of a block into two or more further blocks, for all further blocks other than the first further block among the further blocks in the coding order, the removal of specific coding parameter candidates is performed. In particular, coding parameters that are the same as the coding parameters associated with any of the further blocks that would result in one of the supported split patterns when integrated with each further block are removed from the set of coding parameter candidates for each further block. By this method, redundancy between split coding and integration coding is avoided, and the signaling overhead for signaling integration information can be reduced by using a set of coding parameter candidates of a reduced size. Furthermore, the positive effect of combining block splitting with block integration is maintained. That is, the number of achievable split patterns for combining block splitting with block integration increases compared to the case without block integration. The increase in signaling overhead is kept within a reasonable range. Finally, block integration enables combining further blocks across the boundary with the current block, thereby providing a granularity that would not be possible without block integration.

[0008] Applying a slightly different idea of the set of integration candidates, according to a further aspect of the present invention, in a decoder configured to decode a bitstream that signals one of the supported segmentation patterns for the current block of an image where the decoder is configured to be removed, when one of the signaled supported segmentation patterns defines a subdivision into two or more further blocks of the block, for all further blocks except the first further block of the further blocks in the encoding order, from the set of candidate blocks of each further block, identify a candidate block that becomes one of the supported segmentation patterns when integrated with each further block.

[0009] Advantageous embodiments of the present invention are the subject of the appended dependent claims.

[0010] Preferred embodiments of the present application will be described in more detail below with reference to the figures.

Brief Description of the Drawings

[0011]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5A

Figure 5B

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11A

Figure 11B

Figure 12

[0012] Regarding the following description, whenever the same reference numerals are used in relation to different figures, the description associated with each element presented in relation to one of these figures is also applicable to the other figure, provided that the description does not conflict with the remaining description of the other figure. Please note this point.

[0013] FIG. 1 shows an encoder 10 according to an embodiment of the present invention. The encoder 10 is configured to encode an image 20 into a bitstream 30. Of course, the image 20 can be part of a video if the encoder is a video encoder.

[0014] The image 20 includes a block 40 that is to be encoded by the encoder 10 currently. As shown in FIG. 1, the image 20 may include two or more blocks 40. For example, the image 20 may be subdivided into a regular arrangement of blocks 40 such that, as shown in FIG. 1 by way of example, the blocks 40 are arranged in rows and columns. However, any other subdivision of the image 20 into blocks 40 may also be possible. In particular, the subdivision of the image 20 into blocks 40 may be fixed, i.e., may be known to the decoder by default, or may be signaled to the decoder within the bitstream 30. In particular, the blocks 40 of the image 20 may have different sizes. For example, in this case, a multi-tree subdivision such as a quadtree subdivision may be applied to the image 20 so as to obtain blocks 40 that form the leaf blocks of the multi-tree subdivision, or the image 20 may be regularly pre-subdivided into regularly arranged tree root blocks.

[0015] In any case, the encoder 10 is configured to signal in the bitstream 30 one of the supported partitioning patterns for the current block 40. That is, the encoder 10 determines, for example, which of the supported partitioning patterns should be used for the current block 40 in order to, for example, further partition the partition block 40 if it is beneficial from the perspective of rate distortion optimization and to adapt the granularity at which specific encoding parameters are set within the current block 40 of the image 20. As will be outlined in more detail below, the encoding parameters can indicate prediction parameters such as, for example, inter prediction parameters. This kind of inter prediction parameter can include, for example, a reference image index, a motion vector, and the like. The supported partitioning patterns can include, for example, a non-partitioning mode, i.e., an option where the current block 40 is not further subdivided, a horizontally partitioning mode, i.e., an option where the current block 40 is subdivided into an upper or top and a lower or bottom part along a horizontally extended line, and a vertically partitioning mode, i.e., an option where the current block 40 is vertically subdivided into a left part and a right part along a vertically extended line. In addition to this, the supported partitioning patterns can also include an option where the current block 40 is further regularly subdivided into four further blocks each considered to be a quarter of the current block 40. Furthermore, the partitioning can relate to all blocks 40 of the image 20 or only to its appropriate subset, such as those having a specific encoding mode associated with it, such as the inter prediction mode. Additionally, the set of blocks for which integration is to be applied to the partitioning of the blocks can be further limited by bitstream signaling for each block 40 for which integration can be performed as to whether integration is available for the partitioning of the blocks. Naturally, this kind of signaling can also be done individually for each possible integration candidate partition.Furthermore, various subsets of the supported partitioning modes can be available for block 40, either in combination or individually, depending on the level of subdivision of block 40, e.g., the block size, if it is a multi-tree subdivision leaf block.

[0016] That is, while the subdivision of image 20 into blocks to obtain block 40 in particular can be fixed or signaled in the bitstream, the partitioning pattern that will currently be used for block 40 is signaled in bitstream 30 in the form of partitioning information. Thus, the partitioning information can be considered, in this way, as an extension of a kind of subdivision of image 20 into block 40. On the other hand, a further relationship of the original granularity of the subdivision of image 20 into block 40 can still be maintained. For example, encoder 10 can be configured to signal, in bitstream 30, each part of image 20 or the encoding mode used for block 40 at the granularity defined by block 40, while encoder 10 is configured to vary the encoding parameters of each encoding mode within each block 40 at an increased (finer) granularity defined by each partitioning pattern selected for each block 40. For example, the encoding modes signaled at the granularity of block 40 can distinguish between intra prediction modes such as time-inter prediction mode, inter-view prediction mode, etc., and inter prediction modes. The type of encoding parameters associated with one or more sub-blocks (partitions) resulting from the partitioning of each block 40 depends on the encoding mode assigned to each block 40. For example, for an intra-encoded block 40, the encoding parameters can include the spatial direction as to which image content of the previously decoded part of image 20 is used to fill each block 40. In the case of an inter-encoded block 40, the encoding parameters can in particular include motion vectors for motion compensation prediction.

[0017] FIG. 1 shows, by way of example, the current block 40 as being subdivided into two further (smaller) blocks 50 and 60. In particular, a vertical split mode is shown by way of example. The smaller blocks 50 and 60 can also be referred to as sub-blocks 50 and 60, or partitions 50 and 60, or prediction units 50 and 60. In particular, when one of the signaled supported split patterns defines the subdivision of the current block 40 into two or more further blocks 50 and 60, the encoder 10, for all further blocks other than the first further block of the further blocks 50 and 60 in encoding order, excludes from the set of encoding parameter candidates for each further block, encoding parameter candidates having the same encoding parameters as the encoding parameters associated with any of the further blocks that would be one of the supported split patterns when integrated into each further block. More precisely, for each of the supported split patterns, the encoding order is defined among the resulting one or more partitions 50 and 60. In the case of FIG. 1, the encoding order is shown, by way of example, by an arrow 70 that defines that the left partition 50 is encoded before the right partition 60. In the case of a horizontal split mode, it can be defined that the upper partition is encoded before the lower partition. In any case, the encoder 10, for the second partition 60 in the encoding order 70, excludes from the set of encoding parameter candidates for each second partition 60, encoding parameter candidates having the same encoding parameters as the encoding parameters associated with the first partition 50, in order to avoid that, as a result of this integration, both partitions 50 and 60 would actually equally result in selecting the non-split mode for the current block 40 with a low encoding rate.

[0018] More precisely, the encoder 10 is configured to use block integration in an efficient manner along with block splitting. As far as block integration is concerned, the encoder 10 determines each set of coding parameter candidates for each of the partitions 50 and 60. The encoder can be configured to determine a set of coding parameter candidates for each of the partitions 50 and 60 based on the coding parameters associated with the previously decoded blocks. In particular, at least some of the coding parameter candidates in the set of coding parameter candidates may be equal to, i.e., adopted from, the coding parameters of the previously decoded partitions. Additionally, or alternatively, at least some of the coding parameter candidates can be obtained from coding parameter candidates associated with two or more previously coded partitions by an appropriate combination such as a median value, an average, etc. However, if the encoder 10 determines a reduced set of coding parameter candidates and one or more such coding parameter candidates remain after removal, and depending on one non-removed or selected coding parameter candidate to set the coding parameters associated with each partition, the encoder 10 is configured to perform a selection of one of the remaining non-removed coding parameter candidates for each non-first partition 60. Thus, the encoder 10 is configured to perform the removal such that coding parameter candidates that would lead to an efficient recombination of the partition 50 and the partition 60 are removed. That is, in this way, a collection of syntaxes in which an efficient splitting situation is coded more complexly than simply signaling this split using only the split information is effectively avoided.

[0019] Furthermore, as the set of encoding parameter candidates becomes smaller, the amount of auxiliary information required to encode the integration information into the bitstream 30 can be reduced due to the smaller number of elements in these candidate sets. In particular, since the decoder can determine and then reduce the set of encoding parameter candidates in the same way as the encoder in FIG. 1 does, the encoder 10 in FIG. 1 can, for example, use fewer bits to insert syntax elements into the bitstream 30 and utilize the reduced set of encoding parameter candidates, defining which of the remaining encoding parameter candidates that were not removed will be used for integration. Naturally, the introduction of syntax elements into the bitstream 30 can be completely suppressed if the number of remaining encoding parameter candidates for each partition is simply 1. In any case, through integration, that is, by setting the encoding parameters associated with each partition depending on the remaining one, or the selected one, of the encoding parameter candidates that were not removed, the encoder 10 can suppress the complete re-insertion of the encoding parameters for each split into the bitstream 30, thereby also reducing the auxiliary information. According to some embodiments of the present application, the encoder 10 may be configured to signal, in the bitstream 30, refinement information for refining the remaining one, or the selected one, of the encoding parameter candidates for each partition.

[0020] As described above, according to the description of FIG. 1, the encoder 10 is configured to determine integration candidates to be excluded by comparing those encoding parameters with the encoding parameters of the partition, and the integration thereby results in another supported split pattern. For example, if the encoding parameters of the left partition 50 form one element of a set of encoding parameter candidates for the right partition 60, this method of processing the encoding parameter candidates efficiently excludes at least one encoding parameter candidate in the case shown in FIG. 1. However, additional encoding parameter candidates can also be excluded if they are equal to the encoding parameters of the left partition 50. However, according to other embodiments of the present invention, the encoder 10 removes the candidate block or blocks from this set of candidate blocks that would result in one of the supported split patterns when integrated into each partition, so as to determine a set of candidate blocks for each partition after the second in the encoding order. In a sense, this means the following. The encoder 10 is such that each element of the candidate set is associated with it in that the candidate adopts each encoding parameter of its associated partition, and has exactly one partition of either the current block 40 or the previously encoded block 40. It can be configured to determine integration candidates for each of the partition 50 or the partition 60 (i.e., the first and the next in the encoding order). For example, each element of the candidate set can be equal to one of this type of encoding parameter of the previously encoded partition, i.e., can be adopted from among them, or at least can be obtained from the encoding parameters of just one such previously encoded partition by additional scaling or refinement using additional transmitted refinement information, etc.However, the encoder 10 may also be configured to add further elements or candidates to such a candidate set, i.e., obtained from combinations of the encoding parameters of one or more previously encoded partitions, or by modification taking only the encoding parameters of one motion parameter list, etc., to add encoding parameter candidates obtained from the encoding parameters of one previously encoded partition. For the "combined" elements, there is no 1:1 relationship between the encoding parameters of each candidate element and each partition. According to the first variant of the description of FIG. 1, the encoder 10 can be configured to exclude all candidates from the entire candidate set, the encoding parameters of which are equal to the encoding parameters of partition 50. According to the latter variant of the description of FIG. 1, the encoder 10 can be configured to exclude only the elements of the candidate set associated with partition 50. Bringing both views into agreement, the encoder 10 can be configured to exclude candidates from a portion of the candidate set that shows a 1:1 relationship with some (e.g., adjacent) previously encoded partitions without extending the removal (and search for candidates with equal encoding parameters) to the remaining part of the candidate set having the encoding parameters obtained by combination. However, of course, if one combination leads to a redundant representation, this can be solved by removing the redundant encoding parameters from the list or by performing a redundancy check on the combined candidates.

[0021] After describing the encoder according to an embodiment of the present invention, with reference to FIG. 2, the decoder 80 according to an embodiment will be described. The decoder 80 in FIG. 2 is configured to decode a bitstream 30 that signals one of the supported split patterns for the current block 40 of the image 20 as described above. When the signaled supported split pattern defines the subdivision of the current block 40 into two or more partitions 50 and 60, the decoder 80 is configured to exclude from the set of encoding parameter candidates each partition encoding parameter candidate having an encoding parameter that is the same as or equal to the encoding parameter associated with any of the partitions, for all partitions other than the first partition 50 among the partitions in the encoding order 70, that is, for the partition 60 in the example shown in FIGS. 1 and 2. This, when integrated with each partition, becomes one of the supported split patterns, that is, one of the supported split patterns that is not signaled in the bitstream 30 but is nevertheless supported.

[0022] That is, the decoder function largely coincides with that of the encoder described with respect to FIG. 1. For example, the decoder 80 can be configured to set the encoding parameters associated with each partition 60 depending on one of the encoding parameter candidates that were not removed when the number of encoding parameter candidates that were not removed is not zero. For example, the decoder 80 sets the encoding parameters of each partition 60 to be equal to one of the encoding parameter candidates that were not removed, regardless of the presence or absence of additional refinement and / or regardless of the presence or absence of scaling according to the temporal distance to which each encoding parameter is related. For example, the encoding parameter candidate for integration among the candidates that were not removed may have a reference picture index related to another one different from the reference picture index explicitly signaled in the bitstream 30 for the partition 60. In that case, the encoding parameters of the encoding parameter candidate can define motion vectors related to the reference picture index respectively, and the decoder 80 can be configured to scale the motion vectors of the finally selected encoding parameter candidate that was not removed according to the ratio between both reference picture indexes. Thus, according to the above-described modification example, the encoding parameters according to integration include motion parameters, while the reference picture index is separated therefrom. However, as described above, according to another embodiment, the reference picture index can also be part of the encoding parameters according to integration.

[0023] The fact that the integration behavior can be limited to the inter-predicted block 40 also equally applies to the encoder of FIG. 1 and the decoder of FIG. 2. Therefore, the decoder 80 and the encoder 10 can be configured to support intra and inter prediction modes for the current block 40 and perform candidate integration and removal only when the current block 40 is encoded in the inter prediction mode. Therefore, only the encoding / prediction parameters of the partitions encoded before this kind of inter prediction can be used to determine / construct the candidate list.

[0024] As already described above, the encoding parameter may be a prediction parameter, and the decoder 80 can be configured to use the prediction parameters of partition 50 and partition 60 to obtain a prediction signal for each partition. Of course, the encoder 10 also performs the derivation of the prediction signal in the same way. However, in addition, the encoder 10 sets the prediction parameter in the bitstream 30 together with all other syntax elements in order to obtain some optimizations in the sense of appropriate optimization.

[0025] Furthermore, as already explained above, the encoder can be configured to insert an index into the non-excluded encoding parameter candidates for each partition only if the number of non-excluded encoding parameter candidates for each partition is greater than 1. Therefore, when the number of non-excluded encoding parameter candidates is greater than 1, for example, depending on the number of non-excluded encoding parameter candidates for partition 60, the decoder 80 can be configured to simply expect the bitstream 30 to include a syntax element that defines which of the non-excluded encoding parameter candidates is used for integration. However, if the candidate set is less than 2 in total, as described above, by expanding the candidate list / set using the combined encoding parameter, that is, the parameter obtained by combining the encoding parameters of one or more or two or more previously encoded partitions, the performance of reducing the candidate set can be generally prevented from occurring by restricting it to the candidate obtained by adopting or derived from the encoding parameter of exactly one previously encoded partition. Similarly, the reverse is also possible, that is, generally, usually, excluding all encoding parameter candidates having the same value as that of the partition which becomes the other supported split pattern.

[0026] Regarding the decision, the decoder 80 operates as the encoder 10 does. That is, the decoder 80 can be configured to determine a set of candidate encoding parameters for (one or more) partitions after the first partition 50 in the encoding order 70 based on the encoding parameters associated with the previously decoded partition. That is, the encoding order is defined not only among the partitions 50 and 60 of each block 40 but also among the blocks 40 of the image 20 itself. All the partitions encoded before the partition 60 can thus serve as a reference for determining a set of candidate encoding parameters for any of the subsequent multiple partitions, such as the partition 60 in the case of FIG. 2. As described above, the encoder and the decoder can limit the determination of the set of candidate encoding parameters to partitions that are in specific spatial and / or temporal adjacency. For example, the decoder 80 can be configured to determine a set of candidate encoding parameters for the non-first partition 60 based on the encoding parameters associated with the previously decoded partition adjacent to each non-first partition, and this kind of partition may be located outside and inside the current block 40. Naturally, the determination of the integration candidates can also be performed for the first partition in the encoding order. Only the removal is not performed.

[0027] Consistent with the description of FIG. 1, the decoder 80 can be configured to determine a set of candidate encoding parameters for each non-first partition 60 from the first set of previously decoded partitions, excluding those encoded in the intra prediction mode.

[0028] Furthermore, when the encoder has introduced segmentation information into the bitstream to segment the image 20 into the block 40, the decoder 80 can be configured to reverse the segmentation of the image 20 into this kind of encoded block 40 according to the segmentation information in the bitstream 30.

[0029] Regarding FIGS. 1 and 2, it should be noted that the residual signal for block 40 can currently be transmitted via bitstream 30 with a granularity that can be different from the granularity defined by the partition with respect to the coding parameters. For example, the encoder 10 in FIG. 1 can be configured to subdivide block 40 into one or more transform blocks in a way parallel to or independent of the division into partitions 50 and 60. The encoder can signal each transform block subdivision of block 40 with additional subdivision information. Next, the decoder 80 can be configured to reverse this further subdivision of block 40 into one or more transform blocks based on the additional subdivision information of the bitstream and derive the residual signal of the current block 40 from the bitstream in units of these transform blocks. The significance of the transform block division is that the transform such as DCT in the encoder and the corresponding inverse transform such as IDCT in the decoder are each performed within each transform block of block 40. To reconstruct the image 20 as block 40, the encoder 10 combines the prediction signal and the residual signal obtained by applying the coding parameters in each of partitions 50 and 60, for example, by adding them. However, it should be noted that the residual coding does not include the transform and the inverse transform respectively, and the prediction residual is encoded, for example, instead in the spatial domain.

[0030] Before explaining further possible details of the following further embodiments, possible internal structures of the encoders and decoders of FIGS. 1 and 2 are described with respect to FIGS. 3 and 4. FIG. 3 shows, by way of example, how encoder 10 is internally constructed. As shown in FIG. 3, encoder 10 can include subtractor 108, transformer 100, and bitstream generator 102, which can perform entropy encoding as shown in FIG. 3. Element 108, element 100, and element 102 are connected in series between input 112 that receives image 20 and output 114 that outputs the aforementioned bitstream 30. In particular, subtractor 108 has its non-inverting input connected to input 112, transformer 100 is connected between the output of subtractor 108 and the first input of bitstream generator 102, and then bitstream generator 102 has an output connected to output 114. The encoder 10 of FIG. 3 further includes an inverse transformer 104 and an adder 110, which are connected in series to the output of transformer 100 in the order described. Encoder 10 further includes a predictor 106 connected between the output of adder 110, a further input of adder 110, and the inverting input of subtractor 108.

[0031] The elements of FIG. 3 interact with each other as follows. Predictor 106 is applied to the inverted input of subtractor 108 to predict a portion of image 20 using the result of the prediction, i.e., the prediction signal. The output of subtractor 108 then indicates the difference between the prediction signal and each part of image 20, i.e., the residual signal. The residual signal follows the transform coding of transformer 100. That is, transformer 100 can perform a transform such as a DCT and subsequent quantization of the transformed residual signal, i.e., the transform coefficients, to obtain transform coefficient levels. Inverse transformer 104 reconstructs the final residual signal output by transformer 100 to obtain a reconstructed residual signal corresponding to the residual signal input to transformer 100, excluding information loss due to quantization by transformer 100. The addition of the reconstructed residual signal and the prediction signal as the output by predictor 106 results in the reconstruction of each part of image 20 and is sent from the output of adder 110 to the input of predictor 106. Predictor 106 operates in various modes such as the intra prediction mode, the inter prediction mode, etc. as described above. The prediction mode and the corresponding coding or prediction parameters applied by predictor 106 to obtain the prediction signal are sent by predictor 106 to entropy coder 102 for insertion into the bitstream.

[0032] A possible embodiment of the internal structure of decoder 80 of FIG. 2 corresponding to the possibilities shown in FIG. 3 for the encoder is shown in FIG. 4. As shown in the figure, decoder 80 can include a bitstream extractor 150, an inverse transformer 152, and an adder 154 that can be implemented as an entropy decoder as shown in FIG. 4, which are connected between the input 158 and the output 160 of the decoder in the order described. Further, the decoder of FIG. 4 includes a predictor 156 connected between the output of adder 154 and its further input. Entropy decoder 150 is connected to the parameter input of predictor 156.

[0033] To briefly explain the function of the decoder in FIG. 4, the entropy decoder 150 is for deriving all the information contained in the bit stream 30. The entropy coding scheme used may be variable length coding or arithmetic coding. Thereby, the entropy decoder 150 reverts from the bit stream conversion coefficients level indicating the residual signal and sends it to the inverse converter 152. Further, the entropy decoder 150 reverts all the coding modes and related coding parameters from the bit stream and sends it to the predictor 156. In addition, the split information and the integration information are derived from the bit stream by the derivator 150. The inversely transformed, i.e., reconstructed, residual signal and the prediction signal as obtained from the predictor 156 are combined, such as by being added by the adder 154, and then the reconstructed signal thus reverted is output at the output 160 and sent to the predictor 156.

[0034] As is apparent from comparing FIGS. 3 and 4, elements 152, element 154, and element 156 functionally correspond to elements 104, element 110, and element 106 in FIG. 3.

[0035] In the above description with respect to FIGS. 1-4, several different possibilities were presented regarding the possible subdivision of image 20 and the corresponding granularity when varying some of the parameters related to the encoded image 20. This kind of possibility is also explained with respect to FIGS. 5a and 5b. FIG. 5a shows a part of image 20. According to the embodiment of FIG. 5a, the encoder and decoder are initially configured to subdivide image 20 into tree root blocks 200. One such tree root block is shown in FIG. 5a. The subdivision of image 20 into the tree root block is made regularly in rows and columns as indicated by the dotted lines. The size of the tree root block 200 can be selected by the encoder and signaled to the decoder by the bitstream 30. Alternatively, the sizes of these tree root blocks 200 can be fixed by default. The tree root block 200 is subdivided using quadtree partitioning to result in the above-mentioned distinct blocks 40, which may be referred to as encoding blocks or encoding units. These encoding blocks or encoding units are drawn by the thin solid lines in FIG. 5a. Thereby, the encoder adds subdivision information to each tree root block 200 and inserts the subdivision information into the bitstream. This subdivision information indicates how the tree root block 200 is to be subdivided into blocks 40. At the granularity of these blocks 40, and in units thereof, the prediction mode varies within image 20. As described above, each block 40, or each block having a specific prediction mode such as an inter prediction mode, is accompanied by partition information regarding which supported partition pattern is used for each block 40. In the case shown for FIG. 5a, for many encoding blocks 40, the non-partitioning mode is selected such that the encoding block 40 coincides with the spatially corresponding partition. In other words, the encoding block 40 is, at the same time, a partition having each set of associated prediction parameters. The type of prediction parameter then depends on the mode associated with each encoding block 40. However, other encoding blocks are shown to be further subdivided, for example.The encoding block 40 at the upper right corner of the tree root block 200 is shown to be divided into, for example, four partitions, while the encoding block at the lower right corner of the tree root block 200 is illustratively shown to be vertically subdivided into two partitions. The subdivision for the division into partitions is indicated by a dotted line. Fig. 5a also shows the encoding order within the partitions thus defined. As shown in the figure, a depth-first traversal order is used. Across the tree root block boundary, the encoding order can be continued in the scanning order in which the rows of the tree root block 200 are scanned row by row from top to bottom of the image 20. By this method, it is possible to have the opportunity to maximize the possibility that a particular partition has a partition encoded before its upper and left boundaries. Each block 40, or each block having a specific prediction mode such as an inter prediction mode, can have an integration switch indicator in the bitstream indicating whether integration operates for the corresponding partition therein. It should be noted that, with the exception that this rule is made only for the smallest possible block size of the block 40, the division of the block into partitions / prediction units can be limited to a division into at most two partitions. This can avoid redundancy between the subdivision information for subdividing the image 20 into the block 40 and the division information for dividing the block 40 into partitions when using quadtree subdivision to obtain the block 40. Alternatively, only division into one or two partitions, including or not including asymmetric ones, can be allowed.

[0036] Fig. 5b shows a subdivision tree. For the solid line, the subdivision of the tree root block 200 is shown, while the dotted line represents the division of the leaf blocks of the quadtree subdivision which are the encoding blocks 40. That is, the division of the encoding blocks represents an extension of a kind of quadtree subdivision.

[0037] As already described above, each encoding block 40 can be subdivided in parallel with the transform block such that the transform block can indicate different subdivisions of each encoding block 40. For each of these transform blocks not shown in FIGS. 5a and 5b, the transform for transforming the residual signal of the encoding block can be performed separately.

[0038] Further embodiments of the present invention will be described below. While the above embodiments focused on the relationship between block integration and block splitting, the following description also includes aspects of the present application regarding other encoding principles known in current codecs such as the SKIP / DIRECT mode. Nevertheless, the following description is not to be regarded as merely explaining another embodiment, i.e., an embodiment separated from the above. Rather, the following description also clarifies the possible embodiments regarding the above embodiments. Therefore, the following description uses the reference numerals of the figures already shown above, and as a result, each possible embodiment described below also defines possible variations of the above embodiments. Most of these variations can be individually transferred to the above embodiments.

[0039] In other words, embodiments of the present application describe a method for reducing the rate of auxiliary information in image and video encoding applications by integrating samples, i.e., syntactic elements associated with a particular set of blocks, in order to transmit related encoding parameters. Embodiments of the present application can consider, in particular, the combination of the integration syntactic elements with the splitting of parts of the image into different splitting patterns, and the combination with the SKIP / DIRECT mode where the encoding parameters are derived from the spatial and / or temporal neighborhood of the current block. In that context, the above embodiments can be modified to perform the integration for samples, i.e., sets of blocks, in combination with different splitting patterns and the SKIP / DIRECT mode.

[0040] Furthermore, before explaining these variations and details, an overview regarding image and video codecs is shown.

[0041] In image and video coding applications, a sample array associated with an image is typically partitioned into a particular set (or sets) of samples that can represent arbitrarily shaped regions, triangles, or other collections of samples, including rectangular or square blocks, or other sets of samples that can contain arbitrarily shaped regions, triangles, or other shapes. The subdivision of the sample array can be fixed by syntax, or the subdivision can be signaled (at least in part) within the bitstream. To keep the auxiliary information rate for signaling the subdivision information small, the syntax typically allows only a limited number of choices, such as simple partitioning of blocks into smaller blocks. Commonly used partitioning schemes include partitioning a square block into four smaller square blocks, or into two rectangular blocks of the same size, or into two rectangular blocks of different sizes. Here, the actual partitioning used is signaled within the bitstream. A sample set is associated with specific coding parameters that can define, for example, prediction information or residual coding mode. In video coding applications, partitioning is often done for motion representation. All samples of a block (within a partitioning pattern) are associated with the same set of motion parameters, which can include parameters that define the type of prediction (e.g., list 0, list 1, or bi-prediction; and / or translational or affine prediction or prediction using different motion models), parameters that define the reference images used, parameters that typically define the motion of the block relative to the reference images that are sent as a difference to the predictor (e.g., displacement vectors, affine motion parameter vectors, motion parameter vectors for other motion models), the accuracy of the motion parameters (e.g., 1 / 2 sample or 1 / 4 sample accuracy), parameters that define the weighting of reference sample signals (e.g., for illumination compensation), or parameters that define the interpolation filter used to obtain the motion-compensated prediction signal for the current block. Assume that for each sample set, individual coding parameters are sent (e.g., to define prediction and / or residual coding).To obtain improved coding efficiency, the present invention shows a method and specific embodiments for integrating two or more sample sets into a so-called group of sample sets. All sample sets of this kind of group share the same coding parameters. And it can be transmitted together with one of the sample sets in the group. By doing so, the coding parameters do not need to be transmitted individually for each sample set of the group of sample sets. Instead, the coding parameters are transmitted only once for the entire group of sample sets. As a result, the auxiliary information rate for transmitting the coding parameters is reduced and the overall coding efficiency is improved. As an alternative approach, additional refinements for one or more of the coding parameters can be transmitted for one or more of the sample sets of the group of sample sets. The refinement can be applied to all sample sets of the group or only to the sample set to which it is transmitted.

[0042] Embodiments of the present invention relate particularly to combinations of block splitting and integration processing for various sub-blocks 50, 60 (as described above). Typically, an image or video encoding system supports various splitting patterns for block 40. As an example, a square block can be not split, or split into four square blocks of the same size, two rectangular blocks of the same size (where the square block is split vertically or horizontally), or rectangular blocks of different sizes (vertically or horizontally). The typical partitioning patterns described are shown in FIG. 6. In addition to the above description, splitting can also involve splitting at one or more levels. For example, a square sub-block can also optionally be further split using the same splitting pattern. The problem that arises when this kind of splitting process is combined with an integration process that enables a (square or rectangular) block to be integrated with, for example, one of its adjacent blocks is that the splitting that occurs can be obtained by different combinations of the splitting pattern and the integration signal. Therefore, the same information can be transmitted in the bitstream using different coding words, which clearly does not reach an optimal state with respect to coding efficiency. As a simple example, consider a square block that is not further split (as shown in the upper left corner of FIG. 6). This splitting can be directly signaled by sending a syntax element that this block 40 is not subdivided. However, the same pattern can also be signaled by sending a syntax element that defines that this block is subdivided into, for example, two vertically (or horizontally) arranged rectangular blocks 50, 60. Then, integration information can be sent that defines that the second of these rectangular blocks is integrated with the first rectangular block, resulting in exactly the same splitting as when sending a signal that the block is not further split. The same can be obtained by first defining that the block is subdivided into four square sub-blocks and then sending integration information that efficiently integrates all four of these blocks. This concept clearly does not reach an optimal state (because we have different coding words to signal the same thing).

[0043] Embodiments of the present invention relate to concepts and possibilities for increasing coding efficiency for combinations with concepts for integration using the concept of reducing auxiliary information rate and thus providing different partitioning patterns for blocks. Looking at the example of the partitioning pattern in FIG. 6, the "simulation" of a block that was not further partitioned by either of the partitioning patterns using two rectangular blocks can be avoided when the rectangular block is prohibited from being integrated with the first rectangular block (i.e., excluded from the bitstream syntax specification). Looking deeper into the problem, it is also possible to "simulate" a pattern that was not subdivided by integrating the second rectangle with other adjacent (i.e., non-first rectangular blocks) related to the same parameters (e.g., information for defining a prediction) as the first rectangular block. Embodiments of the present invention condition the transmission of integration information in a way that the transmission of certain integration parameters is excluded from the bitstream syntax when these integration parameters result in a pattern that can also be obtained by signaling one of the supported partitioning patterns. As an example, if the current partitioning pattern defines the subdivision into two rectangular blocks as shown in FIGS. 1 and 2 before sending integration information for the second block, i.e., block 60 in the cases of FIGS. 1 and 2, it can be checked which of the possible integration candidates have the same parameters (e.g., parameters for defining a prediction signal) as the first rectangular block, i.e., block 50 in the cases of FIGS. 1 and 2. Then, all candidates having the same motion parameters (including the first rectangular block itself) are removed from the set of integration candidates. The coded word or flag sent to signal the integration information is adapted to the resulting set of candidates. If the set of candidates becomes empty by the parameter check, the integration information is not sent. If the set of candidates consists of exactly one entry, only whether that block is to be integrated is signaled, and since that candidate is obtained on the decoder side etc., it does not need to be signaled.Regarding the above example, the same concept is also used for a partitioning pattern that divides a square block into four smaller square blocks. Here, the transmission of the integration flag is adapted in a way that neither a partitioning pattern that does not define subdivision nor two partitioning patterns that define subdivision into two rectangular blocks of the same size can be achieved by a combination of integration flags. Although the above example using a specific partitioning pattern has been described with the most common concept, it is clear that the same concept (avoiding the specification of a particular partitioning pattern by a combination with integration information corresponding to other partitioning patterns) can be used for other sets of partitioning patterns.

[0044] For a concept where only partitioning is allowed, the advantage of the present invention described is that much greater freedom is provided with respect to signaling the partitioning of an image into parts related to the same parameters (for example, for defining a prediction signal). As an example, additional partitioning patterns resulting from the integration of square blocks of a larger subdivided block are represented in FIG. 7. However, it should be noted that a very large number of resulting patterns can be achieved by integration with further adjacent blocks (outside the previously subdivided blocks). With only a few coding words used to signal the partitioning and integration information, various partitionabilities are provided, and the coder can select the best option in terms of rate distortion (for example, by minimizing a particular rate distortion measure) due to a certain coder complexity. The advantage of an approach where only one partitioning pattern (for example, subdivision into four blocks of the same size) is provided in combination with an integration approach is that commonly used patterns (such as rectangles of different sizes) can be signaled by short coding words instead of several subdivision and integration flags.

[0045] Another aspect that needs to be considered is that the integration concept is similar in the sense that it is in the SKIP or DIRECT mode found in video coding design. In the SKIP / DIRECT mode, basically, the motion parameters are not transmitted for the current block but are derived from spatial and / or temporal neighbors. In a particular efficient concept of the SKIP / DIRECT mode, a list of motion parameter candidates (reference frame index, displacement vector, etc.) is generated from spatial and / or temporal neighbors, and an index to this list that specifies which of the candidate parameters is selected is transmitted. For a bi-predicted block (or multiple hypothesis frames), another candidate can be signaled for each reference list. The possible candidates can include the block above the current block, the block to the left of the current block, the block to the upper left of the current block, the block to the upper right of the current block, the median predictor of various combinations of these candidates, blocks placed at the same position in one or more previous reference frames (or, other already encoded blocks, or combinations obtained from already encoded blocks). When combining the integration concept with the SKIP / DIRECT mode, it must be ensured that both the SKIP / DIRECT mode and the integration mode should not include the same candidates. This can be achieved by different configurations. The SKIP / DIRECT mode can be enabled only for specific blocks (e.g., using a size larger than a specified size, or only for square blocks, etc.) (e.g., using more candidates than the integration mode), and the integration mode for these blocks may not be supported. Or, the SKIP / DIRECT mode can be removed, and all candidates (including parameters indicating combinations of parameters for spatially / temporally adjacent blocks) are added to the integration mode as candidates. This option has already been described above with respect to FIGS. 1 - 5. The increased candidate set may be used only for specific blocks (such as those of a size larger than a predetermined minimum size, or square blocks, etc.). Here, a reduced candidate set is used for other blocks.Alternatively, as a further variation, the integration mode is used for a reduced candidate set (e.g., only the upper and left neighbors), and additional candidates (e.g., the upper left mode, blocks placed at the same position, etc.) are used for the SKIP / DIRECT mode. Also, in this kind of configuration, the SKIP / DIRECT mode may only be possible for certain blocks (those larger than a given minimum size, or square blocks, etc.), while the integration mode is possible for a larger set of blocks. The advantage of this kind of combination is that multiple options for signaling the reuse of already transmitted parameters (e.g., to define prediction) are provided for different block sizes. As an example, more options are provided for larger square blocks. This is because here the additional bitrate spent supplies an increase in rate distortion efficiency. For smaller blocks, a smaller set of options is given. An increase in the candidate set will not result in a gain in rate distortion efficiency here for a small ratio of samples, as each bit required to signal the selected candidate is concerned.

[0046] As described above, embodiments of the present invention also supply the coder with more degrees of freedom to generate a bitstream because the integrated approach significantly increases the number of possibilities for selecting a partition for the sample array of the image. Since the coder can choose from more options, for example, to minimize a specific rate-distortion metric, the coding efficiency can be improved. As an example, some of the additional patterns (e.g., the pattern of FIG. 7) that can be shown by a combination of subdivision and integration can be additionally tested (using the corresponding block sizes for motion estimation and mode decision), and the best patterns provided by purely subdividing (FIG. 6), and by subdivision and integration (FIG. 7) can be selected based on a specific rate-distortion metric. In addition, for each block, it can be tested whether integration using any of the already encoded candidate sets results in a reduction in a specific rate-distortion metric, and the corresponding integration flag is set during the encoding process. In summary, there are several possibilities for operating the coder. In a simple approach, the coder can first determine the maximum subdivision of the sample array (as the highest-level coding scheme). Then, for each sample set, it can check whether integration using other sample sets or other groups of sample sets reduces a specific rate-distortion cost metric. Here, the prediction parameters associated with the integrated group of sample sets can be re-estimated (e.g., by performing a new motion search). Or, the prediction parameters already determined for the current sample set and the candidate sample sets (or groups of sample sets) for integration can be evaluated for the considered group of sample sets. In a more extensive approach, a specific rate-distortion cost metric can be evaluated for additional candidate groups of sample sets. As a special case, when testing various possible subdivision patterns (see, e.g., FIG. 6), some or all of the patterns that can be shown by a combination of subdivision and integration (see, e.g., FIG. 7) can be additionally tested.That is, for all patterns, specific motion estimation and mode determination processing is performed, and the pattern that obtains the smallest rate distortion measure is selected. This processing can also be combined with the above-described low-complexity processing. As a result, for the resulting block, it is additionally tested whether integration with already encoded blocks (e.g., outside the patterns of FIGS. 6 and 7) causes a reduction in the rate distortion measure.

[0047] In the following, for example, with respect to the encoders of FIGS. 1 and 3 and the decoders of FIGS. 2 and 4, some possible detailed embodiments for the embodiments outlined above will be described. As already mentioned above, it is usable in image and video encoding. As described above, a particular set of samples of an image or an array of samples for an image can be decomposed into blocks, which are associated with particular encoding parameters. An image usually consists of a plurality of sample arrays. In addition, an image can be associated with additional auxiliary sample arrays. And it can define, for example, transparency information or a depth map. The sample arrays of an image (including auxiliary sample arrays) can be classified into one or more so-called plane groups. Here, each plane group consists of one or more sample arrays. The plane groups of an image can be encoded independently or, if the image is associated with a plurality of plane groups, by prediction from other plane groups of the same image. Each plane group is usually decomposed into blocks. The blocks (or corresponding blocks of the sample arrays) are predicted by inter-image prediction or intra-image prediction. The blocks can have different sizes and can be square or rectangular. The partitioning of an image into blocks can also be fixed by syntax or it can be signaled (at least partially) in the bitstream. Often, a syntax element that signals the subdivision with respect to blocks of a defined size is transmitted. This kind of syntax element can define whether a block is subdivided into smaller blocks and, for example, for prediction purposes, how it is associated with the encoding parameters. An example of a possible partitioning pattern is shown in FIG. 6. For all samples of a block (or the corresponding block of the sample array), the decoding of the associated encoding parameters is defined in a particular way.In an example, all samples of a block are predicted using the same set of a reference index (identifying a reference image in a set of already encoded images), motion parameters (defining a measure of the motion of the block between the reference image and the current image), parameters for defining an interpolation filter, prediction parameters such as an intra prediction mode, etc. The motion parameters can be indicated by a displacement vector having a horizontal component and a vertical component, or by a higher-order motion parameter such as an affine motion parameter consisting of six components. It is also possible for multiple sets of specific prediction parameters (e.g., the reference index and the motion parameters) to be associated with one block. In that case, for each set of these specific prediction parameters, one intermediate prediction signal for the block (or the corresponding block of the sample array) is generated, and the final prediction signal is constructed by a combination including superimposing the intermediate prediction signals. The corresponding weighting parameters and optionally a constant offset (added to the weighted sum) can be fixed with respect to the image, or with respect to the reference image, or the set of reference images, or they can be included in the set of prediction parameters for the corresponding block. The difference between the original block (or the corresponding block of the sample array) and their prediction signals (also called the residual signal) is usually transformed and quantized. Often, a two-dimensional transformation is applied to the residual signal (or the corresponding sample array for the residual block). For transform coding, the block (or the corresponding block of the sample array) for which a specific set of prediction parameters was used can be further divided before applying the transformation. The transform block can be made equal to or smaller than the block used for prediction. It is also possible for the transform block to include two or more of the blocks used for prediction. Different transform blocks can have different sizes, and the transform block can represent a square or rectangular block.In the above example with respect to FIGS. 1 to 5, it is noted that the first subdivided leaf node, i.e., the coding block 40, can be further divided in parallel, on the one hand, into partitions that define the granularity of the coding parameters and, on the other hand, into transformation blocks to which the two-dimensional transformation is individually applied. After the transformation, the resulting transformation coefficients are quantized, and so-called transformation coefficient levels are obtained. The transformation coefficient levels, as well as the prediction parameters and, if any, the subdivision information, are entropy-coded.

[0048] In the state-of-the-art image and video coding standards, the possibility of subdividing an image (or group of planes) into blocks supplied by the syntax is very limited. Usually, it can only be specified whether a block of a predefined size can be subdivided into smaller blocks (and, in some cases, how). For example, the maximum block size in H.264 is 16x16. A 16x16 block, also called a macroblock, is divided into macroblocks in the first step of each image. For each 16x16 macroblock, it can be signaled whether it is encoded as a 16x16 block, or as two 16x8 blocks, or as two 8x16 blocks, or as four 8x8 blocks. When a 16x16 block is subdivided into four 8x8 blocks, each of these 8x8 blocks can be encoded as one 8x8 block, or as two 8x4 blocks, or as two 4x8 blocks, or as four 4x4 blocks. The small set of possibilities for defining the subdivision in the state-of-the-art image and video coding standards blocks has the advantage that the auxiliary information rate for signaling the subdivision information can be kept low, but has the disadvantage that the bit rate required to transmit the prediction parameters for the blocks can become significant, as will be explained below. The auxiliary information rate for signaling the prediction information usually represents a significant amount of the overall bit rate for the blocks. And when this auxiliary information is reduced, the coding efficiency can be increased. And this can be achieved, for example, by using larger block sizes. It is also possible to increase the set of supported subdivision patterns compared to H.264. For example, the subdivision patterns shown in FIG. 6 can be supplied for square blocks of all sizes (or selected sizes). The real image or picture of a video sequence consists of arbitrarily shaped objects having specific properties. For example, an object of this kind or a part of an object is characterized by a unique texture or a unique motion.And, typically, the same set of prediction parameters can be used for this kind of object or a part of the object. However, the object boundary usually does not coincide with the possible block boundaries for large prediction blocks (e.g., 16x16 macroblocks in H.264). The encoder typically determines a subdivision (among a limited set of possibilities) that results in a minimum of a particular rate-distortion cost measure. For arbitrarily shaped objects, this results in a large number of small blocks. This description also holds when a large number of the above (described as such) splitting patterns are supplied. It should be noted that the amount of splitting patterns should not be too large. This is because a large amount of auxiliary information and / or the computational complexity of the encoder / decoder is required to signal and process these patterns. Thus, arbitrarily shaped objects often result in a large number of small blocks for splitting. And since each of these small blocks is associated with a set of prediction parameters that need to be transmitted, the auxiliary information rate can be a significant part of the overall bitrate. However, since some of the small blocks still represent the same region of the object or a part of the object, the prediction parameters for a large number of the resulting blocks are the same or very similar. Intuitively, when the syntax is extended in a way that not only allows for splitting the blocks but also allows for integrating two or more blocks obtained after splitting, the coding efficiency can be increased. As a result, a group of blocks encoded by the same prediction parameters is obtained. The prediction parameters for this kind of group of blocks only need to be encoded once. In the above examples of FIGS. 1 to 5, for example, when integration occurs, i.e., when the reduced set of candidates does not become zero, the coding parameters for the current block 40 are not transmitted. That is, the encoder does not transmit the coding parameters associated with the current block, and the decoder does not expect the bitstream 30 to contain the coding parameters for the current block 40. Rather, according to that particular embodiment, only the refinement information can be transmitted for the integrated current block 40.The determination of the candidate set, its reduction and integration, etc. are performed for the other coding blocks 40 in the image 20. The coding blocks somehow form groups of coding blocks along the coding chain. Therein, the coding parameters for these groups are transmitted only once in total in the bitstream.

[0049] If the bitrate saved by reducing the number of coding prediction parameters is greater than the bitrate additionally spent for coding the integration information, the described integration results in an increased coding efficiency. It must be further stated that the described syntax extension (for integration) supplies the coder with additional degrees of freedom when selecting the division of the image or group of planes into blocks. The coder is not restricted to first performing the subdivision and then checking whether some of the resulting blocks have the same set of prediction parameters. As a simple variant, the coder can first determine the subdivision as a state-of-the-art coding technique. Next, for each block, it can be checked whether the integration with one of its adjacent blocks (or the already determined group of related blocks) reduces the rate-distortion cost metric. Herein, the prediction parameters associated with the new group of blocks can be re-estimated (e.g., by performing a new motion search), or the prediction parameters already determined for the current block and the adjacent block or group of blocks can be evaluated for the new group of blocks. The coder can also directly check the (subset of) patterns supplied by the combination of the division and integration. That is, the motion estimation and mode decision can be made in the resulting shape as already described above. The integration information can be signaled on a block basis. In practice, the integration can also be interpreted as the estimation of the prediction parameters for the current block. Here, the estimated prediction parameters are set equal to the prediction parameters of one of the adjacent blocks.

[0050] In this regard, it should be noted that combinations of different segmentation patterns and integration information can result in the same shape (associated with the same parameters). This is clearly not an optimal state since the same message can be transmitted by different combinations of encoded words. To avoid (or reduce) this drawback, embodiments of the present invention show the concept of prohibiting the same shape (associated with a particular set of parameters) from being signaled by different segmentation and integration syntax elements. Thus, for all blocks of previously subdivided blocks other than the first in the encoding order, in encoders and decoders such as 10 or 50, for all integration candidates, it is checked whether the integration results in a pattern that can be signaled by segmentation without integration information. All candidate blocks for which this is true are removed from the set of integration candidates, and the transmitted integration information is applied to the resulting candidate set. If no candidates remain, the integration information is not transmitted. If only one candidate remains, a flag is transmitted that defines whether or not that block is integrated, etc. As a further example of this concept, preferred embodiments are described later. The advantage of the described embodiment regarding the concept where only segmentation is allowed is that much greater freedom is given to signal the segmentation of an image to parts (e.g., to define a prediction signal) associated with the same parameters. The advantage compared to an approach where only one segmentation pattern (e.g., subdivision into four blocks of the same size) is combined with the integration approach and supplied is that frequently used patterns (such as rectangles of different sizes) can be signaled by short encoded words instead of several segmentation and integration flags.

[0051] Modern video coding standards such as H.264 also include specific inter - coding modes called SKIP and DIRECT modes. In those, the parameters that define the prediction are fully derived from spatially and / or temporally adjacent blocks. The difference between SKIP and DIRECT is that the SKIP mode further signals that the residual signal is not sent. In various proposed improvements to the SKIP / DIRECT modes, instead of one candidate (such as in H.264), a list of possible candidates is derived from the spatial and / or temporal neighborhood of the current block. The possible candidates can include the block above the current block, the block to the left of the current block, the block to the upper - left of the current block, the block to the upper - right of the current block, the median predictor of various of these candidates, blocks placed at the same position in one or more previous reference frames (or other already - coded blocks, or combinations obtained from already - coded blocks). With respect to combinations with the integrated mode, it must be ensured that both the SKIP / DIRECT mode and the integrated mode do not contain the same candidates. This can be achieved by various configurations as described above. The advantage of the described combinations is that multiple options for signaling the reuse of already - transmitted parameters (for example, to define the prediction) are provided for different block sizes.

[0052] One advantage of an embodiment of the present invention is to reduce the bit rate required to transmit prediction parameters by integrating adjacent blocks into groups of blocks. Here, each group of blocks is associated with a unique set of encoding parameters, such as prediction parameters or residual encoding parameters. The integrated information is signaled in the bitstream, (if any, in addition to the segmentation information). By combining various partitioning patterns with the SKIP / DIRECT mode and sending the corresponding integrated information, it can be ensured that neither the SKIP / DIRECT mode nor any of the supplied patterns are "simulated". The advantage of an embodiment of the present invention is the increased encoding efficiency resulting from the reduced auxiliary information rate for the encoding parameters. Embodiments of the present invention are applicable in image and video encoding applications. Therein, a set of samples is associated with specific encoding or prediction parameters. The integration process described here can also be extended to three dimensions or higher dimensions. For example, blocks of a group of several video images can be integrated into one group of blocks. It can also be applied to 4D compression of light field encoding. On the one hand, it can also be used for compression of 1D signals. Here, the 1D signal is partitioned and a given partition is integrated.

[0053] Embodiments of the present invention also relate to a method for reducing the auxiliary information rate in image and video coding applications. In image and video coding applications, a particular set of samples (which can represent any collection of rectangular or square blocks or arbitrarily shaped regions or other samples) is typically associated with a particular set of coding parameters. For each of these sample sets, the coding parameters are included in the bitstream. The coding parameters can indicate prediction parameters that define how the corresponding set of samples is predicted using previously encoded samples. The partitioning of the sample array of an image into sample sets can be fixed by syntax or signaled by corresponding subdivision information in the bitstream. Multiple partitioning patterns can be allowed for blocks. The coding parameters for a sample set are sent in a predefined order given by the syntax. Embodiments of the present invention also show a way to signal for the current set of samples that is integrated (e.g., for prediction purposes) with one or more other sample sets into a group of sample sets. Thus, a possible set of values for the corresponding integration information is applied to the partitioning pattern used in a way that a particular partitioning pattern cannot be indicated by a combination of other partitioning patterns and corresponding integration data. The coding parameters for a group of sample sets need to be sent only once. In a particular embodiment, when the current sample set is integrated with a sample set (or group of sample sets) for which the coding parameters have already been sent, the coding parameters for the current sample set are not sent. Instead, the coding parameters for the current sample set are set equal to the coding parameters of the sample set (or group of sample sets) with which the current sample set is integrated. As an alternative approach, additional refinements for one or more of the coding parameters can be sent for the current sample set.The refinement can be applied to all sample sets of the group or only to the sample sets to which it is sent.

[0054] In a preferred embodiment, for each set of samples, the set among all previously encoded sample sets is referred to as the "set of causal sample sets". The set of samples that can be used for integration with the current set of samples is referred to as the "set of candidate sample sets" and is always a subset of the "set of causal sample sets". The way this subset is formed can be known to the decoder or it can be specified within the bitstream. In any case, the encoder 10 and the decoder 80 determine the candidate set that will be reduced. If a particular current set of samples is encoded and that set among the candidate sample sets is not empty, the current set of samples in the samples is integrated with one of the sample sets in this set of candidate sample sets or (if so, / if multiple candidates exist) signaled for any of them. Otherwise, integration cannot be used for this block. Candidate blocks that result in a shape where the integration is also directly specified by the split pattern are excluded from the candidate set to avoid the same shape being shown by different combinations of split information and integration data. That is, by removing each candidate as described above with respect to FIGS. 1 to 5, the candidate set is reduced.

[0055] In a preferred embodiment, a set of some candidate sample sets is a set of zero or more sample sets that includes at least a specific non-zero number of samples (which can be one, two, or more) that indicate the direct spatial adjacency of every sample within the current set of samples. In other preferred embodiments of the present invention, the set of candidate sample sets additionally (or exclusively) includes a set of samples that have the same spatial position, i.e., are included by both the candidate sample set and the current sample set currently being affected by integration, but are included in different images, and includes a specific non-zero number of samples (which can be one, two, or more). In other preferred embodiments of the present invention, the set of candidate sample sets can be derived from previously processed data within the current image or in other images. The derivation method can include spatial direction information such as a specific direction or a conversion coefficient related to the image gradient of the current image, or it can include temporal direction information such as adjacent motion displays. From this type of data and other data (if any) and auxiliary information available to the receiver, the set of candidate sample sets can be derived. The removal of candidates (from the original candidate set) that result in the same shape that can be represented by a specific segmentation pattern is derived in the same way in the encoder and decoder, and as a result, the encoder and decoder derive the final candidate set for integration in exactly the same way.

[0056] In a preferred embodiment, the considered set of samples is a rectangular or square block. Then, the integrated set of samples indicates a collection of rectangular and / or square blocks. In other preferred embodiments of the present invention, the considered set of samples is an arbitrarily shaped image region, and the integrated set of samples indicates a collection of arbitrarily shaped image regions.

[0057] In a preferred embodiment, one or more syntax elements are sent for each set of samples. And they define whether the set of samples will be integrated with other sets of samples (which may be part of an already integrated group of sample sets) and which set of candidate sample sets will be used for integration. However, if the candidate set is empty (e.g., for removal of candidates that would cause a split signaled by different split patterns without integration), the syntax elements are not sent.

[0058] In a preferred embodiment, one or two syntax elements are sent to define integration information. The first syntax element defines whether the current set of samples will be integrated with other sets of samples. The second syntax element, which is only sent if the first syntax element defines that the current set of samples will be integrated with other sets of samples, defines which of the set of candidate sample sets will be used for integration. In a preferred embodiment, the first syntax element is only sent if the derived set of candidate sample sets is not empty (after potential removal of candidates that would cause a split signaled by different split patterns without integration). In other preferred embodiments, the second syntax element is only sent if the derived set of candidate sample sets contains one or more sets of samples. In a further preferred embodiment of the present invention, the second syntax element is only sent if at least two sets of samples in the derived set of candidate sample sets are associated with different coding parameters.

[0059] In a preferred embodiment of the present invention, the integration information for a set of samples is encoded before the prediction parameters (or more generally, the specific coding parameters associated with the set of samples). The prediction or coding parameters are only sent if the integration information signals that the current set of samples will not be integrated with other sets of samples.

[0060] In another preferred embodiment, after a subset of the prediction parameters (or, more generally, the specific encoding parameters associated with the sample set) has been transmitted, the integrated information for the set of samples is encoded. The subset of prediction parameters can be composed of, for example, one or more reference image indices, or one or more components of the motion parameter vector, or a reference index, as well as one or more components of the motion parameter vector. The already transmitted subset of prediction or encoding parameters can be used to derive a (reduced) set of candidate sample sets. As an example, a measure of the difference between the already encoded prediction or encoding parameters and the corresponding prediction or encoding parameters of the original set of candidate sample sets can be calculated. And only those sample sets for which the calculated difference measure is less than or equal to a predetermined or derived threshold are included in the final (reduced) set of candidate sample sets. The threshold can be obtained based on the calculated difference measure. Or, as another example, only those sets of samples that minimize the difference measure are selected. Or, only one set of samples is selected based on the difference measure. In the latter case, the integrated information can be reduced in a way that only specifies whether the current set of samples is integrated into one of the candidate sets of samples.

[0061] The following preferred embodiments are described for a set of samples showing rectangular and square blocks, but it can be extended directly to arbitrarily shaped regions or other sets of samples.

[0062] 1. Derivation of the first set of candidate blocks The derivation of the first set of samples described in this section relates to the derivation of the first set of candidates. Some of all the candidate blocks may be removed later by analyzing the relevant parameters (e.g., prediction information) and the removal of those candidate blocks that result in the final segmentation that can also be obtained by using other segmentation patterns. This process is described in the next subsection.

[0063] In a preferred embodiment, the first set of candidate blocks is formed as follows. Starting from the sample position at the upper left of the current block, its adjacent sample position to the left and its adjacent sample position at the top are derived. The first set of candidate blocks can have at most only two elements, i.e., those blocks from a set of cause blocks that includes only one of the two sample positions. Thus, the first set of candidate blocks can have, as its elements, only the two directly adjacent blocks of the sample position at the upper left of the current block.

[0064] In another preferred embodiment of the present invention, the set of first candidate blocks is encoded before the current block and is given by all blocks that include one or more samples indicating a direct spatial adjacency of any sample of the current block (the direct spatial adjacency may be limited to the direct left neighbor and / or the direct upper neighbor and / or the direct right neighbor and / or the direct lower neighbor). In another preferred embodiment of the present invention, the set of first candidate blocks additionally (or exclusively) includes blocks that are at the same position as any of the samples of the current block but are included in a different (already encoded) image and include one or more samples. In another preferred embodiment of the present invention, the first candidate set of blocks represents a subset of the above-described set of (adjacent) blocks. The subset of candidate blocks may be fixed, signaled, or derived. The derivation of the subset of candidate blocks can take into account decisions made for other blocks of the image or of other images. As an example, blocks associated with the same (or very similar) encoding parameters as other candidate blocks may not be included in the first candidate set of blocks.

[0065] In a preferred embodiment of the present invention, the set of first candidate blocks is derived with respect to one of the above embodiments but has the following limitation. Only blocks using motion compensation prediction (inter prediction) can be elements of the set of candidate blocks. That is, intra-encoded blocks are not included in the (first) candidate set.

[0066] As already mentioned above, it is possible to expand the list of candidates by additional candidates for block integration such as combined bi-prediction integration candidates, unscaled bi-prediction integration candidates, and zero motion vectors.

[0067] The derivation of the first set of candidate blocks is similarly performed by both the encoder and the decoder.

[0068] 2. Derivation of the Last Set of Candidate Blocks After deriving the initial candidate set, the relevant parameters of the candidate blocks within the initial candidate set are analyzed such that integration candidates that would result in a split that can be shown by using different split patterns are removed. If the sample arrays that can be integrated are of different shapes and / or sizes, the same split that can be indicated by at least two different encoding words may exist. For example, if the encoder determines to split a sample array into two sample arrays, this split will be reversed by integrating the two sample arrays. To avoid this kind of redundant description, the set of candidate blocks for integration is constrained according to certain block shapes and splits that are allowed. On the one hand, the allowed shapes of the sample arrays can be constrained according to the specific candidate list used for integration. The two functions of splitting and integration need to be designed together so that when combined, redundant description is avoided.

[0069] In a preferred embodiment of the invention, a set of splitting modes (or split modes) shown in FIG. 6 is supported for square blocks. When a square block of a specific size is split into four small square blocks of the same size (the lower left pattern in FIG. 6), the same set of split patterns can be applied to the resulting four square blocks so that hierarchical splitting can be defined.

[0070] After deriving the first set of candidate blocks, the candidate list is reduced as follows. - If the current block is not further split (the upper left pattern in FIG. 6), the initial candidate list is not reduced. That is, all initial candidates represent the final candidates for integration. ​- When the current block is split into exactly two blocks of any size, one of these two blocks is encoded before the other, which is determined by the syntax. For the first encoded block, the first candidate set is not reduced. However, for the second encoded block, all candidate blocks having the same relevant parameters as the first block are removed from the candidate set (which includes the first encoded block). - When the block is split into four square blocks of the same size, the first candidate list for the first three blocks (in the encoding order) is not reduced. All blocks in the first candidate list also exist in the final candidate list. However, for the fourth (last) block in the encoding order, the following applies. - When a block in a different row from the current block (in the splitting pattern illustrated in the lower left of FIG. 6) has the same relevant parameters (e.g., motion parameters), all candidates having the same motion parameters as the already encoded block in the same row as the current block are removed from the candidate set (which includes the blocks in the same row). - When a block in a different column from the current block (in the splitting pattern illustrated in the lower left of FIG. 6) has the same relevant parameters (e.g., motion parameters), all candidates having the same motion parameters as the already encoded block in the same column as the current block are removed from the candidate set (which includes the blocks in the same column).

[0071] In a low-complexity variation of the embodiment (using the splitting pattern of FIG. 6), the reduction of the candidate list is performed as follows. - When the current block is not further split (the upper left pattern of FIG. 6), the first candidate list is not reduced. That is, all first candidates represent the final candidates for integration. - When the current block is now split into exactly two blocks of any size, one of these two blocks is encoded before the other, which is determined by the syntax. For the first encoded block, the first candidate set is not reduced. However, for the second encoded block, the first encoded block of the split pattern is removed from the candidate set. - When the block is split into four square blocks of the same size, the first candidate list for the first three blocks (in encoding order) is not reduced. All blocks of the first candidate list are also present in the final candidate list. However, for the fourth (last) block in encoding order, the following applies. - When signaling that integration information is integrated into the first encoded block of that row for a block in another row that is encoded later (than the current block), blocks in the same row as the current block are removed from the candidate set. - When signaling that integration information is integrated into the first encoded block of that column for a block in another column that is encoded later (than the current block), blocks in the same column as the current block are removed from the candidate set.

[0072] In other preferred embodiments, the same split pattern as shown in FIG. 6 is supported without the pattern of splitting a square block into two rectangular blocks of the same size. Except for the pattern of splitting the block into four square blocks, candidate list reduction proceeds as described by any of the above embodiments. Here, all initial candidates are allowed for all sub-blocks, or only the candidate list for the last encoded sub-block is constrained as follows. If the three previously encoded blocks are associated with the same parameters, all candidates associated with these parameters are removed from the candidate list. In the low-complexity version, if these three sub-blocks are combined together, the last encoded sub-block cannot be integrated into any of the three previously encoded sub-blocks.

[0073] In another preferred embodiment, different sets of split patterns for blocks (or other forms of sample array sets) are supported. For sample array sets that are not split, all candidates in the first candidate list can be used for integration. When the sample array is split into exactly two sample arrays, for the first sample array in the encoding order, all candidates in the first candidate set are inserted into the final candidate set. For the second sample array in the encoding order, all candidates having the same relevant parameters as the first sample array are removed. Or, in a low-complexity variation, only the first sample array is removed from the candidate set. For split patterns that split the sample array into more than two sample arrays, the candidate removal depends on whether other split patterns can be simulated by the current split pattern and the corresponding integration information. The candidate removal process follows the concepts explicitly shown above, but takes into account the actually supported candidate patterns.

[0074] In another preferred embodiment, when the SKIP / DIRECT mode is supported for a particular block, the integration candidates that are also the current candidates for the SKIP / DIRECT mode are removed from the candidate list. This removal can replace or can be used in conjunction with the candidate block removal described above.

[0075] 3. Combination with SKIP / DIRECT Mode The SKIP / DIRECT mode can be supported for all or only certain block sizes and / or block shapes. A set of candidate blocks is used for the SKIP / DIRECT mode. The difference between SKIP and DIRECT is whether residual information is sent. The parameters of SKIP and DIRECT (e.g., for prediction) are presumed to be equal to any of the corresponding candidates. Candidates are selected by sending an index to the candidate list.

[0076] In a preferred embodiment, the candidate list for SKIP / DIRECT can include different candidates. An example is shown in FIG. 8. The candidate list can include the following candidates (the current block is denoted by Xi): - Median (median) (between the left, upper, and corner) - Left block (Li) - Above block (Ai) - Corner blocks (in order: upper right (Ci1), lower left (Ci2), upper left (Ci3)) - Different but Collocated blocks of the already encoded image

[0077] In a preferred embodiment, the candidates for integration include Li (left block) and Ai (above block). Selecting these candidates for integration requires a small amount of auxiliary information to signal which block the current block is integrated into.

[0078] The following symbols are used to describe the following embodiments: - set_mvp_ori is the set of candidates used for the SKIP / DIRECT mode. This set is composed of {Median, Left, Above, Corner, Collocated}. Here, Median is the median (the median in the ordered set of Left, Above, and Corner), and collocated is given by the closest reference frame and scaled according to the temporal distance. - set_mvp_comb is the set of candidates used for the SKIP / DIRECT mode in combination with the block integration process.

[0079] For a preferred embodiment, the combination between the SKIP / DIRECT mode and the block integration mode can be processed by the original set of candidates. This means that the SKIP / DIRECT mode has the same set of candidates as when it operates alone. The interest in combining these two modes comes from their complementarity in signaling auxiliary information between frames. Despite the fact that both of these modes use neighboring information to improve the signaling of the current block, block integration processes only the left and upper neighbors, while the SKIP / DIRECT mode processes up to five candidates. The main complementarity lies in the different approaches to processing neighboring information. The block integration process holds the entire set of its neighboring information for all reference lists. This means that block integration holds not only its motion vectors for each reference list, but also all the auxiliary information from these neighbors. On the other hand, the SKIP / DIRECT mode processes the prediction parameters separately, for each reference list, and transmits an index to the candidate list for each reference list. That is, for a bi-predicted image, two indexes are transmitted to signal candidates for reference list 0 and candidates for reference list 1.

[0080] In another preferred embodiment, a combined set of candidates called set_mvp_comb can be found for the SKIP / DIRECT mode in combination with the block integration mode. This combined set is part of the original set (set_mvp_ori), and since the list of candidates set_mvp_comb is reduced, it enables a reduction in signaling for the SKIP / DIRECT mode. The candidates that must be removed from the original list (set_mvp_ori) could have been redundant by the block integration process or were not used much.

[0081] In another preferred embodiment, the combination between the SKIP / DIRECT mode and the block integration process can be processed by a combined set of candidates (set_mvp_comb) that is the original set (set_mvp_ori) without Median. Due to the low efficiency observed for Median for the SKIP / DIRECT mode, that reduction of the original list results in an improvement in coding efficiency.

[0082] In another preferred embodiment, the combination of the SKIP / DIRECT mode and the block integration can be processed by a combined set of candidates (set_mvp_comb) that is the original set (set_mvp_ori) having only Corner and / or Collocated as candidates.

[0083] In another preferred embodiment, the combination of the SKIP / DIRECT mode and the block integration process can be processed by a combined set of candidates that is set_mvp_ori having only Corner and Collocated. As already mentioned, despite the complementarity between the SKIP / DIRECT mode and the block integration, the candidates that need to be removed from the list can overlap with the candidates for the block integration process. These candidates are Left and Above. The combined set of candidates (set_mvp_comb) is reduced to only two candidates, Corner and Collocated. The SKIP / DIRECT mode using this candidate set set_mvp_comb combined with the block integration process gives a highly efficient increase in signaling auxiliary information between frames. In this embodiment, the SKIP / DIRECT mode and the integration mode do not share any candidate blocks.

[0084] In a further embodiment, slightly different combinations of SKIP / DIRECT and integrated modes can be used. It is possible to enable the SKIP / DIRECT mode only for specific blocks (e.g., blocks having a size larger than a specified size, or only for square blocks, etc.) and not support the integrated mode for these blocks (e.g., for more candidates than the integrated mode). Or, the SKIP / DIRECT mode can be disabled, and all candidates (including parameters indicating combinations of parameters for spatially / temporally adjacent blocks) are added to the integrated mode as candidates. This option is illustrated in FIGS. 1 to 5. The increased candidate set can be used only for specific blocks (e.g., blocks having a size larger than a specified minimum size, or only for square blocks, etc.). Here, for other blocks, a reduced candidate set is used. Or, as a further variation, the integrated mode is used with a reduced candidate set (e.g., only the upper and left neighbors), and additional candidates (e.g., the upper left neighbor, blocks arranged at the same position, etc.) are used for the SKIP / DIRECT mode. Also, in this type of configuration, the SKIP / DIRECT mode can be allowed only for specific blocks (e.g., blocks having a size larger than a specified minimum size, or for square blocks, etc.), while the integrated mode is allowed for a larger set of blocks.

[0085] 4. Transmission of integrated information Regarding the preferred embodiments, and in particular regarding the embodiments of FIGS. 1 to 5, the following may apply. Assume that only two blocks are considered as candidates, including the sample to the left and the sample adjacent above the sample in the upper left of the current block. (After the removal of candidates as described above) If the set of final candidate blocks is not empty, one flag called merge_flag is signaled that defines whether the current block is integrated into any of the candidate blocks. If merge_flag is equal to 0 (for "false"), this block is not integrated into one of its candidate blocks and all coding parameters are sent normally. If merge_flag is equal to 1 (for "true"), the following applies. If the set of candidate blocks contains only one block, this candidate block is used for integration. Otherwise, the set of candidate blocks contains exactly two blocks. If the prediction parameters of these two blocks are the same, these prediction parameters are used for the current block. Otherwise (if the two blocks have different prediction parameters), a flag called merge_left_flag is signaled. If merge_left_flag is equal to 1 (for "true"), the block containing the adjacent sample position to the left of the sample position in the upper left of the current block is selected from the set of candidate blocks. If merge_left_flag is equal to 0 (for "false"), the other (i.e., the adjacent above) block in the set of candidate blocks is selected. The prediction parameters of the selected block are used for the current block. In other embodiments, a combined syntax element that signals the integration process is sent. In other embodiments, merge_left_flag is sent regardless of whether the two candidate blocks have the same prediction parameters.

[0086] Note that the syntax element merge_left_flag could be named merge_index as its function is to perform a selected one-indexing among the non-removed candidates.

[0087] In other preferred embodiments, two or more blocks can be included in the set of candidate blocks. The integration information (i.e., whether the block is integrated and if so, into which candidate block) is signaled by one or more syntax elements. Here, the set of coded words depends on the number of candidates in the final candidate set and is selected in the same way by the coder and decoder. In one embodiment, the integration information is transmitted using one syntax element. In other embodiments, one syntax element defines whether the block is integrated with any of the candidate blocks (compared to the above merge_flag). This flag is transmitted only if the set of candidate blocks is not empty. The second syntax element signals which of the candidate blocks is used for integration. It is transmitted only if the first syntax element signals that the current block is integrated with one of the candidate blocks. In a preferred embodiment of the present invention, the second syntax element is transmitted only if the set of candidate blocks includes one or more candidate blocks and / or any of the candidate blocks has different prediction parameters from the others. The syntax can depend on how many candidate blocks are given and / or how the different prediction parameters are related to the candidate blocks.

[0088] It is possible to add a set of candidates for block integration as done for DIRECT mode.

[0089] As described in other preferred embodiments, the second syntax element integration index may be transmitted only if the list of candidates contains one or more candidates. This requires deriving the list before parsing the integration index, preventing these two processes from being executed in parallel. To allow for increased parsing throughput and to make the parsing process more robust with respect to transmission errors, this dependency can be removed by using a fixed codeword for each index value and a fixed number of candidates. If this number is not reached by candidate selection, auxiliary candidates can be derived to complete the list. These additional candidates can include so-called combined candidates, which are constructed from the possible motion parameters of different candidates that may already be in the list, and zero motion vectors.

[0090] In other preferred forms, a syntax for signaling which of the candidate blocks are set is applied simultaneously at the encoder and decoder. For example, if three selections of blocks for integration are given, those three selections are only present in the syntax and are considered for entropy coding. The probability for all other selections is considered to be 0, and the entropy codec is adjusted simultaneously at the encoder and decoder.

[0091] The predicted parameters determined as a result of the integration process can indicate the entire set of predicted parameters associated with the block, or they can indicate within a subset of these predicted parameters (e.g., the predicted parameters for one hypothesis of a block for which multiple hypothesis prediction is used).

[0092] In a preferred embodiment, the syntax elements related to the integration information are entropy coded using context modeling. The syntax elements can consist of the above-mentioned merge_flag and merge_left_flag.

[0093] In a preferred embodiment, one of the three context models is used to encode the merge_flag. The used context model merge_flag_ctx is obtained as follows. When the set of candidate blocks contains two elements, the value of merge_flag_ctx is equal to the sum of the merge_flag values of the two candidate blocks. When the set of candidate blocks contains one element, the value of merge_flag_ctx is equal to twice the merge_flag value of this one candidate block.

[0094] In a preferred embodiment, merge_left_flag is encoded using one probability model.

[0095] Different context model encodings for merge_idx(merge_left_flag) can be used.

[0096] In other embodiments, different context models may be used. Non-binary syntax elements can be mapped to a sequence of binary symbols (bins). The context model for some syntax elements or bins of syntax elements can be derived based on the already transmitted syntax elements of adjacent blocks or the number of candidate blocks or other metrics, while other syntax elements or bins of syntax elements can be encoded with a fixed context model.

[0097] 5. Encoder Operations Since the integrated approach will, of course, with increased signaling overhead, significantly increase the number of possibilities for selecting a partition for the sample array of the image, the inclusion of the integration concept supplies the encoder with greater freedom for the generation of the bitstream. Some or all of the additional patterns that can be shown by the combination of subdivision and integration (e.g., the pattern of FIG. 7 when the partition pattern of FIG. 6 is supported) can be additionally tested (using the corresponding block sizes for motion estimation and mode decision), and the best of the patterns supplied by simply subdividing (FIG. 6) and by subdividing and integrating (FIG. 7) can be selected based on a particular rate distortion metric. Additionally, for each block, it can be tested whether integration with any of the already encoded candidate sets results in a reduction of a particular rate distortion metric, and then whether the corresponding integration flag is set during the encoding process.

[0098] In another preferred embodiment, the encoder can first determine the highest subdivision of the sample array (as in the highest level of the code system). Then, for each sample set, it can be checked whether integration with other sample sets or other groups of sample sets reduces a particular rate distortion cost metric. In this case, the prediction parameters associated with the integrated group of sample sets can be re-estimated (e.g., by performing a new motion search), or the prediction parameters already determined for the current sample set and candidate sample sets (or groups of sample sets) for integration can be evaluated for the considered group of sample sets.

[0099] In other preferred embodiments, a particular rate distortion cost metric can be evaluated for additional candidate groups of the sample set. As a particular example, when testing various possible partitioning patterns (see, e.g., FIG. 6), some or all of the patterns that can be represented by a combination of partitioning and merging (see, e.g., FIG. 7) can be further tested. That is, for all of the patterns, a particular motion estimation and mode determination process is performed, and the pattern that follows the smallest rate distortion metric is selected. This process can also be combined with the low-complexity process described above, such that, for the resulting blocks, it is further tested whether integration with already encoded blocks (e.g., other than the patterns of FIGS. 6 and 7) results in a reduction of the rate distortion metric.

[0100] In other preferred embodiments, the encoder tests various patterns that can be represented by partitioning and merging in order of priority, which tests as many patterns as possible given certain real-time requirements. The priority can also be changed based on already encoded blocks or selected partitioning patterns.

[0101] One way to transfer the embodiments outlined above to a specific syntax is described below with respect to the following figures. In particular, FIGS. 9-11 show different parts of the syntax that utilize the embodiments outlined above. In particular, according to the embodiments outlined below, the image 20 is first updated to an encoded tree block whose image content is encoded using the syntax coding_tree shown in FIG. 9. As shown there, for example, for entropy_coding_mode_flag = 1 regarding context adaptive binary arithmetic coding or other specific entropy coding modes, the quadtree subdivision of the current encoded tree block is signaled within the syntax part coding_tree by a flag called split_coding_unit_flag at symbol 400. As shown in FIG. 9, according to the embodiments described below, the tree root block is subdivided to be signaled by split_coding_unit_flag in depth-first traversal order as shown in FIG. 9a. Whenever a leaf node is reached, it indicates an encoding unit that is immediately encoded using the syntax function coding_unit. This can be seen in the conditional clause 402 that checks whether the current split_coding_unit_flag is set. If so, the function coding_tree is called recursively, leading to further transmission / derivation of further split_coding_unit_flags in the encoder and decoder, respectively. If not, i.e., if split_coding_unit_flag = 0, the current sub-block of the tree root block 200 in FIG. 5a is a leaf block, and the function coding_unit in FIG. 10 is called at 404 to encode this encoding unit.

[0102] In the embodiments described above, the above options are used according to whether any integration is simply available for images for which the inter-prediction mode can be used. That is, the intra-coded slice / image does not use integration in any case. This can be seen from FIG. 10 where the flag merge_flag is sent at 406 only when the slice type is not equal to the intra-image slice type. According to this embodiment, integration relates only to prediction parameters related to inter-prediction. According to this embodiment, merge_flag is signaled for all coding units 40 and also signals to the decoder a particular partitioning mode for the current coding unit, i.e., the non-partitioning mode. Accordingly, the function prediction_unit is called at 408 with respect to meaning the current coding unit as a prediction unit. However, this is not the only possibility to switch to the integration option. Rather, if the merge_flag related to all coding units is not set at 406, the prediction type of the coding unit of the non-intra-image slice depends on it and is signaled at 410 by the syntax element pred_type by calling the function prediction_unit for any partitioning of the current coding unit, for example at 412, if the current coding unit is not further partitioned. In FIG. 10, only four different partitioning options are shown, but other partitioning options shown in FIG. 6 can equally be available. Another possibility is that the partitioning option PART_NxN is not available but others are. The relationship between the names from the partitioning mode used in FIG. 10 to the partitioning options shown in FIG. 6 is shown in FIG. 6 by each subscript under the individual partitioning options. The function prediction_unit is called for each partition, for example for partitions 50 and 60 in the coding order described above. The function prediction_unit begins by checking the merge_flag at 414. If the merge_flag is set, the merge_index necessarily follows at 416.The check in step 414 is for checking whether the merge_flag associated with all the coding units as signaled in 406 is set. If not, the merge_flag is signaled further in 418, and if the latter is set, the merge_index follows in 420 which indicates the integration candidate for the current partition. Also, the merge_flag is simply signaled in 418 for the current partition if the current prediction mode of the current coding unit is the inter prediction mode (see 422).

[0103] As can be seen from FIG. 11, according to this embodiment, the transmission of the prediction parameters used for the current prediction unit of 424 is performed only if integration is not used for the current prediction unit.

[0104] The above description of the embodiments of FIGS. 9 to 11 already shows most of the functions and meanings, but some further information is shown below.

[0105] merge_flag[x0][y0] defines whether the inter prediction parameters for the current prediction unit (see 50 and 60 in the figure) are derived from adjacent inter predicted partitions. The array indices x0, y0 define the position (x0, y0) of the top-left luma sample of the considered prediction block (see 50 and 60 in the figure) in relation to the top-left luma sample of the image (see 20 in the figure).

[0106] merge_idx[x0][y0] defines the integration candidate index in the integration candidate list. Here, x0, y0 define the position (x0, y0) of the top-left luma sample of the considered prediction block in relation to the top-left luma sample of the image.

[0107] Although not specifically shown in the foregoing description of FIGS. 9-11, the integration candidates or list of integration candidates are determined in this embodiment using, by way of example, the coding parameters or the prediction parameters of spatially adjacent prediction units / partitions. Rather, the list of candidates is also formed by using the prediction parameters of temporally adjacent partitions of temporally adjacent and previously coded images. Further, combinations of prediction parameters of spatially and / or temporally adjacent prediction units / partitions are used and included in the list of integration candidates. Of course, only a subset thereof may be used. In particular, FIG. 12 shows one possibility of determining the spatial adjacency, i.e., spatially adjacent partitions or prediction units. FIG. 12 shows, by way of example, prediction unit or partition 60 and pixels B0-B2 and A0, A1. It is located directly adjacent to the boundary 500 of partition 60, i.e., B2 is diagonally adjacent to the upper left pixel of partition 60, B1 is adjacent vertically above and to the right of the upper right pixel of partition 60, B0 is diagonally located at the upper right pixel of partition 60, A1 is adjacent to the left and lower left pixel of partition 60, and A0 is diagonally located at the lower left pixel of partition 60. Partitions containing at least one of pixels B0-B2 and A0, A1 form a spatial proximity, and their prediction parameters form integration candidates.

[0108] To perform the above removal of those candidates leading to other partitioning modes that were also available, the following function can be used.

[0109] In particular, if any of the following situations are true, candidate N, i.e., the coding / prediction parameters resulting from the prediction unit / partition covering pixel N=(B0,B1,B2,A0,A1), i.e., the position (xN,yN), is removed from the candidate list (see FIG. 6 for the partitioning mode PartMode and the corresponding partitioning index PartIdx indexing each partition within the coding unit).

[0110] - The PartMode of the current prediction unit is PART_2NxN, PartIdx is equal to 1, and the prediction units covering the luma positions (xP, yP-1) (PartIdx = 0) and (xN, yN) (Cand.N) have the same motion parameters. mvLX[xP,yP-1]==mvLX[xN,yN] refIdxLX[xP,yP-1]==refIdxLX[xN,yN] predFlagLX[xP,yP-1]==predFlagLX[xN,yN]

[0111] - The PartMode of the current prediction unit is PART_Nx2N, PartIdx is equal to 1, and the prediction units covering the luma positions (xP-1, yP) (PartIdx = 0) and (xN, yN) (Cand.N) have the same motion parameters. mvLX[xP-1,yP]==mvLX[xN,yN] refIdxLX[xP-1,yP]==refIdxLX[xN,yN] predFlagLX[xP-1,yP]==predFlagLX[xN,yN]

[0112] - The PartMode of the current prediction unit is PART_NxN, PartIdx is equal to 3, and the prediction units covering the luma positions (xP-1, yP) (PartIdx = 2) and (xP-1, yP-1) (PartIdx = 0) have the same motion parameters. mvLX[xP-1,yP]==mvLX[xP-1,yP-1] refIdxLX[xP-1,yP]==refIdxLX[xP-1,yP-1] predFlagLX[xP-1,yP]==predFlagLX[xP-1,yP-1] And the prediction units covering the luma positions (xP, yP-1) (PartIdx = 1) and (xN, yN) (Cand.N) have the same motion parameters. mvLX[xP, yP - 1] == mvLX[xN, yN] refIdxLX[xP, yP - 1] == refIdxLX[xN, yN] predFlagLX[xP, yP - 1] == predFlagLX[xN, yN]

[0113] - The PartMode of the current prediction unit is PART_NxN, the PartIdx is equal to 3, and the prediction units covering the luma positions (xP, yP - 1) (PartIdx = 1) and (xP - 1, yP - 1) (PartIdx = 0) have the same motion parameters. mvLX[xP, yP - 1] == mvLX[xP - 1, yP - 1] refIdxLX[xP, yP - 1] == refIdxLX[xP - 1, yP - 1] predFlagLX[xP, yP - 1] == predFlagLX[xP - 1, yP - 1] And the prediction units covering the luma positions (xP - 1, yP) (PartIdx = 2) and (xN, yN) (Cand.N) have the same motion parameters. mvLX[xP - 1, yP] == mvLX[xN, yN] refIdxLX[xP - 1, yP] == refIdxLX[xN, yN]

[0114] Regarding this point, it should be noted that the position or location (xP, yP) means the point that is the maximum pixel of the current partition / prediction unit. That is, according to the first item, for all coding parameter candidates, it is checked how each coding parameter of the adjacent prediction unit, that is, prediction unit N, is derived by directly adopting it. However, other additional coding parameter candidates can be similarly checked for whether it is equal to the coding parameter of each prediction unit that appears regarding whether it will obtain other split patterns supported by the syntax. According to the embodiment just described, the equality of the coding parameters includes checking the equality of the motion vector, that is, mvLX, the reference index, that is, refIxLX, and the prediction flag predFlagLX that indicates that the motion vector and reference index related to the reference list X of X with its parameter being 0 or 1 are used for inter prediction.

[0115] It should also be noted that the aforementioned possibility for removing coding parameter candidates of adjacent prediction units / partitions is also applicable when supporting the asymmetric split mode shown in the right half of FIG. 6. In that case, the mode PART_2NxN can indicate a mode of all horizontal subdivisions, and PART_Nx2N can correspond to a mode of all vertical subdivisions. Further, the mode PART_NxN can be excluded from the supported split modes or split patterns, in which case only the first two removal checks need to be performed.

[0116] Regarding FIGS. 9 to 12 of the embodiment, it should also be noted that it is possible to exclude intra-predicted partitions from the candidate list, that is, their coding parameters are naturally not included in the candidate list.

[0117] Furthermore, it should be noted that three contexts can be used for merge_flag and merge_index.

[0118] Although several aspects have been described in relation to an apparatus, it will be apparent that these aspects also indicate a description of a corresponding method. Here, a block or device corresponds to a method step or the function of a method step. Similarly, aspects described in relation to a method step also indicate a description of a corresponding block or item or function of a corresponding apparatus. Some or all of the method steps can be performed (or by using a hardware device) by a hardware device such as, for example, a microprocessor, a programmable computer, or an electronic circuit. In some embodiments, some or one or more of the most important method steps can be performed by such a device.

[0119] Depending on specific implementation requirements, embodiments of the present invention can be implemented in hardware or in software. This embodiment can be performed using a digital storage medium having electronically readable control signals stored thereon that cooperate (or can cooperate) with a programmable computer system so that each method is performed, such as a floppy (registered trademark) disk, DVD, Blue?Ray, CD, ROM, PROM, EPROM, EEPROM or FLASH memory. Thus, the digital storage medium may be computer-readable.

[0120] Some embodiments according to the present invention include a data carrier having electronically readable control signals that can cooperate with a programmable computer system so that one of the methods described herein is performed.

[0121] Generally, embodiments of the present invention can be implemented as a computer program product having program code. And when the computer program product operates on a computer, the program code operates to execute one of the methods. The program code can be stored, for example, on a machine-readable carrier.

[0122] Other embodiments include a computer program stored on a machine-readable carrier for performing one of the methods described herein.

[0123] Thus, in other words, an embodiment of the method of the present invention is a computer program having program code for performing one of the methods described herein when the computer program runs on a computer.

[0124] Thus, a further embodiment of the method of the present invention is a data carrier (or digital storage medium or computer-readable medium) containing a computer program for performing one of the methods described herein recorded thereon. The data carrier, digital storage medium or recording medium is generally tangible and / or non-transitory.

[0125] Thus, a further embodiment of the method of the present invention is a sequence of a data stream or signal representing a computer program for performing one of the methods described herein. The sequence of the data stream or signal can be configured to be transferred via a data communication connection, for example via the Internet.

[0126] Further embodiments include processing means, such as a computer or a programmable logic device, configured or adapted to perform one of the methods described herein.

[0127] Further embodiments include a computer having installed thereon a computer program for performing one of the methods described herein.

[0128] A further embodiment according to the present invention includes an apparatus or system configured to transfer (e.g., electronically or optically) a computer program for performing one of the methods described in the present specification to a receiver. The receiver may be, for example, a computer, a mobile device, a storage device, etc. The apparatus or system may include, for example, a file server for transferring the computer program to the receiver.

[0129] In some embodiments, a programmable logic device (e.g., a field programmable gate array) can be used to perform some or all of the functions of the methods described in the present specification. In some embodiments, a field programmable gate array can cooperate with a microprocessor to perform one of the methods described in the present specification. Usually, the method is preferably executed by any hardware device.

[0130] The above-described embodiments are merely illustrative for the principles of the present invention. It is understood that modifications and changes to the apparatus and details described in the present specification will be apparent to other persons skilled in the art. Therefore, it is intended to be limited only by the scope of the impending claims and not by the specific details shown by the description and illustration of the embodiments in the present specification.

Claims

1. A decoder (80) configured to decode a bitstream (30) that signals one of the supported segmentation patterns for the current block (40) of an image (20), wherein the decoder is when the signaled one of the supported segmentation patterns specifies subdividing the current block (40) into two or more further block subdivisions (50, 60), for each block subdivision of the current block other than the first block subdivision in the encoding order (70), configured to determine a set of coding parameter candidates for each block of the current block, where at least some of the coding parameter candidates are adopted from the coding parameters of only one previously decoded block subdivision, and the at least some coding parameter candidates are equal to the coding parameters of the only one previously decoded block subdivision, and the decoder is also configured to perform the determination such that a coding parameter candidate equal to the coding parameter associated with any of the block subdivisions that becomes one of the supported segmentation patterns when integrated with each block subdivision (60) is excluded from the set of coding parameter candidates for each block subdivision (60), the determination of the set of coding parameter candidates for each block subdivision (60) of the current block is from a combination of the coding parameters of two or more previously decoded block subdivisions, or by modification, from the coding parameters of one previously decoded subdivision further comprising deriving at least some further coding parameter candidates, characterized in that it is a decoder.

2. The decoder (80) according to claim 1, wherein when the number of the coding parameter candidates not excluded is not zero, the decoder is configured to set the coding parameter associated with each block subdivision (60) depending on one of the coding parameter candidates not excluded.

3. The decoder (80) according to claim 1, wherein when the number of the coding parameter candidates not excluded is not zero, the decoder is configured to set the coding parameter associated with each block subdivision equal to the coding parameter candidate not excluded. Claim 4 The decoder supports an intra prediction mode and an inter prediction mode for the current block, and only when the current block is encoded in the inter prediction mode, the integration and the determination of the set of the encoding parameter candidates are such that an encoding parameter candidate equal to an encoding parameter associated with any of the block partitions that will become one of the supported partition patterns when integrated with each block partition (60) is excluded from the set of the encoding parameter candidates for each block partition (60), The decoder according to claim 2 or claim 3, configured to perform. Claim 5 The encoding parameter is a prediction parameter, and the decoder is configured to use the prediction parameter of each block partition (60) to derive a prediction signal for each block partition (60). The decoder according to any one of claims 2 to 4. Claim 6 When the number of non-excluded encoding parameter candidates for each block partition is greater than 1, the decoder simply expects in the bitstream (30) to include a syntax element that defines which of the non-excluded encoding parameter candidates will be used for integration depending on the number of the non-excluded encoding parameter candidates for each block partition. The decoder according to any one of claims 1 to 5, configured as such. Claim 7 The decoder (80) is configured to determine a set of the encoding parameter candidates for each block partition based at least in part on the encoding parameters associated with previously decoded block partitions that are adjacent to each block partition and located inside and outside the current block. The decoder according to any one of claims 1 to 6. Claim 8 The decoder (80) is configured to determine a set of the encoding parameter candidates for each block partition from those other than those encoded in the intra prediction mode among the first set of previously decoded blocks. The decoder according to any one of claims 1 to 7. Claim 9 The decoder (80) is configured to divide the image (20) into encoding blocks according to the segmentation information included in the bit stream (30), and the encoding block includes the current block. The decoder according to any one of claims 1 to 8.

10. The decoder (80) is further configured to further divide the current block (40) into one or more transformation blocks according to further segmentation information included in the bit stream, and derive the residual signal of the current block (40) from the bit stream (30) in units of the transformation blocks. The decoder according to claim 9.

11. A depth map is associated with the image as additional information. The decoder according to any one of claims 1 to 10.

12. An encoder (10) configured to encode an image (20) into a bit stream (30), the encoder is configured to signal in the bit stream (30) one of the supported partitioning patterns for the current block (40), when the one signaled among the supported partitioning patterns defines dividing the current block (40) into two or more block partitions (50, 60), for each block partition other than the first block partition in the encoding order (70) among the block partitions of the current block, for each block partition (60) of the current block, it is configured to determine a set of encoding parameter candidates, where at least some of the encoding parameter candidates are adopted from the encoding parameters of only one previously encoded block partition, and the at least some of the encoding parameter candidates are equal to the encoding parameters of the only one previously encoded block partition, and the encoder is configured to perform the determination such that an encoding parameter candidate equal to the encoding parameter associated with any of the block partitions that will become one of the supported partitioning patterns when integrated with each block partition is excluded from the set of encoding parameter candidates for each block partition (60), the determination of the set of encoding parameter candidates for each block partition (60) of the current block from a combination of the encoding parameters of two or more previously encoded block partitions, or from the encoding parameters of one previously encoded partition by modification further comprising deriving at least some additional encoding parameter candidates Encoder.

13. The encoder according to claim 12, wherein a depth map is associated with the image as additional information.

14. A method for decoding a bitstream (30) that signals one of the supported partition patterns for the current block (40) of an image (20), the method comprising: When the signaled one of the supported partition patterns defines subdividing the current block (40) into two or more block partitions (50, 60), For each block partition other than the first block partition in the encoding order (70) of the block partitions of the current block, For each block partition (60) of the current block, a step of determining a set of encoding parameter candidates, where At least some of the encoding parameter candidates are adopted from the encoding parameters of only one previously decoded block partition, and the at least some of the encoding parameter candidates are equal to the encoding parameters of the only one previously decoded block partition, step; Performing the determination such that an encoding parameter candidate equal to the encoding parameter associated with any of the block partitions that becomes one of the supported partition patterns when integrated with each block partition (60) is excluded from the set of encoding parameter candidates for each block partition (60); including The determination of the set of encoding parameter candidates for each block partition (60) of the current block is from a combination of the encoding parameters of two or more previously decoded block partitions, or from the encoding parameters of one previously decoded partition by modification further comprising deriving at least some additional encoding parameter candidates Method.

15. A method for encoding an image (20) into a bitstream (30), the method comprising: signaling, within a bitstream (30), one of the supported split patterns for a current block (40); when the signaled one of the supported split patterns defines subdividing the current block (40) into two or more block splits (50, 60); for each block split of the block splits of the current block, other than the first block split in the encoding order (70); determining, for each block split (60) of the current block, a set of candidate encoding parameters, wherein at least some of the candidate encoding parameters are adopted from the encoding parameters of only one previously encoded block split and are equal to the encoding parameters of the only one previously encoded block split; performing the determining such that any candidate encoding parameter equal to an encoding parameter associated with one of the block splits that, when integrated with each block split, results in one of the supported split patterns is excluded from the set of candidate encoding parameters for each block split (60); comprising the determining of the set of candidate encoding parameters for each block split (60) of the current block is from a combination of the encoding parameters of two or more previously decoded block splits, or by modification, from the encoding parameters of one previously encoded split further comprising deriving at least some additional candidate encoding parameters; A method. Claims 16 A computer program having program code for performing the method according to claim 14 or claim 15 when the computer program is executed on a computer.

Citation Information

Patent Citations

  • Methods and apparatus for context-dependent merging in skip / direct mode for video encoding and decoding.

    JP2010524397A

  • Coding method, decoding method, coding apparatus, decoding apparatus, image processing system, coding program, and decoding program

    WO2003026315A1

  • Method and apparatus for context dependent merging for SKIP-direct modes for video encoding and decoding

    WO2008127597A2

  • Video and depth coding

    WO2009091383A2