Inter-prediction concepts using tile independence constraints
By employing decoder-side tile boundary recognition and enforcement, the solution ensures tile-independent video coding maintains encoding efficiency by preventing motion information from crossing tile boundaries, addressing inefficiencies in existing technologies.
Patent Information
- Application Number
- JP2023112911
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2018-11-26
- Filing Date
- 2023-07-10
- Publication Date
- 2025-07-31
- Estimated Expiration
- 2039-11-25
AI Technical Summary
Existing video coding technologies face challenges in maintaining encoding efficiency while achieving tile-independent coding, as tile splitting often disrupts encoding dependencies, leading to inefficiencies at tile boundaries.
The decoder recognizes tile boundaries and enforces tile independence constraints by mapping or redirecting motion information to ensure that motion vectors and prediction residuals do not cross tile boundaries, using irreversible mappings and clipping techniques to maintain encoding efficiency.
This approach maintains encoding efficiency by minimizing bit rate and reducing coding penalties at tile boundaries, ensuring that motion information is within the current tile, thus preserving the integrity of tile-independent coding.
Smart Images

Figure 0007716447000001 
Figure 0007716447000002 
Figure 0007716447000003
Abstract
Description
[Technical Field]
[0001] This application relates to inter-coding concepts for use in block-based codecs, such as hybrid video codecs, and in particular to tile-based coding, i.e. concepts that allow independent coding of tiles into which an image is spatially subdivided. [Background technology]
[0002] Existing applications, such as 360° video services based on the MPEG OMAF standard, rely heavily on spatial video division or segmentation techniques. In such applications, spatial video segments are transmitted to a client and decoded jointly to adapt to the client's current viewing direction. Another related application that relies on spatial segmentation of video planes is the parallelization of encoding and decoding operations, for example, to leverage the multi-core capabilities of modern computing platforms.
[0003] One such spatial segmentation technique, known as tiling, is implemented in HEVC and divides the image plane into segments that form a rectangular grid. The resulting spatial segments are coded independently for entropy coding and intra prediction. Furthermore, there are means to demonstrate that spatial segments are also coded independently for state-of-the-art inter prediction. In some applications, such as those mentioned above, it is important to have constraints in all three areas: entropy coding, intra prediction, and inter prediction.
[0004] However, the ever-evolving technology of video coding brings about new coding tools, many of which are related to the field of inter-prediction, i.e., many tools involve new dependencies on previously coded images or on different areas within the image currently being coded, so that proper attention must be paid to how to ensure independence in all these areas.
[0005] Heretofore, it should be noted that it is the encoder that sets the encoding parameters using the above available encoding tools so that the independent encoding of the tiles of the video is observed. The decoder "relies" on each guarantee signaled to the decoder via the bitstream by the encoder.
[0006] There should be value in a concept that, by realizing tile-independent encoding, suppresses a decrease in encoding efficiency due to the breakdown of the encoding dependencies caused by tile splitting, while only slightly changing the operation of the codec along the boundaries. SUMMARY OF THE INVENTION PROBLEMS TO BE SOLVED BY THE INVENTION
[0007] Therefore, it would be preferable to have a concept that enables video encoding in such a way that tile-independent coding is achieved, but with a method that reduces the loss of encoding efficiency usually associated with tile-dependency disruptions, while only slightly changing the operation of the codec along the tile boundaries. MEANS FOR SOLVING THE PROBLEM
[0008] This object is achieved by the subject matter of the independent claims of the present application.
[0009] Generally speaking, the findings of the present application are that the obligation to comply with tile-independent coding of video is partially inherited among coders, or in other words, if the decoder provides partial Co-Attention to this so that the coder can utilize Co-Attention, a more efficient method for enabling tile-independent coding of video material becomes possible. More precisely, according to an embodiment of the present application, the decoder is equipped with tile boundary recognition. That is, the decoder operates depending on the position of the boundary between tiles into which the video is spatially divided. In particular, this tile boundary recognition is also related to the derivation of motion information from the data stream by the decoder. This "recognition" enables the decoder to recognize the signaled motion information transmitted in the data stream that would lead to violating the tile independence requirement when applied as signaled, and thus, when such signaled motion information that would violate tile independence is used for inter prediction, map it to an acceptable motion information state corresponding to motion information that does not lead to violating tile independence. The coder may rely on this operation, that is, recognize the redundancy of the signable motion information state resulting from the decoder's recognition, in particular, the compliance or enforcement of the tile independence constraint by the decoder. In particular, the coder can utilize the enforcement / compliance of the tile independence constraint of the decoder to select, from among the signable motion information states that lead to the same motion information on the decoder side due to the operation of the decoder, a state that requires a lower bit rate, for example, a state associated with the signaled motion information prediction residual being zero.Thus, on the encoder side, the motion information of a particular inter-predicted block is determined such that the patch from which the particular inter-predicted block is to be predicted is within the boundary of the tile in which the particular inter-predicted block is contained, i.e., the tile within which the particular inter-predicted block is located, and obeys the constraint of not crossing the tile boundary; however, when encoding the motion information of the particular inter-predicted block into a data stream, the encoder takes advantage of the fact that the derivation from the data stream will be dependent on the tile boundary, i.e., requires the above-mentioned tile boundary recognition.
[0010] According to embodiments of the present application, the motion information includes the motion vectors of specific inter-prediction blocks, and the tile boundary recognition in the decoder processing is related to the motion vectors. In particular, according to embodiments of the present application, the decoder enforces tile independence constraints regarding the motion vectors that are predictively encoded. That is, according to these embodiments, when determining the motion vectors, the decoder, on the one hand, based on the motion vector predictor / prediction on one side and, on the other hand, the motion information prediction regarding the residual transmitted in the data stream for the inter-prediction block, ensures that the patch from which the inter-prediction block is predicted does not cross the boundary of the tile in which the inter-prediction block is located. That is, the decoder performs the above compliance / enforcement by using an irreversible mapping, for example, instead of mapping all combinations of motion information prediction and motion information prediction residual to their sum, to generate the finally used motion vectors. This mapping redirects all possible combinations of motion vector prediction and motion vector prediction residual whose sum results in a motion vector associated with a patch that crosses the current tile boundary, i.e., the boundary of the tile in which the current inter-prediction block is located, to the motion vectors for which the associated patch does not cross the current tile boundary. As a result, the encoder can utilize the ambiguity in signaling a specific motion vector for a given inter-prediction block. For example, it can select the signaling of the motion vector prediction residual that results in the lowest bitrate for this inter-prediction block. This may be, for example, a difference of zero motion vectors.
[0011] According to the above-mentioned variation of the idea of providing the decoder with a function that can at least partially employ tile boundary recognition to enforce tile independence constraints that are usually the responsibility of only the encoder, the decoder applies the enforcement of tile independence constraints not to the motion information obtained as a result of a combination of motion information prediction and motion information prediction residue, but to one or more motion information predictors of a specific inter-prediction block. Both the encoder and the decoder enforce tile independence for one or more motion information predictors such that both use the same one or more motion information predictors. The ambiguity of signaling and the possibility of minimizing the bit rate using the decoder are not issues here. However, by creating in advance one or more motion information predictors for a specific inter-prediction block, i.e., before using one or more motion information predictors for motion information prediction coding / decoding, instead of wasting one or more motion information predictors that indicate patch positions beyond the current tile boundary where non-zero motion information prediction residue signaling is required to predictively encode motion information for the inter-prediction block to redirect conflicting motion information predictors to patch positions inside the current tile, it becomes possible to adjust or "concentrate" the available one or more motion vector predictors for a specific inter-prediction block to indicate only patch positions within the current tile. Again, the motion information may be a motion vector here.
[0012] Although related to the latter modification, and yet different from that modification, a further embodiment of the present application aims to avoid motion information prediction candidates for certain inter-prediction blocks whose direct, i.e., zero motion information prediction residual, application leads to violating the tile independence constraint. That is, in addition to the aforementioned modification, such motion information prediction candidates are simply not used for loading into the motion information prediction candidate list of the currently predicted inter-prediction block. The encoder and decoder operate in the same way. No redirection is performed. These predictors are simply aborted. That is, the establishment of the motion information prediction candidate list will be performed by the same tile boundary recognition method. In this way, all members that can be signaled for the currently encoded inter-prediction block are concentrated on non-conflicting motion information prediction candidates. Thereby, a complete list becomes signalable at points in the data stream, e.g., the signaling possible state of such pointers for motion information prediction candidates that are prohibited from being signaled to comply with the tile independence constraint when the indicated motion information prediction candidate conflicts with either this constraint or the motion information prediction candidate at the previous rank position becomes "useless" without waste.
[0013] Similarly, a further embodiment of the present application aims to avoid loading candidates into the motion information prediction candidate list whose origin is located outside the block where the current tile, i.e., the current inter-prediction block, is located. Thus, according to these embodiments, the decoder and the encoder check whether the inter-prediction block is adjacent to a predetermined side, such as the lower side and / or the right side, of the current tile. If so, the first block in the motion information reference image is identified, and the motion information prediction candidate derived from the motion information of this first block is loaded into the list. If not, the second block in the motion information reference image is identified, and instead, the motion information prediction candidate derived from the motion information of the second block is loaded into the motion information prediction candidate list. For example, the first block may be a block that inherits the position at the same location as the first alignment position inside the current inter-prediction block, while the second block is a block that is outside the inter-prediction block, i.e., offset from the inter-prediction block along the direction perpendicular to the predetermined side and located at the same location as the second predetermined position.
[0014] According to a further variation of the above idea of the present application, the construction / establishment of the motion information prediction candidate list is performed in such a way that motion information prediction candidates whose origin is likely to conflict with the tile independence constraint, i.e., whose origin is outside the current tile, are shifted towards the end of the motion information prediction candidate list. In this method, the signaling of pointers to the motion information prediction candidate list on the coder side is not overly restricted. In other words, the pointer that signals the motion information prediction candidate to be actually used for the current inter-prediction block, which is transmitted for a specific inter-prediction block, indicates this motion information candidate actually used according to the rank position within the motion information prediction candidate list. By shifting the motion information prediction candidates in the list that may not be available because their origin is located outside the current tile to the end, or at least towards the back, i.e., to occur at a higher rank, all the motion information prediction candidates preceding them in the list are still signalable by the coder, and thus the range of those motion information prediction candidates available for predicting the motion information of the current block is still large compared to the case where such "problematic" motion information prediction candidates are not shifted towards the end of the list. According to some embodiments related to the above aspect, the reading of the motion information prediction candidate list in such a way that "problematic" motion information prediction candidates are shifted towards the end of the list is performed by a tile boundary recognition method in the coder and decoder. In this way, the slight coding efficiency penalty associated with shifting this potentially more affected motion information prediction candidate towards the end of the list is restricted to the area of the image of the video along the tile boundary. However, according to an alternative, the shift of the "problematic" motion information prediction candidates towards the end of the list is performed regardless of whether the current block is along the tile boundary or not. Although slightly reducing the coding efficiency, the latter alternative may improve robustness and simplify the coding / decoding procedure. The "problematic" motion information prediction candidates may be candidates derived from blocks in the reference image or candidates derived from motion information history management.
[0015] According to a further embodiment of the present application, the decoder and the encoder further use the predicted motion vectors, whose motion information is used to form temporal motion information prediction candidates in a list, to indicate blocks in the motion information reference picture, by enforcing tile independence constraints regarding the predicted motion vectors, to determine the temporal motion information prediction candidates in a tile boundary recognition method. The predicted motion vectors are clipped to stay within the current tile or to indicate a position within the current tile. The availability of such candidates is thus guaranteed within the current tile such that compliance with the tile independence constraints of the list is maintained. According to an alternative concept, instead of clipping the motion vector, a second motion vector is used when the first motion vector points outside the current tile. That is, the second motion vector is used to place the block for which the temporal motion information prediction candidate is formed based on the motion information when the first motion vector points outside the tile.
[0016] According to yet another embodiment, in the above idea of providing tile boundary recognition to the decoder to assist in enforcing tile independence constraints, the decoder supports motion compensation prediction followed when motion information is encoded in the data stream for a particular inter prediction block, and the decoder derives a motion vector for each sub-block of the sub-blocks into which this inter prediction block is divided, and depending on the position of the boundary between tiles, performs either the derivation of the sub-block motion vectors, or the prediction of each sub-block using the derived motion vectors, or both. In this way, cases where such an effective coding mode cannot be used by the encoder because it would conflict with the tile independence constraints are considerably reduced.
[0017] According to a further aspect of the present application, a codec that supports motion-compensated bidirectional prediction and includes a bidirectional optical flow tool in an encoder and a decoder to improve motion-compensated bidirectional prediction complies with tile-independent coding by providing automatic deactivation of the bidirectional optical flow tool in the encoder and decoder when the application of the tool leads to a conflict with the tile independence constraint, or by using boundary padding to determine a region of a patch that is predicted using the bidirectional optical flow tool from outside the current tile from which a particular bidirectional predictive inter-prediction block is predicted. Another aspect of the present application relates to the loading of a motion information predictor candidate list that uses a motion information history list for storing previously used motion information. This aspect can be used regardless of whether tile-based coding is used or not. This aspect attempts to provide a video codec with higher compression efficiency by giving a selection of motion information predictor candidates from among the motion information history lists in which the currently loaded entry in the candidate list will be filled depending on the motion information predictor candidates that have been loaded into the motion information predictor candidate list so far. The purpose of this dependency is to select motion information entries in the history list that are likely to be further away from the motion information predictor candidates that have been loaded into the candidate list so far. A suitable distance metric can be defined, for example, based on the motion vectors respectively constituted by the motion information entry in the history list and the motion information predictor candidate in the candidate list, and / or the reference picture index constituted by them. By this method, the loading of the candidate list using history-based candidates results in a more advanced "update" of the resulting candidate list such that the likelihood that the coder finds a suitable candidate within the candidate list from the perspective of excellent distortion optimization is higher compared to the case where history-based candidates are simply selected according to the rank in the motion information history list, i.e., how newly the candidate was entered into the motion information history list. This concept can actually be applied to any other motion information predictor candidate that is currently to be selected from among the set of motion information predictor candidates for loading into the motion information predictor candidate list.
[0018] Advantageous aspects of the present invention are the subject matter of the dependent claims. Preferred embodiments of the present application will be described below with reference to the figures.
Brief Description of the Drawings
[0019]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12a
Figure 12b
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
Figure 20
Figure 21
Figure 22
[0020] In the following description of the figures, first, an encoder and a decoder of a block - based predictive codec for encoding an image of a video are presented to form an example of an encoding framework in which an embodiment of an inter - prediction codec can be incorporated. The aforementioned encoder and decoder are described in relation to FIGS. 1 to 3. Thereafter, an embodiment of the inter - prediction concept of the present application is presented. Those concepts can be combined or used as a combination. In particular, all the concepts described hereinafter may be incorporated in combination into the encoder and decoder of FIGS. 1 and 2 respectively, and the embodiments described in FIGS. 4 and subsequent figures may also be used to form encoders and decoders that do not operate according to the encoding framework underlying the encoders and decoders of FIGS. 1 and 2.
[0021] FIG. 1 exemplarily shows an apparatus for predictively encoding a video 11 consisting of an image sequence 12 into a data stream 14 using transform-based residual coding. This apparatus, i.e., the encoder, is designated using reference numeral 10. FIG. 2 shows a corresponding decoder 20, i.e., apparatus 20, configured to predictively decode a video 11' consisting of an image sequence 12' from the data stream 14, also using transform-based residual decoding. The apostrophe is used to indicate that the video 11' and image 12' reconstructed by the decoder 20 deviate from the image 12 initially encoded by the apparatus 10 with respect to the coding loss introduced by quantization of the prediction residual signal. FIGS. 1 and 2 exemplarily use transform-based prediction residual coding, but embodiments of the present application are not limited to this type of prediction residual coding. This also applies to other details described with respect to FIGS. 1 and 2, as will be explained below.
[0022] The encoder 10 is configured to perform a spatial-spectral transform on the prediction residual signal and encode the prediction residual signal thus obtained into the data stream 14. Similarly, the decoder 20 is also configured to decode the prediction residual signal from the data stream 14 and perform a spectral-spatial transform on the prediction residual signal thus obtained.
[0023] Internally, the coder 10 may include a prediction residual signal former 22 that generates a prediction residual 24 to measure the deviation of the original signal, i.e., the prediction signal 26 from the current image 12. The prediction residual signal former 22 may be, for example, a subtractor that subtracts the prediction signal from the original signal, i.e., the current image 12. In that case, the coder 10 further includes a converter 28 that performs a spatial-spectral conversion on the prediction residual signal 24 and then obtains a prediction residual signal 24' in the spectral domain that is quantized by a quantizer 32 also included in the coder 10. The prediction residual signal 24'' thus quantized is encoded into the bitstream 14. For this purpose, the coder 10 may optionally include an entropy coder 34 that entropy-encodes the converted and quantized prediction residual signal into the data stream 14. The prediction residual 24 is decoded from the data stream 14 and is generated by the prediction stage 36 of the coder 10 based on the prediction residual signal 24'' that is decodable from the data stream 14. For this purpose, the prediction stage 36 may internally include an inverse quantizer 38 that inverse-quantizes the prediction residual signal 24'' to obtain a prediction residual signal 24''' in the spectral domain corresponding to the signal 24' excluding quantization loss as shown in FIG. 1, followed by an inverse converter 40 that performs an inverse conversion, i.e., a spectral-spatial conversion, on the latter prediction residual signal 24''' to obtain a prediction residual signal 24'''' corresponding to the original prediction residual signal 24 excluding quantization loss. Then, a combiner 42 of the prediction stage 36 recombines the prediction signal 26 and the prediction residual signal 24'''' by addition or the like so as to obtain a reconstructed signal 46, i.e., a reconstruction of the original signal 12. The reconstructed signal 46 may correspond to the signal 12'. Then, a prediction module 44 of the prediction stage 36 generates the prediction signal 26 based on the signal 46 using, for example, spatial prediction, i.e., intra prediction, and / or temporal prediction, i.e., inter prediction.
[0024] Similarly, decoder 20 may be internally configured with components corresponding to prediction stage 36 and interconnected to correspond to prediction stage 36. In particular, entropy decoder 50 of decoder 20 may entropy decode quantized spectral domain prediction residual signal 24″ from the data stream, whereupon an interconnected inverse quantizer 52, an inverse transformer 54, a combiner 56, and a prediction module 58, operating in the manner described above with respect to the modules of prediction stage 36, recover a reconstructed signal based on prediction residual signal 24″, such that the output of combiner 56 is reconstructed signal, i.e., image 12′, as shown in FIG.
[0025] Although not specifically mentioned above, it is readily apparent that the encoder 10 may set some coding parameters, including, for example, prediction modes, motion parameters, etc., according to some optimization scheme, such as, for example, to optimize some rate- and distortion-related criteria, i.e., encoding cost. For example, the encoder 10 and decoder 20 and corresponding modules 44, 58 may each support different prediction modes, such as intra-coding and inter-coding modes. The granularity at which the encoder and decoder switch between these prediction mode types may correspond to the subdivision of the image 12 and the image 12′ into coding segments or coding blocks, respectively. In units of these coding segments, for example, the image may be subdivided into intra-coded blocks and inter-coded blocks. The intra-coded blocks are predicted based on the respective blocks' spatial, already coded / decoded neighbors. Several intra-coding modes may exist and may be selected for each intra-coded segment, including a directional intra-coding mode or an angular intra-coding mode, according to which each segment is filled by extrapolating neighboring sample values to the respective intra-coded segment along a specific direction specific to the directional intra-coding mode. Intra-coding modes may also include one or more additional modes, such as a DC coding mode in which the prediction of each intra-coded block follows to assign a DC value to all samples in each intra-coded segment, and / or a planar intra-coding mode in which the prediction of each block follows to approximate or determine a spatial distribution of sample values described by a two-dimensional linear function over the sample positions of each intra-coded block by deriving the slope and offset of the plane defined by the two-dimensional linear function based on neighboring samples. In comparison, inter-coded blocks may be, for example, temporally predicted. In inter-coded blocks, motion information may be signaled in the data stream.The motion information may include vectors indicating the spatial displacement of portions of previously coded images of the video to which image 12 belongs, where the previously coded / decoded images are sampled to obtain the prediction signal for each inter-coded block. More complex motion models may also be used. This means that in addition to the residual signal coding included in data stream 14, such as entropy-coded transform coefficient levels representing the quantized spectral-domain prediction residual signal 24", data stream 14 may also optionally encode further parameters therein, such as coding mode parameters for assigning coding modes to various blocks, prediction parameters of some of the blocks, such as motion parameters for inter-coded blocks, and parameters controlling and signaling the subdivision into blocks of each of images 12 and 12'. Decoder 20 uses these parameters to subdivide the image in the same way as encoder 10, assigning the same prediction modes to the blocks and making the same predictions to obtain the same prediction signals.
[0026] Figure 3 shows the relationship between one of the reconstruction signals, i.e., the reconstructed image 12’, and the combination of the other, the prediction residual signal 24’’’’ signaled in the data stream and the prediction signal 26. As already described above, this combination may be an addition. The prediction signal 26 is shown in Figure 3 as the subdivision of the image area into the intra-coded blocks illustrated using hatching and the inter-coded blocks illustrated without hatching. The subdivision can be any subdivision, such as the subdivision of the rows and columns of the blocks of the image area or the regular subdivision into blocks, or the multi-tree subdivision into leaf blocks of various sizes of the image 12 such as the quad-tree subdivision. A mixture of them is shown in Figure 3. The image area is first subdivided into the rows and columns of the tree root blocks and then they are further subdivided according to the recursive multi-tree subdivision. Also in this case, the data stream 14 may encode therein the intra-coding mode for each intra-coded block 80 that assigns one of several supported intra-coding modes to each intra-coded block 80. For the inter-coded block 82, the data stream 14 may encode therein motion information such as motion information including one or more motion vectors. Details will be described below. Generally speaking, the inter-coded block 82 is not limited to being temporally coded only. Alternatively, the inter-coded block 82 may be any block predicted from a previously coded part beyond the current image 12 itself, such as a previously coded image of the video to which the image 12 belongs, or, when the encoder and the decoder are respectively scalable encoders and decoders, an image of another view or a hierarchically lower layer. The prediction residual signal 24’’’’ in Figure 3 is also shown as the subdivision into the blocks 84 of the image area. These blocks may also be called transform blocks to distinguish these blocks from the coding blocks 80 and 82.Substantially, FIG. 3 shows that the encoder 10 and the decoder 20 can use two different subdivisions of the blocks, namely, one subdivision into each of the encoding blocks 80 and 82, and another subdivision into block 84, for the images 12 and 12' respectively. Both subdivisions may be the same, i.e., each of the encoding blocks 80 and 82 may form the conversion block 84 simultaneously. However, FIG. 3 shows a case where, for example, any boundary between the two blocks 80 and 82 overlaps the boundary between the two blocks 84, in other words, the subdivision into the conversion block 84 forms an extension of the subdivision into the encoding blocks 80 / 82 such that each block 80 / 82 coincides with one of the conversion blocks 84 or coincides with a cluster of the conversion blocks 84. However, the subdivisions may be determined or selected independently of each other such that the conversion block 84 can also cross the block boundary between the blocks 80 / 82 as an alternative. As far as the subdivision into the conversion block 84 is concerned, the same description as presented for the subdivision into the blocks 80 / 82 applies accordingly, i.e., the block 84 may be the result of a regular subdivision into blocks arranged in rows and columns of the block / image area, the result of a recursive multi-tree subdivision of the image area, or a combination thereof or any other kind of blocking. Incidentally, it should be noted that the blocks 80, 82, and 84 are not limited to being of secondary shape, rectangular, or any other shape.
[0027] In the embodiments described below, the inter-prediction block 104 is typically used to explain the specific details of each embodiment. This block 104 may be one of the inter-prediction blocks 82. Any other blocks referred to in the subsequent figures may be either of the blocks 80 and 82.
[0028] 3 shows that the reconstructed signal 12′ is directly obtained by combining the prediction signal 26 with the prediction residual signal 24′″. However, it should be noted that, according to alternative embodiments, multiple prediction signals 26 may be combined with the prediction residual signal 24′″ to result in the image 12′.
[0029] 3, the transform segments 84 have the following significance: The transformer 28 and the inverse transformer 54 perform their transforms in units of these transform segments 84. For example, many codecs use some kind of DST or DCT for all transform blocks 84. Some codecs allow for skipping the transform for some of the transform segments 84, so that the prediction residual signal is directly coded in the spatial domain. However, according to embodiments described below, the encoder 10 and the decoder 20 are configured so that they support multiple transforms. For example, the transforms supported by the encoder 10 and the decoder 20 may include: DCT-II (or DCT-III), where DCT stands for Discrete Cosine Transform DST-IV, DST stands for Discrete Sine Transform DCT-IV DST-VII Identity Transformation (IT)
[0030] Naturally, the transformer 28 will support all forward transform versions of these transforms, and the decoder 20 or inverse transformer 54 will support their corresponding backward or inverse versions. Inverse DCT-II (or Inverse DCT-III) Reverse DST-IV Inverse DCT-IV Reverse DST-VII Identity Transformation (IT)
[0031] In the following description, details are shown on how the encoder 10 and the decoder 20 can perform inter prediction. All the other modes described above, such as the intra prediction mode, can also be supported individually or in total additionally. Residual encoding can also be performed in different ways, such as in the spatial domain.
[0032] As already described above, FIGS. 1 to 3 are presented as examples of an encoder that performs block-based video encoding using one or more of the concepts detailed below. The details described below can be transferred to the encoder of FIG. 1 or to another block-based video encoder, such that the other block-based video encoder differs from the encoder of FIG. 1, for example, in that it does not support intra prediction, or in that the subdivision into blocks 80 and / or 82 is performed in a different way than illustrated in FIG. 3, or even in that this encoder does not use transform prediction residual encoding, for example, but instead directly encodes the prediction residual in the spatial domain. Similarly, the decoder according to the embodiments of the present application can also perform block-based video decoding from the data stream 14 using any inter prediction encoding concept further outlined below, but, for example, differs from the decoder 20 of FIG. 2 in that it does not support intra prediction, or in that the image 12' is subdivided into blocks in a different way than described with respect to FIG. 3, and / or in that the prediction residual from the data stream 14 is not derived in the transform domain, for example, but in the spatial domain.
[0033] However, before describing the embodiments, the ability to encode video in a tiled form and the method of encoding / decoding tiles independently of each other are described. Further, a detailed description of additional encoding tools that can be selectively used according to the embodiments described below and that nevertheless maintain the independence of the tile encoding from each other is described, which has not been discussed so far.
[0034] Thus, after describing potential implementations of a block-based video encoder and decoder, the concept of tile-independent coding is illustrated using FIG. 4. Tile-based independent coding may represent just one of multiple coding configurations for an encoder. That is, the encoder may be configured to adhere to tile-independent coding or may be configured to not order tile-independent coding. Of course, a video encoder may also unavoidably apply tile-independent coding.
[0035] According to tile-independent coding, the images of the video 11 are divided into tiles 100. In Fig. 4, only three images 12a, 12b, and 12c of the video 11 are exemplarily shown, each of which is exemplarily divided into six tiles 100, although division into any other number of tiles, such as two or more tiles, is possible. Furthermore, while Fig. 4 shows that the tile division into tiles 100 is done in such a way that the tiles 100 within an image are regularly arranged in rows and columns, with each tile 100 being a rectangular portion of the respective image and with the boundaries between the tiles 100 forming straight lines 102 passing through the images 12a-12c, it should be noted that alternative solutions exist in which the tiles have another shape. 3 into blocks for encoding and decoding each image, respectively, may be aligned with tile boundaries 102, and none of the coding, prediction and / or transformation blocks may cross tile boundaries 102, but rather may be exclusively in one of the tiles 100 of the respective image. For example, the tile division may be made to be aligned with the division into blocks into tree-root blocks of image 12, i.e., also representatively shown as those to which block 104 belongs, which are individually subdivided into blocks 80 / 82 by a recursive multi-tree division by the encoder, the individual subdivisions of which are signaled to the decoder in the data stream.
[0036] Tile-independent encoding means that, as exemplified for block 104 in FIG. 4, the encoding of each part of a tile is done in a way independent of other tiles in which block 104 is not located internally. This independence of encoding is not only related to other tiles of the same image, i.e., other tiles within the same image of which block 104 is a part, but the independence within the image is shown in FIG. 4 using exemplary arrow 106, which is shown as being interrupted to indicate that the encoding of block 104 does not depend on any other tile within the same image. The independence of encoding also relates to encoding dependencies regarding other images, and all images 12a - 12c of video 11 are divided into tiles 100 in the same way. Thus, each tile 100 of a particular image 12a has a tile at the same location corresponding to all other images. "Tile" thus describes not only the spatial segment of a particular image, such as segment of image 12a in FIG. 4 containing block 104, but also the spatio-temporal segment of video 11 consisting of all tiles at the same location in all images of the video that are at the same location as this tile within the particular image. In FIG. 4, this spatio-temporal segment is shown by a dashed line with respect to the tile containing block 104. When encoding block 104, the encoder will thus also limit the encoding dependencies with respect to images other than image 12a containing block 104 in such a way that no encoding dependencies are created with respect to tiles outside the spatio-temporal segment 108. In FIG. 4, this is shown by an interrupted arrow 110 pointing to another tile within reference image 12b from the tile containing block 104.
[0037] The concepts and embodiments further outlined below present the possibility of how to avoid encoding dependencies between different tiles of different images and how to ensure this between the encoder and the decoder, which would otherwise be caused, for example, by motion vectors pointing far away, leaving tiles at the same location in the reference image, and / or which would otherwise be caused by deriving motion information prediction from regions beyond the boundaries of the tile to which block 104 belongs.
[0038] However, before describing the embodiments of the present application, as an example that tends to cause tile interdependencies between different tiles of different images, which can be implemented in a block-based video encoder and a block-based video decoder according to an embodiment of the present application, and thus is likely to conflict with the constraints shown in FIG. 4 using the interrupted arrow 110, a specific encoding tool will be described.
[0039] Typical state-of-the-art video encoding makes extensive use of inter-prediction between frames to improve encoding efficiency. In the encoding of inter-prediction constraints, the inter-prediction process usually needs to be adapted to comply with spatial segment boundaries, which is achieved by encoding the significant motion information prediction residuals or MV differences per bitrate (regarding the available motion information predictors or MV predictors).
[0040] Typical state-of-the-art video encoding heavily relies on the concept of collecting motion vector predictors in so-called candidate lists at different stages of the encoding process, which also includes a mixed deformation of the MV predictor, for example, in the merge candidate list.
[0041] Typical candidates added to such a list are the MVs of blocks in the same location from temporal motion vector prediction (TMVP), which are adjacent to the same spatial position of the current encoding block but are the bottom-right blocks within the reference frame of most of the current blocks. The only exceptions are the blocks at the image boundaries and the blocks at the lower boundary of the CTU rows (at least in HEVC) where the blocks in the actual same location are used to derive so-called center same-location candidates because there is no bottom-right block within the image.
[0042] There are use cases where the use of such candidates can be a problem. For example, when tiles are used and partial decoding of the tiles is performed, their availability can change. This happens because the TMVP used in the encoder may belong to another tile that is not currently being decoded. This process also affects all candidates after the index of the candidates in the same location. Therefore, decoding such a bitstream can lead to an encoder / decoder mismatch.
[0043] Video coding standards such as AVC and HEVC rely on high-precision motion vectors that are more accurate than integer pixels. This means that the motion vector can refer to samples of the previous image at sub-sample positions rather than integer sample positions. Therefore, the video coding standard defines an interpolation process that derives the value of the sampling position of the sub-sample by an interpolation filter. In the case of HEVC, for example, an 8-tap interpolation filter (4 taps for chroma) is used for the luma component.
[0044] As a result of the sub-pixel interpolation process used in many video coding specifications, when using a motion vector MV2 that indicates a non-integer sample position with a length (horizontal or vertical component) shorter than the length of the motion vector MV1 that indicates an integer sample position in the same direction, samples of neighboring tiles are not required for the use of the MV1 motion vector, but due to the sub-sample interpolation process, it can happen that samples reaching further away from the origin (e.g., of neighboring tiles) are required.
[0045] Conventionally, video coding has used a motion model with only translation, i.e., a rectangular block is displaced according to a two-dimensional motion vector to form a motion compensation predictor for the current block. Such a model cannot represent rotational or zoom motions that are common in many video sequences. Therefore, efforts have been made to slightly extend the conventional translational motion model, such as what is called an affine motion in JEM. As shown in FIG. 5, in this affine motion model, two (or three) so-called control point motion vectors (v0, v1) are used for each block 104 to model rotational and zoom motions. The control point motion vectors v0, v1 are selected from among the MV candidates of adjacent blocks. The resulting motion compensation predictor is derived by calculating a sub-block based motion vector field and depends on the conventional translational motion compensation of the rectangular sub-blocks 112. FIG. 5 exemplarily shows a predictor or patch 114 (dashed line) obtained as a result of predicting the upper right sub-block (solid line) according to the respective sub-block motion vectors therefrom.
[0046] Motion vector prediction is a technique for deriving motion information from temporally or spatially-temporally correlated blocks. In one such procedure called Alternative TMVP (ATMVP) as shown in FIG. 6, a default MV candidate (e.g., the first merge candidate in the merge candidate list) from a list of blocks 104 in the current image 12a called the temporal vector 116 is used to determine the correlated block 118 in the reference frame 12b.
[0047] In other words, for the current prediction block 104 within the CTU, a prediction block 118 at the same location in the reference image 12b is determined. To place the block 118 at the same location, the motion vector MVTemp 116 is selected from among the spatial candidates in the current image 12a. The block at the same location is determined by adding the motion vector MVTemp 116 to the position 12a of the current prediction block.
[0048] For each sub-block 122 of the current block 104, the motion information of each sub-block 124 of the correlation block 118 is then used for the inter-prediction of the current sub-block 122, and the inter-prediction of the current sub-block 122 also incorporates additional procedures such as MV scaling of the time difference between the involved frames.
[0049] Bi-directional optical flow (BIO) uses sample-by-sample motion refinement that is performed in addition to motion compensation for each block of bi-directional prediction. The procedure is shown in FIG. 7. The refinement is based on minimizing the difference between image patches (A, B) in two reference frames 12b, 12c (Ref0, Ref1) based on the optical flow theorem that is based on the horizontal and vertical gradients of each image patch. The horizontal and vertical gradient images of each image patch in A and B, i.e., the measure of the directional change of the luma intensity in the image, are calculated to derive the enhanced offset for each sample of the prediction block.
[0050] A non-redundant FiFo buffer of the last used MVs is maintained to fill the MV candidate list with candidates more promising than the zero motion vector, which is called history-based motion vector prediction (HMVP). A new concept is introduced into current state-of-the-art video coders such as the test model of VVC. The state-of-the-art coder adds the HMVP candidates to the motion vector candidate list after the candidates at the same location (per sub-block or block) or spatial candidates.
[0051] The general functions of the block-based video coder and decoder according to the examples described above with respect to FIGS. 1 to 3 are explained, and the purpose of tile-independent coding, as well as several coding tools whose use is particularly easy to scale down or becomes very ineffective when tile-independent coding is not specifically taken into account individually, are explained as a whole. In the following, embodiments will be described that comply with tile independence and yet enable the very effective use of the coding tools as exemplified above.
[0052] FIG. 8 illustrates a block-based video encoder 10 and a block-based video decoder 20 according to an embodiment of the present application; as already mentioned above, the reuse of reference numbers from FIGS. 1 and 2 does not imply that the encoder 10 and decoder 20 of FIG. 8 must be interpreted as described above with respect to FIGS. 1 and 2, although this may or may not be the case. The encoder 10 encodes an image 11 into a data stream 14, and the decoder 20 decodes a reconstruction of the image 11 from the data stream 14. The images 12 of the image 11 are divided into tiles 100 as described with respect to FIG. 4. Coding independence between the tiles 100 is not, according to FIG. 8, a task that is only taken into account by the encoder 10 when encoding the image 11 into the data stream 14. Rather, the decoder 20 also recognizes tile boundaries 102 that run through the image 12 of the image 11 and separate neighboring tiles 100 from each other. In particular, as already described above with respect to FIG. 4, the tiles 100 also extend in the temporal dimension because each image 12 is divided into tiles 100 in the same way, such that each tile of one image has a tile that is co-located in any other image of the video 11. Thus, motion information used to predict an inter-predicted block in one image 12 should remain within the current tile, i.e., within the same tile of the respective inter-predicted block. In other words, the motion information of such an inter-predicted block of the current image, which positions a patch in the reference image from which the respective inter-predicted block is to be predicted, should not cause inter-tile dependencies. According to FIG. 8, the decoder alleviates, at least to some extent, the situation of the encoder 10, by deriving motion information of a given inter-predicted block 104 of a current image 12a of the video 11, which positions a patch 130 in the reference image 12b of the video from which the given inter-predicted block 104 is to be predicted from the data stream 14 such that the position of the boundary 102 in the image 12 of the video 11 is dependent. The motion information so derived may be defined as or may include, for example, motion vectors 132. Encoder 10 may further rely on this decoder operation and may operate as follows.Still, the encoder 10 determines the motion information 132 of a given inter-prediction block of the current picture 12a in such a way that the patch 130 remains inside and the inter-prediction block 104 does not cross the boundary 102 of the tile in which it is located. More precisely, the encoder 10 determines the motion information 132 such that the patch is within the boundary of the tile at the same location in the reference picture 12b, i.e., the same tile in which the block 104 is located. The encoder 10 determines the motion information 132 in this way and uses this motion information for the inter-prediction of the block 104. However, when encoding the motion information into the data stream 14, the encoder utilizes the function of the decoder 40, more precisely, the recognition of the decoder 20 regarding the boundary 102. Thus, the encoder 10 encodes the motion information 132 into the data stream 14 such that the derivation of the motion information 132 from the data stream 14 is performed depending on the position of the boundary 102 between the tiles 100, i.e., requires a dependency on the position of the boundary 102 between the tiles 100.
[0053] According to the embodiment described next, the decoder 20 uses the aforementioned tile boundary recognition, i.e., the dependence on the derivation of motion information from the data stream for the boundary position of the tile, to enforce tile independence regarding the signal state of the motion information for a given inter-prediction block 104. Refer to FIG. 9. The data stream 14 includes a portion or syntax portion 140 that signals the motion information of block 104. This portion 140 thus determines the signaling state of the motion information of block 104. Later, it will be explained that portion 140 may include, for example, the difference of motion vectors, i.e., the motion vector prediction residual, and optionally, a motion predictor index indicating a list of motion predictors. A reference picture index identifying the reference picture 12b to which the motion vector is related may also optionally be included in portion 140. However precisely the motion information portion 140 is described, as long as the signaling state of the motion information of block 104 leads to the corresponding patch 130 within the same tile 100a from which block 104 is predicted, this signaling state will not conflict with the tile independence constraint, and thus, the decoder 20 may leave this state as it is and use this state as the final state of the motion information, i.e., the state used for the inter-prediction of block 104 from patch 130. Two exemplary signaling states are shown in FIG. 9. The state indicated by 1 enclosed in a circle is associated with motion information 132a that clearly leads to patch 130a inside tile 100a. The other state indicated by 2 enclosed in a circle is associated with motion information 132b and leads to patch 130b that crosses boundary 102 from the current tile 100a to neighboring tile 100b. The overlapping portion of patch 130b that overlaps with neighboring tile 100b is shown hatched in FIG. 9. The predicted block 104 from patch 130b will introduce inter-tile dependence because samples of neighboring tile 100b are used for the prediction of block 104.Thus, according to the example of Figure 9, decoder 20 checks whether the signaling state of portion 140 is in a state like state 2 of Figure 9, where the patch is at least partially outside tile 100a, and if so, redirects the motion information from the signaling state to a redirection state that leads patch 130 to stay within boundary 102 of current tile 100a. The dashed-dotted line of Figure 9 shows patch 130'b positioned by redirection state 2' of redirected motion information 130b'.
[0054] There are various ways in which such redirection can be performed, as will be described in more detail below. In particular, as will become clear from the following description, patches 130a, 130b, and 130'b may be of different sizes. At least some of them may be of different sizes compared to block 104. This may be caused, for example, by the need to use an interpolation filter to predict block 104 from a reference image in each patch, for example, in the case of corresponding motion information including sub-pixel motion vectors. As will become clear from the description of FIG. 9, the above redirection 142 gives the encoder two options for signaling motion information 130b' inside data stream 14: the encoder can use signal state 2 or state 2'. This "freedom" may lead to a decrease in bit rate. For example, assume that predictive coding is used to signal the motion information. If the predictor for the motion information of block 104 indicates state 2, the encoder can simply use this predictor to signal a zero motion information prediction residual in portion 140 of data stream 14, thereby signaling state 2, which will be redirected to state 2' by the decoder.
[0055] FIG. 10 shows the possibilities already described with respect to FIG. 9. The decoder may use the dependence of the motion information derivation from the data stream at the tile boundary positions and the result of its correction using the prediction of the motion information and the motion information prediction residual signal in the portion 140 of the data stream 14 of block 104 to comply with or enforce the constraints on the tile independence with respect to the signaling state of the motion information. That is, for block 104, the encoder and decoder will use previously encoded blocks, such as based on the motion information used to encode spatially neighboring blocks of block 104 in the current image or blocks located near block 104 in a particular reference image, to provide a motion information prediction, such as a vector, or a motion information predictor 150. The latter may be the reference image 12b including the patch 130 or another image used for motion information prediction. A motion information prediction residual 152, such as a difference in motion vectors, is signaled in the portion 140 of block 104. The decoder thus decodes the motion information prediction residual 152 from the data stream 14 and determines the motion information (such as a motion vector) to be used for block 104 based on the motion information prediction 150 and the motion information prediction residual 152. For example, in the case of a motion vector, the sum of the motion vectors 150 and 152 indicates a signaled motion vector 154 that conforms to the signaling state of the motion information of block 104, and the motion vector 154 is the footprint or patch 130 that is to be predicted therefrom by placing the foot of block 104 at a predetermined position, such as a position 156 in a reference image that is the same place as a predetermined alignment position, such as the upper left corner 158 as shown in FIG. 10, of the corner of block 104. The head of the vector 154 will then indicate, for example, the upper left corner of the footprint of block 104. The patch 130 may be enlarged or otherwise deformed compared to the pure footprint by filtering, such as interpolation filtering, or other techniques used to derive the prediction of block 104 from the patch 130, as described above. According to the embodiments described next, the prediction-based decoded motion information 154 is subject to the compliance / enforcement of the constraints.
[0056] That is, the decoder observes or enforces a constraint on the motion information 154 such that the patch 130 does not cross the boundary 102 of the tile 108 to which the block 104 belongs. As described with respect to FIG. 9, the decoder uses an irreversible mapping for this observance / enforcement. This mapping maps the possible signaling states, here the possible combinations of the motion information prediction 150 and the motion information prediction residual 152 of FIG. 10, to non-conflicting states of the motion information 154. The irreversible mapping maps each possible combination to the final motion information such that the patch 130 arranged according to such final motion information remains within the boundary 102 of the current tile 100a. As a result, depending on the position 104 with respect to the tile boundary of the current tile and / or the direction and size of the motion information prediction of 150, there are possible combinations of the motion information prediction 150 and the motion information prediction residual 152, i.e., possible signaling states, that are mapped to the same final motion information 154, i.e., the same final state, but these different combinations share, for example, the same motion information prediction 150 and have different residuals 152 internally. As a result, the encoder can freely select the motion information prediction residual 152. FIG. 10 shows, for example, that the motion information prediction residuals 152' can each lead to substantially the same motion information 154 due to the observance / enforcement of the decoder-side constraint, i.e., the redirection 142 of the above irreversible mapping.
[0057] In such cases where the encoder can freely select any motion information prediction residual from among those that lead to substantially the same motion information to be applied to the block 104 for performing inter prediction, the encoder 10 may use the one for which the motion information prediction residual is zero because this setting is likely to lead to the lowest bit rate.
[0058] FIG. 11 shows a patch 130 from which the block 104 is to be inter-predicted, where the patch 130 is arranged according to specific motion information such as a specific motion vector 132, indicating that it may have a different shape compared to the block 104. Usually, the patch 130 is larger than a mere footprint 160 of the block 104 in the reference image, displaced according to the motion information 132 with respect to its position in the current image. In particular, when actually performing the prediction of the block 104 from the patch 130, the decoder and the coder can predict each sample 162 of the block 104 by a mathematical combination of the samples 164 within each part of the patch 160. Usually, the part of the patch 130 that contributes to the mathematical combination of a specific block sample 162 is arranged at the position of each sample 162 and its vicinity, displaced according to the motion information 132, that is, at or around the displaced sample position 162' in the reference image 12b. For example, the mathematical combination may be a weighted sum of the samples of the part corresponding to each sample from the block 104. This part of each sample 162 can be determined, for example, by the kernel of an interpolation filter when the motion vector 132 is a sub-pixel motion vector, thus by obtaining the sub-pixel displacement of the block 104 with respect to the sub-pixel sample position in the reference image 12b. Interestingly, for example, in the case of motion information 132 including a full-pixel motion vector, the patch 130 may then be the same size as the block 104, and the interpolation filter may not be necessary such that each sample 162 is predicted by setting it equal to one corresponding sample 164 within the patch 130. Another type or origin of the contribution to the enlargement of the footprint 160 with respect to the patch 130 may alternatively or additionally be derived from other inter-prediction tools such as the BIO tool, which will be more specifically mentioned below. In particular, according to BIO, the predictor derived from the patch 130 may be one or two hypotheses of the block 104 that will be predicted bidirectionally in that case.However, in order to obtain this predictor, a provisional version of the predictor derived with interpolation in the case of sub-pixel motion vectors and without an interpolation filter in the case of full-pixel motion vectors is determined such that this provisional version is expanded by n samples in all directions with respect to the size of the actual block 104 (n > 0, such as 1). This is done to determine local luminance gradients at multiple positions distributed over the area of block 104 of this provisional predictor and to linearly combine the two hypotheses within the area of the block, i.e., the provisional predictor obtained from reference image 12b and another provisional predictor obtained similarly from another reference image, so that the predictor for the final bidirectional prediction varies over the area of the block depending on the local gradients to obtain the predictor for the final bidirectional prediction. Referring to FIG. 22, it is explained that the contribution of the latter to the patch expansion of BIO can be ignored and that the BIO tool can alternatively be deactivated depending on whether the BIO expansion contribution of the n-sample width of the patch expansion (n > 0) leads to a patch 130 that crosses the boundary 102 or not. That is, this redirection procedure is performed provisionally ignoring the BIO tool, and the BIO tool recognizes the boundary and is deactivated by the coder instead of the coder. Further, another mode that can be used in addition to the BIO tool is that motion information defines motion vectors of the motion vector field within block 104 at two different corners of block 104, and based on this, an affine motion model tool that calculates sub-block motion vectors of sub-blocks of block 104 using the affine motion model, including the two motion vectors of the current block 104. The resulting sub-block patches can be considered to form a structured patch, such as that shown in FIG. 20, consisting of sub-block patches, and again, the enforcement of tile independence can enforce that the structured patch as a whole, i.e., none of the individual sub-block patches, crosses the boundary 102.With respect to the redirect and forced procedures described with respect to FIGS. 9 and 10 respectively, the functions described above mean that patches 130 of various patch sizes are considered with respect to both the size of the patch of the actually signaled motion information that will be conditionally redirected and the motion information settings where "problematic" motion information is redirected or remapped. That is, although not explicitly mentioned with respect to FIG. 9, the redirection of a particular motion information state can lead to a change in patch size, and a state that has not yet been redirected, such as state 2 in FIG. 9, can have a patch size different from the state brought about by redirect 142, i.e., 2'.
[0059] There are various possibilities regarding how the aforementioned irreversible mapping or redirection is performed. This applies to both sides, namely, on the one hand, the state that receives a redirect when forming the input of the irreversible mapping, and on the other hand, the redirect state of the output of the irreversible mapping. FIG. 12a shows one possibility. FIG. 12a illustrates or exemplifies four different settings 1, 2, 3, and 4 that have not yet been redirected or subjected to an irreversible mapping by showing the corresponding patches with respect to the boundary 102 that separates the current tile 100a on the one hand and the neighboring tile 100b on the other hand. In the first three states 1, 2, and 3, the patch 130 is smaller than in state 4. In particular, FIG. 12a shows that in states 1, 2, and 3, the patch 130 coincides with the footprint 160 of the inter-prediction block. In state 4, the patch 130 is enlarged with respect to the footprint 160 shown by the dashed line for this setting 4 in FIG. 12a. The enlargement is due to, for example, the corresponding motion vector in state 4 being a sub-pixel motion vector while in states 1 - 3 it is a full-pixel motion vector. Due to this enlargement, the patch 130 in states 1 - 3 is within the current patch 100a and does not cross the boundary 102, while the patch 130 in the case of state 4, although its corresponding motion vector is between the full-pixel motion vectors of states 2 and 3, extends beyond the boundary 102.
[0060] That is, in any case, compliance / enforcement of the constraints must redirect state 4. According to some embodiments, as will be described later, states 1, 2, and 3 are left as they are because the corresponding patches do not cross boundary 102. What this means is that, in effect, the range of values of the irreversible mapping, i.e., the reservoir or region of the redirect states, allows the patch 130 to approach the boundary 102 of the current tile 100a more than the width 170 of the extended edge portion 172 where the patch 130 is enlarged compared to the footprint 160 of the inter-prediction block 104. In other words, the full-pixel motion vectors leading to patches 130 that do not cross boundary 102, for example, are left unchanged in the redirect process 142. In different embodiments, even those states of the motion information are redirected, and the associated patches that do not extend beyond boundary 102 to neighboring tiles, are not extended with respect to the footprint, and would cross boundary 102 if extended, in the example of FIG. 12a, these are states 1, 2, and 3, and even they will be redirected at least to place a distance from boundary 102 to adapt to width 170. See FIG. 12b, where the patch 130 of state 3 is shown to have a distance from boundary 102 smaller than the extension range 170. Thus, its associated patch 130 or footprint 160 is redirected to the full-pixel state 5 having a distance 171 at least as large as the extension range 171. That is, here state 3 is redirected to state 5 where the corresponding patch 130 does not include the boundary area 172, yet is placed at a distance of at least the same size as width 170, i.e., distance 171, from boundary 102.
[0061] In other words, as a result of the sub-pixel interpolation process used in many video coding specifications, samples may be required that are a plurality of integer sampling positions away from the nearest full-pixel reference block position. Assume that the motion vector MV1 is a full-pixel motion vector and MV2 is a sub-pixel motion vector. More specifically, the position Pos x , Pos y within tile 100a of size B W×B H Assume block 104. Focusing on the horizontal component, to avoid using neighboring tile 100b or the tile on the right, MV1 can be at most (assuming MV1 x >0) Pos x +B W -1 + Mv1 x equal to the sample at the right end of tile 100a to which block 104 belongs, or (assuming MV1 x <0, i.e., the vector points to the left) Pos x + Mv1 x can be made equal to the sample at the left end of tile 100a to which the block belongs. However, for MV2, interpolation overhead from the filter kernel size must be incorporated to avoid reaching within neighboring tiles. For example, considering Mv2 x as the integer part of MV2, MV2 can be at most (assuming MV2 x >0) Pos x +B W -1 + Mv2 x equal to the sample -4 at the right end of the tile to which the block belongs, and (assuming MV2 x <0) Pos x + Mv2 x can only be made equal to the sample +3 at the left end of the tile to which the block belongs.
[0062] "3" and "-4" are here examples of width 170 caused by the interpolation filter. However, other kernel sizes can also be applied.
[0063] However, it is desirable to establish a means to efficiently limit motion compensation prediction at tile boundaries. Some solutions to this problem are described below.
[0064] One possibility is to clip the motion vectors with respect to independent spatial regions within the coded image, i.e., the boundaries 102 of tile 100. Additionally, the motion vectors can be clipped to adapt to the motion vector resolution.
[0065] Assume that the following block 104 is given: Size B W xB H Position Pos x , Pos y Tile Left , Tile Top , Tile Right and Tile Bottom belonging to tile 100a having a boundary 102 defined by Define the following: When MVX >> precision, MVX Int When MVY >> precision, MVY Int When MVX & (2 precision - 1), MVX Frac When MVY & (2 precision - 1), MVY Frac MVX is the horizontal component of the motion vector, MVY is the vertical component of the motion vector, and precision represents the precision of the motion vector. For example, in HEVC, the precision of the MV is 1 / 4 of a sample, and thus precision = 2.
[0066] Clip the full - pixel part of the motion vector MV so that the resulting clipped full - pixel vector with zeros set in the sub - pixel part becomes exactly the footprint or patch within tile 100a. MVX Int = Clip3(Tile Left - Pos x , Tile Right - Pos x -(B W - 1), MVX Int ) MVY Int = Clip3(Tile Top - Pos y , Tile Bottom - Posy -(B H -1),MVY Int )
[0067] Horizontally, MVX (assuming the 8-tap filter of HEVC), which means that the full-pixel part of the motion vector (after clipping) is closer to the boundary 102 than the expansion range 170 Int <=Tile Left -Pos x +3 or MVX Int >=Tile Right -Pos x -(B W -1)-4, then MVX Frac is set to 0.
[0068] Vertically, MVY (assuming the 8-tap filter of HEVC), which means that the full-pixel part of the motion vector (after clipping) is closer to the boundary 102 than the expansion range 170 Int <=Tile Top -Pos y +3 or MVY Int >=Tile Bottom -Pos y -(B H -1)-4, then MVY Frac is set to 0.
[0069] This corresponds to the situation described with respect to FIG. 12a where the full pixel motion vectors in states 1, 2, and 3 are allowed to remain as they are because the corresponding patches do not extend into the neighboring tile 100b. The full pixel motion vectors leading to a patch 130 that extends beyond the boundary 102 are clipped to the nearest full pixel motion vector within the boundary 102 such that the corresponding patch is within the boundary 102 without crossing it. For sub-pixel motion vectors where the enlarged patch 130 crosses the boundary 102, the following procedure is applied: The sub-pixel motion vector is mapped to the nearest full pixel motion vector that is smaller than it and for which the patch remains within tile 100a. That is, the full pixel part of the motion vector is clipped accordingly and the sub-pixel part is set to zero. What this means is that the sub-pixel part of the sub-pixel motion vector is simply set to zero if the full pixel part of the sub-pixel motion vector, i.e., the truncated version of the sub-pixel motion vector, does not lead to a footprint 160 leaving tile 100a. State 4 is thus redirected to state 2.
[0070] Instead of setting the fractional component to zero, the integer pixel part of the motion vector can also be rounded to the nearest neighboring integer pixel position spatially as follows, rather than being averaged to the next smaller integer pixel position as described above. As before, clip the full pixel part of the motion vector MV so that the resulting clipped full pixel vector, when the sub-pixel part is set to zero, results in a footprint or patch that is strictly within tile 100a. MVX Int =Clip3(Tile Left -Pos x ,Tile Right -Pos x -(B W -1),MVX Int ) MVY Int =Clip3(Tile Top -Pos y ,Tile Bottom -Pos y-(B H -1), MVY Int ) (assuming the 8-tap filter of HEVC) MVX Int <= Tile Left -Pos x +3 or Tile Right -Pos x -(B W -1) > MVX Int >= Tile Right -Pos x -(B W -1) - 4 cases MVX Int = MVX Int +(MVX Frac +(1 << (precision - 1)) >> precision) (assuming the 8-tap filter of HEVC) MVX Int <= Tile Left -Pos x +3 or MVX Int >= Tile Right -Pos x -(B W -1) - 4 cases MVX Frac is set to 0. (assuming the 8-tap filter of HEVC) MVY Int <= Tile Top -Pos y +3 or Tile Bottom -Pos y -(B H -1) > MVY Int >= Tile Bottom -Pos y -(B H -1) - 4 cases MVY Int = MVY Int +(MVY Frac +(1 << (precision - 1)) >> precision) (assuming the 8-tap filter of HEVC) MVY Int <= Tile Top -Pos y +3 or MVY Int >= Tile Bottom-Pos y -(B H - In the case of (1)-4 MVY Frac is set to 0.
[0071] That is, the above procedure is changed as follows. The motion vector is rounded to the nearest full-pixel motion vector that does not leave tile 100a. That is, the full-pixel portion can be clipped as necessary. However, if the motion vector was a sub-pixel motion vector, the sub-pixel portion is not simply set to zero. Rather, rounding is performed to the nearest full-pixel motion vector, that is, to a full-pixel motion vector different from the initial sub-pixel motion vector by conditionally clipping the full-pixel portion. This leads, for example, to mapping or redirecting state 4 in FIG. 12a to state 3 instead of state 2.
[0072] Put another way, according to the first alternative described above, the motion vector by which the patch 130 extends between the boundaries 102 is redirected to a full-pixel motion vector by clipping the full-pixel portion of the motion vector so that the footprint 160 remains within tile 100a, in which case the sub-pixel is set to zero. As a second alternative, after clipping the full-pixel portion, rounding to the nearest full-pixel motion vector is performed.
[0073] Another option is to clip the integer and fractional parts of the motion vector together in a way that is more restrictive and ensures that the ultimately generated sub-sample interpolation that incorporates samples from another tile is avoided. This possibility is referred to in FIG. 12b.
[0074] MVX = Clip3((Tile Left -Pos x + 3) << precision, (Tile Right -Pos x -B W - 1 - 4) << precision, MVX) MVY = Clip3((Tile Top - Pos y + 3) << precision, (Tile Bottom - Pos y - B H - 1 - 4) << precision, MVY)
[0075] As is apparent from the above description, MV clipping depends on the block size to which the MV is applied.
[0076] In the description, the juxtaposition of color components with different spatial resolutions was ignored. Therefore, the above clipping procedure does not take into account the chroma format, or only considers the mode 4:4:4 where there is a chroma sample for each luma sample. However, there are two additional chroma formats where the relationship between chroma samples and luma samples is different. 4:2:0, where the chroma has half the luma samples in the horizontal and vertical directions (i.e., one chroma sample for every two luma samples in each direction) 4:2:2, where the chroma has half the samples in the horizontal direction but the same samples in the vertical direction (i.e., one chroma sample per luma sample in the vertical direction, but one chroma sample for every two luma samples in the horizontal direction)
[0077] The chroma sub - pixel interpolation process may use a 4 - tap filter, while the luminance filter uses an 8 - tap filter as described above. As mentioned above, the exact number is not important. The chroma interpolation filter kernel size is half that of the luma, but this may also be changed in alternative embodiments.
[0078] In the case of 4:2:0, the derivation of the integer and fractional parts of the motion vector is as follows. In the case of MVCX>>(precision + 1), MVX Int In the case of MVCY>>(precision + 1), MVY Int MVCX & (2 precision+1In the case of -1), MVX Frac MVCY & (2 precision+1 In the case of -1), MVY Frac
[0079] What this means is that the integer part is half of the corresponding luma part, and the fractional part has finer-grained signaling. For example, in HEVC with precision = 2, in the case of luma, there are 4 sub-pixels that can be interpolated between samples, and in chroma, there are 8.
[0080] This is xPos + MVX Int = Tile Left + 1, and MVX Int When it is defined as luma (not chroma as described above) in the 4:2:0 format, this leads to it falling into integer luma samples rather than fractional chroma samples. Such samples require 1 chroma sample beyond Tile Left by sub-pixel interpolation, which will prevent tile independence. The cases where such problems occur are as follows: 4:2:0 (ChromaType = 1 in HEVC) xPos + MVX Int = Tile Left + 1, MVX Int is defined for luma xPos + (B W - 1) + MVX Int = Tile Right - 2, MVX Int is defined for luma yPos + MVY Int = Tile Top + 1, MVY Int is defined for luma yPos + (B H - 1) + MVY Int = Tile Bottom - 2, MVY Int is defined for luma 4:2:2 (ChromaType = 2 in HEVC) xPos + MVXInt = Tile Left +1, MVX Int is defined for luma xPos + (B W - 1) + MVX Int = Tile Right -2, MVX Int is defined for luma.
[0081] There are two possible solutions. Clip in a restricted way based on ChromaType (Ctype): ChromaOffsetHor = 2 * (Ctype == 1 || Ctype == 2) ChromaOffsetVer = 2 * (Ctype == 1) MVX Int = Clip3(Tile Left - Pos x + ChromaOffsetHor, Tile Right - Pos x - (B W - 1) - ChromaOffsetHor, MVX Int ) MVY Int = Clip3(Tile Top - Pos y + ChromaOffsetVer, Tile Bottom - Pos y - (B H - 1) - ChromaOffsetVer, MVY Int )
[0082] Or, clip in the same way as the previous case without additional changes by chroma, but check whether the following is true: 4:2:0 (ChromaType = 1 in HEVC) xPos + MVX Int = Tile Left +1, MVX Int is defined for luma xPos + (B W - 1) + MVX Int = Tile Right -2, MVXInt is defined for luma yPos + MVY Int = Tile Top + 1, MVY Int is defined for luma yPos + (B H - 1) + MVY Int = Tile Bottom - 2, MVY Int is defined for luma 4:2:2 (ChromaType = 2 in HEVC) xPos + MVX Int = Tile Left + 1, MVX Int is defined for luma xPos + (B W - 1) + MVX Int = Tile Right - 2, MVX Int is defined for luma
[0083] Also, MVX Int or MVY Int is changed so that (one or more) prohibited conditions do not occur (e.g., depending on the fractional part rounded in the nearest direction by +1, or by +1 or -1).
[0084] In other words, the above description has focused only on one color component and ignored cases where different color components of the video image may have different spatial resolutions. However, the juxtaposition of different color components with different spatial resolutions may be taken into account in the enforcement of the tile independence constraint, and this description also applies to the following described variants where the enforcement is applied to, for example, the MI predictor rather than the final MI. To put it bluntly, the motion vector can always be treated as a sub - pixel motion vector if any of the color components requires an interpolation filter for which the patch is expanded accordingly, and similarly, the motion vectors for which redirection is performed are also selected so as to avoid the interpolation filtering that may be required for any of the color components crossing the boundary 102 of the current tile 100a.
[0085] According to an alternative embodiment, the above-described embodiment (compared to FIG. 10) in which compliance / enforcement of the constraint is applied to the finally signaled motion vector 154 is changed insofar as the decoder applies the same concept, i.e., compliance / enforcement of the constraint, i.e., clipping in the case of a motion vector, to the predictor 150 rather than to the final motion vector or motion information 154. Thus, both the encoder and the decoder operate in the same way to enforce that all motion information predictors 150 comply with the constraint that the patch of block 104 is within the current tile 100a when the patches of block 104 are arranged according to their respective motion information predictors 150, i.e., when the motion information prediction residual 152 is zero. All of the above details can be transferred to these alternative embodiments.
[0086] However, insofar as the encoder 10 is concerned, the above alternative embodiment leads to different situations for the encoder. That is, there is no ambiguity or “freedom” for the encoder to make the decoder use the same motion information for block 104 by signaling one of the different motion information prediction residuals. Rather, when the motion information to be used for block 104 is selected, the signaling within part 140 is uniquely determined at least for certain motion information predictors 150. The advantages are as follows. Since the motion information predictor 150 is prevented from leading to a contradiction with the tile independence constraint, there is no need to “redirect” such a motion information predictor 150 by each non-zero motion information prediction residual 152 whose signaling is usually more cost-intensive with respect to the bitrate than the signaling of a zero motion information prediction residual. Further, when establishing a list of motion information predictors for block 104, the encoder is forced, in any case, to select the motion information in such a way that the tile independence constraint is complied with by an automatic synchronous enforcement that all motion information predictors read into such a list of motion information predictors for block 104 do not contradict the tile independence constraint. Thus, it is more likely that any of these available motion information predictors is quite close to the best motion information with respect to rate distortion optimization.
[0087] That is, in the current video coding standard, MV clipping is performed on the final MV, i.e., the difference of motion vectors is, if any, added to the predictor and then executed, which is done for prediction, and here the correction of this prediction using the residual is additionally performed. When the predictor points outside the image (potentially with boundary extension), if the predictor points really far from the boundary, the difference of motion vectors will probably, at least when the block is at the image boundary after clipping, not include the component in the direction in which it is clipped. The difference of motion vectors will make sense only when the resulting motion vector points to a position within the reference image inside the image boundary. However, adding such a large difference of motion vectors for block reference within the image may be too costly compared to clipping the difference of motion vectors to the image boundary.
[0088] Therefore, one embodiment is based on clipping the predictor depending on the block position such that all MV predictors used from neighboring blocks or temporal candidate blocks always point inside the tile in which the block is included, and thus the remaining difference of motion vectors to be signaled to the better predictor becomes smaller and can be signaled more efficiently.
[0089] One embodiment that utilizes the above-mentioned tile independence constraint related to the motion vector predictor is shown in FIG. 13. In particular, FIG. 13 shows what has often been mentioned as a possibility above, namely, that a list 190 of motion information predictor candidates 192 is established for block 104 using a specific order according to specific rules. One of these motion information predictor candidates 192 is then selected by the coder for block 104 and signaled to be selected within the syntax portion 140 for block 104 by respective pointers 193 that indicate the position of the selected motion information predictor candidate 192 with respect to its rank within list 190. The establishment of list 190 is performed in the same way on the coder side and the decoder side, and according to the current embodiment, it involves a tile independence constraint for each motion information predictor candidate 192 loaded into list 190. Thus, all the motion information predictors 192 within list 190 are not associated with patches that do not exclusively exist within the boundaries of the current tile to which block 104 belongs. A further syntax element 194 within portion 140 then indicates the motion information prediction residual in the case of motion information predictor candidate 192 for the motion vector, i.e., the difference 152 of the motion vectors.
[0090] FIG. 14 shows the effect of forcing tile independence on each motion information prediction candidate 192 in list 190. For block 104, patches 130a to 130c of three motion information predictor candidates 192a to 192c in list 190 are shown. However, patch 130c is associated with motion information predictor candidate 192c, which is, however, actually the result of redirect 142. That is, the encoder and decoder actually derived a motion information predictor candidate for block 104 indicated by the dashed line and 192c'. This predictor 192c', however, should have led to a patch 130c' that is not within the boundary 102 of the current tile 100a. Therefore, this motion information predictor candidate 192c' is redirected to become motion information predictor candidate 192c. The effect is as follows. The possibility that motion information predictor candidate 192c' becomes the most effective motion information predictor candidate in terms of the meaning of large distortion is quite low due to the necessity for the encoder to have a non-zero motion information predictor residual. Therefore, including such a candidate 192c' in list 190 has the highest possibility of being only a waste of candidate positions in list 190. Rather, any of predictors 192a to 192c may be selected by pointer 193, and in most cases, it is only necessary to transition to the final motion vector 154 (compared with FIG. 10) so that only a small residual 194 is obtained by the combination of the residual vector 152 (compared with FIG. 10) and the selected predictor such as 192c.
[0091] The concepts of FIGS. 13 and 14 may of course be applied to a single prediction 150 for the current block 104 without establishing list 190 at all.
[0092] An alternative concept regarding the construction / establishment of the motion information predictor candidate list is the subject of the embodiment described next with respect to FIG. 15. As already mentioned above with respect to FIG. 13 as a possibility, the list 190 can be loaded using motion information predictor candidates derived from previously encoded / decoded blocks according to a predetermined rule. These rules can relate, for example, to the way in which the corresponding previously encoded / decoded blocks are arranged. This can be, for example, a block spatially adjacent to the left of the current block 104 or a block spatially adjacent above the current block 104. The derivation of each motion information predictor candidate can then simply be the adoption of the respective motion information used for each block with the same encoding / decoding, i.e., each motion information predictor candidate for the block 104 can be set to be equal to the respective motion information used for each block. Another motion information predictor candidate can be derived from a block in the reference image. Also in this case, this can be the reference image to which one of the candidates 192 is related, or another image. FIG. 15 shows such a motion information predictor candidate primitive 200. The primitive 200 forms a pool of motion information predictor candidates from which the entries in the list 190 are read according to an order 204 called the reading order. According to this order 204, the encoder and decoder derive the primitive 200 and check their availability at 202. The reading order is indicated at 204. The availability can be negated, for example, because the corresponding block of each primitive 200, i.e., the block whose motion information was used for the same predictive encoding / decoding, is not within the same tile 100a as the current block 104. However, according to the embodiment of FIG. 15, the availability check 202 accompanies or includes, additionally or alternatively, the following check, i.e., a check as to whether the corresponding primitive 200 conflicts with the tile independence constraint.That is, instead of checking the availability of the origin of each primitive 200, i.e., the left block, upper block, or a collocated block in the reference picture of the current block 104 whose motion information used for the same predictive encoding / decoding is to be adopted to form the primitive 200, it is checked whether the motion vector predictor is represented by such a primitive 200, i.e., whether the motion information derived from the motion information used to predictively encode / decoded the previously encoded / decoded block associated with the primitive and to be adopted to form the primitive 200 is associated with the patch 130 located within the current tile 100a without crossing its boundary. If not, each primitive is marked as unavailable and the next primitive is proceeded to. FIG. 15 shows, for example, that the second candidate primitive 200 is unavailable and the first two motion information predictor candidates 192 in the list 190 are formed by the first and third candidate primitives 200 in the order 204 accordingly. In other words, another solution to the above problem is to change the derivation of the availability flag for the motion vector list construction process. This can be done in a way that motion vectors using samples outside the target independently encoded region (e.g., the currently encoded motion constraint tile) are set to be unavailable as if the constraintFlag is set.
[0093] For example, the following is the construction process for the motion vector predictor candidate list, mvpListLX. i = 0 if (availableFlagLXA) { mvpListLX[i++] = mvLXA if (availableFlagLXB && (mvLXA!= mvLXB)) mvpListLX[i++] = mvLXB } else if (availableFlagLXB) mvpListLX[i++] = mvLXB if (i < 2 && availableFlagLXCol) mvpListLX[i++] = mvLXCol HMVP while (i < 2) { mvpListLX[i][0] = 0 mvpListLX[i][1] = 0 i++ }
[0094] If mvLXA and mvLXB are available but point outside the given MCTS, it can be seen that perhaps more promising collocated MV candidates or zero motion vector candidates are not input into the list because the list is already full. Therefore, it is advantageous to have a constraintFlag signaling in the bitstream that controls the derivation of MV candidate availability in a way that incorporates the availability of reference samples regarding spatial segment boundaries such as tiles.
[0095] In another embodiment, the availability derivation in the context of bidirectional prediction allows for a more granular description of the available state and can thus further allow for loading a mixed version of partially available candidates into the motion vector candidate list.
[0096] For example, when the current block is bidirectionally predicted in the same way as its bidirectionally predicted spatial neighborhood, the above concept (availability marking depending on the reference sample position within the tile) causes the spatial candidate whose MV0 points outside the current tile to be marked as unavailable, and thus all candidates having MV0 and MV1 are not input to the candidate list. However, MV1 is a valid reference within the tile. To make this MV1 accessible via the motion vector candidate list, MV1 is added to a temporary list of partially available candidates. Combinations of partial candidates within that temporary list can also be added following the final mv candidate list. For example, mix the MV0 of spatial candidate A with the MV1 or zero motion vector or HMVP candidate component of spatial candidate B.
[0097] From the latter clue, it became clear that the process shown in FIG. 15 can be performed by the coder and decoder with a little more sophistication when block 104 is of the bidirectional prediction type. In this case, two motion vectors of two different reference images are transmitted by each candidate primitive 200. One possibility would have been to mark such a candidate primitive 200 as unavailable if one of the motion vectors leads to a contradiction with tile independence. The pair of motion vectors is thus skipped in the reading order and not added to list 190. However, according to an alternative, the tile independence check is performed for each hypothesis individually. That is, for a particular candidate primitive 200, one of the motion vectors may lead to a contradiction with tile independence while the other may not. In that case, the non - conflicting motion vector hypothesis can be added to a particular replacement list of the replacement motion information predictor. If necessary, i.e., when no more candidate primitives 200 are available in the reading order 204, the coder and decoder can form one or more other motion information predictor candidates 192 for list 190 from a single hypothesis motion information predictor and replacement list, such as by combining their pairs or one entry and replacement list with a default motion information predictor such as a zero motion vector.
[0098] Note that in order to conclude the description of FIG. 15, the remaining details described with respect to the foregoing figures may be added. For example, such details may relate to the actual use of the motion information finally signaled to predict the pointer 193 and the residual 194, as well as the patch size and the block 104.
[0099] The embodiments described next handle motion information prediction candidates having their origin in blocks of a specific reference image. Also in this case, this reference image need not be the image from which the actual inter prediction is performed. However, this may be the case. In order to distinguish the reference image used for MI (motion information) prediction from the reference image including the patch 130, an apostrophe is used hereinafter for the former.
[0100] FIG. 16 shows a current inter-prediction block 104 in the current image 12a and its collocated portion in the reference image 12b’, i.e., the non-displaced footprint 104’. It has been found that the temporal motion information predictor candidate (primitive) is cost-effective with respect to RD performance if it is formed using the motion information of the block in the reference image 12b’ that is used in the following manner. Preferably, the block in the reference image 12b’ includes a predetermined position 204’ in the reference image 12b, i.e., a block that includes sample positions collocated with a specific alignment position 204 in the current image 12a that has a predetermined positional relationship to the block 104. In particular, it is a sample position that is in the vicinity of the diagonal of the bottom-right sample of the block 104 but outside the block 104. The position 204 is thus offset and arranged with respect to the block 104 along the horizontal and vertical directions in the example of FIG. 16. The block of the reference image 12b’ that includes the position 204’ is indicated by reference numeral 206 in FIG. 16. Its motion information is used to derive or directly represent the temporal motion information predictor candidate (primitive) if available. However, the availability may not apply, for example, when the positions 204’ and 204 outside the image or block 206 are not inter-prediction blocks but, for example, intra-prediction blocks. In such a case, i.e., in case of non-availability, another alignment position 208 is used to identify the source block or the original block in the reference image 12b. In this case, the alignment position 208 is arranged inside the block 104, such as being centered. The corresponding position in the reference image is indicated by reference numeral 210. The block that includes the position 210 is indicated by reference numeral 212, and this block is used to form the temporal motion information predictor candidate (primitive) by substitution, i.e., using the motion information by which this block 212 was predicted using it.
[0101] To avoid the problems that would occur if the above concept were applied to all blocks 104 within the current image 12a, the following alternative concept is applied. In particular, the encoder and decoder check whether block 104 is adjacent to a predetermined side of the current tile 100a. The predetermined side is, for example, the right boundary of the current tile 100a and / or the lower side of this tile 100a as in the case of this example, and for example, the alignment position 204 is offset with respect to block 104 in both the horizontal and vertical directions. Thus, if block 104 is adjacent to one of these sides, the encoder and decoder identify the source block 206 of the temporal motion information predictor candidate (primitive) and skip the use of alignment position 204 in order to use only alignment position 208 instead. The latter alignment position "hits" only the blocks within the reference image 12b that are within the same tile 100a as the current block 104. For all other blocks that are not adjacent to the right or lower side of the current tile 100a, the derivation of the temporal motion information predictor candidate (primitive) can be performed including the use of alignment position 204.
[0102] However, it should be noted that many variations are possible with respect to the embodiment of FIG. 16. Such possible modifications from the description presented with respect to FIG. 16 are related, for example, to the exact positions of alignment positions 204 and 208.
[0103] In other words, typical state-of-the-art video coding specifications also rely heavily on the concept of collecting motion vector predictors in so-called candidate lists at various stages of the coding process. Below, the concept in the context of MV candidate list construction that enables higher coding efficiency for inter prediction constrained coding will be described.
[0104] The merge candidate list can be interpreted as follows. I = 0 if(availableFlagA1) mergeCandList[i++] = A1 if(availableFlagB1) mergeCandList[i++] = B1 if (availableFlagB0) mergeCandList[i++] = B0 if (availableFlagA0) mergeCandList[i++] = A0 if (availableFlagB2) mergeCandList[i++] = B2 if (availableFlagCol) mergeCandList[i++] = Col
[0105] When slice_type is equal to B and there are not enough candidates in many video codec specifications, the derivation process for combined bidirectional prediction merge MV candidates is performed to fill the candidate list. Candidates are combined by taking various combinations of the L0 component of one MV candidate and the L1 component of another MV candidate. Zero motion vector merge candidates are added when there are not enough candidates in the list.
[0106] When considering the collocated MV candidates (Col), relevant problems of the merge list occur in the context of inter prediction constraints. When the collocated block (collocated not at the lower right but in the center) is not available, it is impossible to know whether there is a Col candidate without parsing the adjacent tiles in the reference picture. Therefore, the merge list is different when decoding all tiles and when decoding only one tile, or the tiles are decoded in an arrangement different from that at the time of encoding. Therefore, there is a possibility that the MV candidate lists on the encoder side and the decoder side do not match, and candidates from a specific on index (col index) cannot be safely used.
[0107] The above-described embodiment of FIG. 16 that avoids the problem of the mismatch in the above MV candidate list has, for the rightmost block and the bottommost block within a tile, a block juxtaposed in the center as a Col candidate belonging to the current tile, rather than a bottom-right juxtaposed block not belonging to the current tile.
[0108] An alternative solution would be to change the list configuration regarding the reading order 204. In particular, when using the concepts described with respect to FIG. 16, it is possible to use a temporal motion information predictor candidate (primitive) relatively early in the list configuration, such as the first primitive 200 for which availability has been checked (compare with FIG. 15). By the tile boundary recognition on the encoder side and the decoder side, it is guaranteed that no list mismatch occurs. However, another possibility is to shift such a temporal motion information predictor candidate (primitive) towards the end of list 190, and shift the combined spatial motion information candidate, such as the combined spatial motion vector, preferably in front of the temporally juxtaposed candidates determined only in an auxiliary way based on block 212 and based on block 206. In relation to the above pseudo-code, this means that Col is a candidate determined mainly only based on the MI of block 206 and the MI of block 212 as a substitute, and is moved to the end of list 190 completed after the current state-of-the-art list configuration. This change in the reading order can be done with tile boundary recognition. That is, the encoder and the decoder change the reading order by shifting the Col candidate to the end of list 190 only for the blocks 104 adjacent to the right side or the bottom side of the current tile, and leave the other reading orders in the other way, that is, the way where Col precedes the combined spatial motion information candidate.
[0109] Of course, for a more unified creation of the motion vector candidate list, it would also be feasible to perform the above reading order change for all blocks 104 within the tile, not only for the blocks adjacent to a particular side.
[0110] FIG. 17 shows the changed reading order for tile boundary recognition. In the left half of FIG. 17, cases are shown where the inter-prediction block 104 is adjacent to one of the sides of the current tile in question, i.e., the right or bottom side, and on the right side of FIG. 17, cases are shown where the block 104 is somewhere within the current tile 100a, or more precisely, somewhere at a distance from a particular side of the current tile 100a.
[0111] Using 200a, motion information predictor candidate primitives are indicated, and these primitives are derived in the aforementioned manner where, for example, in the case of 206 which is an intra coding type, only the motion information of block 212 is used as a substitute, and the motion information of block 206 is preferably derived accordingly. The blocks that form the basis of primitive 200a are also called aligned blocks at 206 or 212. Another motion information predictor candidate primitive 200b is shown. This motion information predictor candidate primitive 200b is derived from the motion information of spatially neighboring blocks 220 that are spatially adjacent to the current block 104 on the upper and left sides, i.e., on the side opposite to a particular side of tile 100a. For example, the average or median value of the motion information of neighboring blocks 220 is used to form the motion information predictor candidate primitive 200b. However, the reading direction 204 is different in both cases. When the block 104 is adjacent to one of the particular sides of the current tile 100a, the combined spatial candidate primitive 200b precedes the temporally juxtaposed candidate primitive 200a, while in the other case, i.e., when the block 104 is not adjacent to any of the particular sides of the current tile 100a, the order is changed such that the temporally juxtaposed candidate primitive 200a reads into list 190 earlier than the combined spatial candidate primitive 200b in the rank order 195 along which the pointer 193 indicates list 190.
[0112] Note that many details not specifically discussed with respect to FIG. 17 may be adopted from any of the foregoing embodiments, such as the embodiment with respect to FIG. 16 related to the primitive 200a. The (one or more) primitives 200b may also be obtained separately.
[0113] FIG. 18 shows a further possibility of efficiently achieving tile-independent coding. Here, accordingly, the motion vector 240 predicted in another way is used by the encoder and decoder to derive a temporal motion information prediction candidate (primitive) for identifying or placing a predetermined block 242 in the reference image 12b, where the motion information, i.e., the motion information used to encode / decode the block 242, is then used. The motion vector 240 can be predicted spatially, for example. The motion vector 240 may be derived from one of the other candidates in the list 190, rather than from the primitive targeted in FIG. 18. According to the embodiment described with respect to FIG. 18, the encoder and decoder check whether the motion vector 240 points outside the current tile 100a with respect to the position of the block 104. In the case of FIG. 18, it is shown that the predicted motion vector 240 points outside the tile 100a. This motion vector 240 is thus clipped by the clipping operation 244 so as to start from the block 104, more precisely from its predetermined alignment position 246, and stay within the boundary 102 of the current tile 100a, i.e., not point outside the tile 100a. The block 248 indicated by the motion vector 250 thus clipped with respect to the block 104 is then used as a basis for deriving a temporal motion information prediction candidate (primitive), such as by simply having the block 248 adopt the motion information predicted using it.
[0114] Another possibility is shown in FIG. 19. Here, the encoder and decoder first test a first predicted motion vector 240. By doing this, the encoder and decoder check whether this first predicted motion vector points outside tile 100a. If so, the encoder and decoder use another predicted motion vector 260, i.e., use it to derive a temporal motion information prediction candidate (primitive) by the block 262 indicated by the second predicted motion vector 260 adopting the predicted motion information. In FIG. 19, for example, it is shown that the first predicted motion vector 240 points outside the current tile 100a such that the block 242 indicated by this motion 240 is outside the current tile 100a in the reference picture 12b. However, the motion vector 260 does not point outside the current tile 100a. In the latter case, each temporal motion information prediction candidate (primitive) can be marked as unusable.
[0115] In other words, in usage scenarios with independently encoded spatial segments such as MCTS, the above ATMVP procedure needs to be restricted to avoid dependencies between MCTS.
[0116] In the sub-block merge candidate list configuration, subblockMergeCandList is configured as follows, and the first candidate SbCol is the sub-block temporal motion vector predictor. i = 0 if (availableFlagSbCol) subblockMergeCandList[i++] = SbCol if (availableFlagA && i < MaxNumSubblockMergeCand) subblockMergeCandList[i++] = A if (availableFlagB && i < MaxNumSubblockMergeCand) subblockMergeCandList[i++] = B if (availableFlagConst1 && i < MaxNumSubblockMergeCand) subblockMergeCandList[i++] = Const1 if (availableFlagConst2 && i < MaxNumSubblockMergeCand) subblockMergeCandList[i++] = Const2 if (availableFlagConst3 && i < MaxNumSubblockMergeCand) subblockMergeCandList[i++] = Const3 if (availableFlagConst4 && i < MaxNumSubblockMergeCand) subblockMergeCandList[i++] = Const4 if (availableFlagConst5 && i < MaxNumSubblockMergeCand) subblockMergeCandList[i++] = Const5 if (availableFlagConst6 && i < MaxNumSubblockMergeCand) subblockMergeCandList[i++] = Const6
[0117] It is advantageous to prevent the correlation block in the reference frame resulting from the temporal vector from belonging to the spatial segment of the current block.
[0118] The positions of the collocated prediction blocks are constrained to be arranged in the following manner. The motion vector mvTemp, i.e., the motion information used to place the collocated sub-blocks in the reference image, is clipped to be contained within the collocated CTU boundary. This ensures that the collocated prediction blocks are placed in the same MCTS region as the current prediction block.
[0119] Alternatively, for the above clipping, when the mvTemp of the spatial candidate does not point to a sample position within the same spatial segment (tile), the next available MV candidate from the current block's MV candidate list is selected until a candidate is found that results in a reference block located within the spatial segment.
[0120] The embodiments described next address another effective encoding tool for reducing the bit rate or efficiently encoding video. According to this concept, the motion information transmitted within the data stream for block 104 enables the decoder and encoder to define more complex motion fields within block 104, i.e., not only constant motion fields but also changing motion fields. This motion information determines, for example, two motion vectors of the motion fields at two different corners of block 104, as exemplarily shown in FIG. 5. The decoder and encoder are thus in a position to derive the motion vector of each of the sub-blocks into which block 104 is divided from the motion information. Although FIG. 20 exemplarily shows a regular division into 4×4 sub-blocks 300, the division into sub-blocks 300 does not have to be regular, nor is the number of sub-blocks limited to 16. The subdivision into sub-blocks 300 can be signaled within the data stream or set by default. Thus, each sub-block generates a motion vector indicating the translational displacement between this sub-block 300 and the corresponding patch 302 within the reference image 12b from which each sub-block 300 is predicted.
[0121] To avoid a contradiction with tile independence, the encoder and decoder operate as follows according to this embodiment. The derivation of the sub-block motion vector 304 based on the motion information transmitted in the data stream exemplified by the two motion vectors 306a and 306b in FIG. 20 is performed in the same way by the encoder and decoder. The resulting sub-block motion vector 304 is assumed to be in a provisional state. According to the first embodiment, the decoder enforces the above-mentioned tile independence described above with respect to FIG. 9 for each of the provisional sub-block motion vectors 304. Thus, for each sub-block, the decoder tests whether the corresponding sub-block motion vector 304 leads to the corresponding patch 302 of each sub-block that extends beyond the boundary of the current tile 100a. If so, the corresponding sub-block motion vector 304 is treated accordingly. Each sub-block 300 is then predicted using its sub-block motion vector 304 that may be redirected. The encoder does the same so that all predictions of the sub-blocks 300 are the same on the encoder side and the decoder side.
[0122] Alternatively, the two motion vectors 306a and 306b are redirected so that none of the sub-block motion vectors 304 lead to tile dependence. Thus, clipping of the vectors 306a and 306b on the decoder side can be used for this purpose. That is, as described above, the decoder treats the pair of motion vectors 306a and 306b as motion information, treats the configuration of the patch 302 of the sub-block 300 as a patch of the block 104, and enforces that the configured patch does not conflict with tile independence. Similarly, when the configured patch reaches another tile, as described above, the motion information corresponding to the motion vector pair 306a and 306b can also be deleted from the candidate list. Or, residual correction is performed by the encoder, that is, the corresponding MV difference 152 is added to the vectors 306a and 306b.
[0123] According to the alternative, for sub-blocks 300 where the corresponding patch 302 crosses the tile boundary of the current tile 100a, prediction is performed in another way, such as using intra prediction. This intra prediction may, of course, be performed after inter predicting other sub-blocks of block 104 whose corresponding sub-block motion vectors 304 do not conflict with tile independence. Furthermore, even the prediction residuals transmitted in the data stream for these non-conflicting sub-blocks may already have been used on the encoder side and the decoder side to reconstruct the interior of these non-conflicting sub-blocks before performing the intra prediction of the conflicting sub-blocks.
[0124] In usage scenarios with independently encoded spatial segments such as MCTS, the above affine motion procedure needs to be restricted to avoid dependencies between MCTS.
[0125] When the MV of the sub-block motion vector field leads to a sample position of a predictor outside the spatial segment boundary, the sub-block MV is individually clipped to a sample position within the spatial segment boundary.
[0126] In addition, when the MV of the sub-block motion vector field leads to a sample position of a predictor outside the spatial segment boundary, the resulting predictor sample is discarded and a new predictor sample is derived from intra prediction techniques, for example, using an angular prediction mode derived in a neighboring adjacent already decoded sample area. The samples of the spatial segment used in this scenario may belong to neighboring blocks as well as sub-blocks surrounding the current block.
[0127] As a further alternative, the motion vector candidates are checked as to the position of the resulting predictor and whether the predictor samples belong to a spatial segment. If this is not the case, the resulting sub-block motion vector is more likely not to point to a sample position within the spatial segment. Therefore, it is advantageous to clip motion vector candidates pointing to sample positions outside the spatial segment boundary to a position within the spatial segment boundary. Alternatively, in such scenarios, the mode using the affine motion model is made unavailable.
[0128] Another concept relates to the construction of history-based motion information prediction candidates (primitives). The history-based motion information prediction candidates can be placed at the end of the aforementioned reading order 204 so as not to cause a mismatch. In other words, when adding an HMVP candidate to the motion vector candidate list, if the availability of temporally collocated MVs changes after encoding due to a change in the tile layout at decoding, the candidate list on the decoder side does not match the encoder side, and the same problem as described above, that the index after Col cannot be used when such a usage scenario is assumed, is introduced. The essence of the concept here is the same as above where, if necessary, the Col candidate is shifted to the end of the list after the HMVP.
[0129] To create the MV candidate list, the above concepts can be executed for all blocks within a tile.
[0130] Note that the placement of the history-based motion information predictor candidates within the reading order may alternatively be performed in a manner that depends on the inspection of the motion information currently included in the history list of the motion information stored in the most recent inter-prediction block. For example, some central tendencies of the motion information included in the history list, such as the median value, average value, etc., can be used by the coder and decoder to detect to what extent the history-based motion information predictor candidates may actually lead to a contradiction with tile independence for the current block. For example, if, according to a certain average, the motion information included in the history list points towards the center of the current tile, or more generally, points to a position sufficiently far from the boundary 102 of the current tile 100a from the current position of block 104, the possibility that the history-based motion information predictor candidates lead to a contradiction with tile independence is sufficiently low. Therefore, in that case, it can be assumed that the history-based motion information predictor candidates (primitives) can be placed earlier in the reading order, on average, compared to the case where the history list points to a location near or beyond the boundary 102 of the current tile 100a. In addition to the measure of central tendency, a measure of the dispersion of the motion information included in the history list can also be used. The greater the dispersion, such as due to distortion, the higher the likelihood that the history-based motion information predictor candidates may cause a contradiction with tile independence and should be placed towards the end of the candidate list 190.
[0131] Figure 21 presents a concept for improving the loading of the motion information predictor candidate list 190 with a motion information predictor candidate 500 selected from a motion information history list 502 that buffers the most recently used motion information, i.e., the set of motion information of the most recently encoded / decoded inter prediction blocks. That is, each entry 504 in the list 502 stores the motion information used for one of the most recently encoded / decoded inter prediction blocks. The buffering in the history list 502 can be performed in a first-in first-out manner by the encoder and decoder. There is also another possibility to keep the content or entries of the history list 502 up-to-date. The concept outlined with respect to Figure 21 is useful both when using tile-based independent coding and when not using tile-based independent coding. Thus, the encoder and decoder can be restricted to the individual tiles 100, i.e., the current tile 100a when inspecting the current coding block 104, for which the motion information is used to fill the entries 504 of the history list 502, or, when not using tile-based independent coding, can be restricted to the collection area of the inter prediction blocks that is the entire image. Further, the concept described with respect to Figure 21 is not necessarily used in the manner described above such that the history-based motion information predictor candidate is shifted towards the end within the motion information predictor candidate list 190, behind the spatial candidates such as at least the combined spatial candidates. Rather, the concept outlined below with respect to Figure 21 simply assumes that there are motion information predictor candidates 192 already loaded in the motion information predictor candidate list 190 at the time 500 when using it to select a history-based motion information predictor candidate from the history list 502. Finally, even when tile-based independent coding is used in the case of Figure 21, in that case, the concept of Figure 21 may or may not be combined with any of the above-described concepts where the decoder enforces tile independence with respect to the finally signaled motion information for the block 104 or with respect to the motion information predictor candidates input to the list 190. In the following description of Figure 21, it is assumed that tile-based independent coding is applied, but the decoder does not enforce tile independence with respect to the motion information predictor candidate 192.
[0132] As shown in FIG. 21, at the time of encoding / decoding the inter prediction block 104, the encoder and decoder establish a motion information predictor candidate list 190. FIG. 21 exemplarily assumes that two motion information predictor candidates 192, indicated as A and B in FIG. 21, have already been loaded into the motion information predictor candidate list 190. This "two" is, of course, just an example. In particular, at the time 500 of selecting a history-based candidate for loading the list 190 with another history-based predictor candidate using it, the number of candidates 192 already loaded in the list 190 for the block 104 depends spatially on the availability of some juxtaposed positions in other reference images and the surrounding inter prediction blocks. However, the fact that FIG. 21 exemplarily shows two candidates 192 as already existing in the list 190 at the time 500 of selecting the history-based candidates of the history list 502 does not mean that the concept of FIG. 21 is limited to cases where at least two candidate primitives precede the history-based candidate primitives in the loading order, nor does it mean that cases where less than two candidates 192 already exist in the list 190 at the time 500 of selecting the history-based candidates for the list 190 can occur.
[0133] FIG. 21 shows the situation at time 500 when deriving a history-based motion information predictor candidate primitive, i.e., selecting a history-based motion information predictor candidate that uses the next empty entry 506 in candidate list 190 to be filled. In particular, FIG. 21 shows the motion information, A and B, defined by the candidates 192 previously read into list 190 thereby, and the motion information in entry 504 of history list 502, in the form of motion vectors, indicated as 1 enclosed in a circle to 5 enclosed in a circle in FIG. 21. For ease of understanding the concept of FIG. 21, all motion vectors A, B, and 1 to 5 are illustratively indicated as referring to the same reference image 12b, but this is not necessarily the case. The motion information of each candidate 192 and entry 504 may actually include a reference image index that indexes the reference image to which the respective motion information refers, and it should be apparent to those skilled in the art that this index may be different between candidates 190 and between entries 504, respectively.
[0134] Instead of simply selecting the most recently entered motion information in the history list 502, the encoder and decoder use the motion information predictor candidates 192 that already exist in the list 190 at the time the selection 500 is made. In the case of FIG. 21, these are the motion information predictor candidates A and B. In particular, in order to make the selection 500, the encoder and decoder determine the difference for each entry 504 in the history list 502 with respect to the motion information predictor candidates 192 that have already been loaded into the list 190 before the selection 500 using it. This difference may depend, for example, on the difference in the motion vectors between each motion information 504 in one history list 502 and the (one or more) motion information predictor candidates 192 in the other list 190. If there are multiple candidates 192 in the list 190, the minimum value of the distance between each motion information 504 in one list 502 and the candidate 192 in the other list 190 may be used, for example, to determine the difference. The difference is illustrated in FIG. 21 for the motion information 1 in the list 502 using double-headed arrows. However, the difference may also depend on the difference in the reference image indexes. There are different possibilities for how to use the difference when making the selection 500. Generally speaking, the purpose is to make the selection 500 in such a way that the motion information 504 selected from the list 502 is sufficiently distinguishable from the motion information predictor candidates 192 that have been loaded into the list 190 so far using it. Therefore, the dependency is designed in such a way that the higher the likelihood of being selected for a particular motion information 504 in the history list 502, the higher the difference with respect to the motion information predictor candidates 192 that already exist in the list 190. For example, the selection 500 can be made in such a way that the selection depends on the difference and the rank of each motion information 504 for which each motion information 504 has been entered into the history list 502. For example, the arrow 508 in FIG. 21 shall indicate the order in which the motion information 504 was entered into the list 502, i.e., the upper entries in the list 502 were entered more recently than the lower entries in the list 502.According to certain embodiments, for example, the coder and decoder exclude from the history list 502 motion information entries 504 whose motion information is less than a predetermined threshold and whose motion information is different from the (one or more) motion information predictors 192 previously read into list 190 using it. To this end, the coder and decoder may use, for example, the above-described example of a distance metric that depends on the difference in motion vectors and / or the difference in reference picture indices, and threshold the differences obtained as a result for a certain threshold parameter. The result of this procedure is an exclusion of certain motion information entries 504 from list 502 such that only a subset of the motion information entries remains from which selection 500 can then be made to form the motion information predictor candidates to be input into list 190 at position 506 using the most recently input motion information entry. In the case of list 190 having another candidate position that must be read using history list 502 in rank order, this procedure may be repeated. In that case, the aforementioned threshold parameter may be lowered so as to result in a less stringent exclusion of the motion information entries within list 502. That is, it is recognized that more similar motion information results in a subset of motion information entries from list 502, from which the next motion information predictor candidate for list 190 is then selected as the most recently input motion information predictor candidate for history list 502. Of course, there are also other possibilities for making selection 500 in a manner that depends on the differences in the motion information entries 504 within list 502 with respect to the (one or more) motion information predictors 192 previously read into candidate list 190 using it.
[0135] In fact, the concept of FIG. 21 leads to a motion information predictor candidate list 190 of the input predicted block 104 that collects motion information predictors 192 that increase the likelihood that the coder will find the optimal candidate among this list 190 for excellent distortion optimization because their mutual differences are large. In that case, the coder may then signal the candidate selected from among the list 190 of block 104 in data stream 14 together with a respective pointer 193 and a specific prediction residual 194.
[0136] Of course, the foregoing possibilities are not limited to the history-based motion information predictors candidates selected from the history list 502. Rather, this process can also be used to select from another reservoir of possible motion information predictors candidates. For example, a subset of the motion information predictors candidates may be selected from the larger set of motion information predictors candidates in the manner described with respect to FIG. 21.
[0137] In other words, the state-of-the-art is a procedure for adding available HMVP candidates only on the criterion that none of the existing list candidates match the HMVP candidates that should potentially be added in the horizontal and vertical components and the reference index. To improve the quality of the HMVP candidates added to the motion vector candidates lists such as the merge candidate list and the sub-block motion vectors, the list insertion of each HMVP candidate is based on an adaptive threshold. For example, if the motion vector candidates list is already filled in the spatial neighborhood A and further candidates B and Col are not available before adding the first eligible HMVP candidate to the next available list entry, the difference between the eligible HMVP candidate and the existing entry (A in the given example) is measured. Each HMVP candidate is added only if the threshold is met; otherwise, the next eligible HMVP is tested, and so on. The above difference measure can incorporate the horizontal and vertical components of the MV and the reference index. In one embodiment, the threshold is adapted for each HMVP candidate, for example, the threshold is lowered.
[0138] Finally, FIG. 22 relates to a case where the codec uses a bidirectional optical flow tool to improve motion compensation bidirectional prediction. FIG. 22 shows a bidirectional predictive coding block 104. Two motion vectors 1541 and 1542, one related to the reference image 12b1 and the other related to the reference image 12b2, are signaled in the data stream of the block 104. The footprints of the block 104 in the reference images 12b1 and 12b2 when displaced according to the respective motion vectors 1541 and 1542 are indicated as 1601 and 1602 in FIG. 22, respectively. The BIO tool determines the luminance gradient direction across the area of the block 104 to combine two hypotheses derived from the corresponding patches 1301 and 1302 in the respective reference images 12b1 and 12b2. Using these luminance gradient directions, the BIO tool changes the way the block 104 is predicted from the two reference images 12b1 and 12b2 in a way that varies across the entire block 104. As a result of the BIO tool, the patches 1301 and 1302 are enlarged relative to the size of the block 104, not only in cases where the respective motion vectors 1541 or 1542 are sub-pixel motion vectors and / or an interpolation filter is required, but also in cases where they are full-pixel vectors. In either case, the patch is enlarged by an additional enlargement of only an n-sample width, where n is an integer with n > 0. This is the result of the operation of the BIO tool. That is, the BIO tool derives from each of the patches 1301 and 1302 respective hypothesis blocks 4021 and 4022 that are enlarged relative to the size of the block 104 by the above-mentioned n-sample width extended edge portions indicated by reference numeral 432 in FIG. 22 so as to be able to determine the local luminance gradient. To generate the enlarged hypothesis blocks 4021 and 4022, the extended edge portions 1721 and 1722 of the patches 1301 and 1302 correspond at least to the corresponding n-sample width enlargement 434 compared to the width of the block 104, and further correspond to an enlargement 436 of a width corresponding to the kernel range of the interpolation filter when the corresponding motion vectors 1541 and 1542 are sub-pixel motion vectors.Therefore, n conforms to the sample width or extension associated with the bidirectional optical flow tool. In FIG. 22, the motion vector 1542 was a sub-pixel vector such that the BIO tool derived the hypothesis block 4022 from the patch 1302 by interpolation 4352, but it is exemplarily assumed that if the motion vector 1541 is a full-pixel motion vector, the BIO tool may derive the block 4021 by sample copy 4351 only. In that case, the BIO tool depends on the local gradient to obtain the predictor 438 of the final bidirectional prediction of block 104 and linearly combines 436 in a way that varies across the area of the block to combine both hypotheses, namely, the hypothesis predictor 4021 obtained from the reference image 12b1 and the predictor 4022 obtained from the reference image 12b2, within the areas 1601 and 1602 of block 104. In particular, for each sample 442 of block 104, the corresponding samples 4401 and 4402 within blocks 4021 and 4022 are added and weighted with respective hypothesis weights such as 1 / 2 or other weights that can optionally sum to 1, and then a constant depending on the luminance gradients determined for the corresponding samples 4401 and 4402 within blocks 4021 and 4022 is added to derive each sample 442. This is done by both the encoder and decoder having the BIO tool.
[0139] As long as the enlargement 436, i.e., the enlargement by interpolation, is concerned, that is, regardless of whether the encoder enforces the tile independence constraint of the decoder regarding the motion vector predictors used to encode the motion vectors 1541 and 1542 into the data stream, neither of the patches 1301 and 1302 crosses the boundary 102, or alternatively, it may be noted that the decoder itself enforces the tile independence constraint regarding the final motion vectors 1541 and 1542. However, an extension of the patches 1301 and 1302 that crosses the boundary 102 by an amount less than the sample width n associated with the bidirectional optical flow tool will still be possible. However, both the encoder and the decoder check whether either of the patches 1301 and 1302 will still cross the boundary 102 due to the additional n-sample-width extension 434. If this is the case, the video encoder and the video decoder deactivate the BIO tool. Otherwise, the BIO tool is not deactivated. Therefore, signaling in the data stream for controlling the BIO tool is not required. In FIG. 22, it is shown that the patch 1302 crosses the boundary 102 of the current tile 100a towards the neighboring tile 100b. Therefore, the BIO tool is deactivated here so that the decoder and the encoder recognize that the patch 1302 crosses the boundary 102. As a supplement, as described above with respect to FIG. 11, it is recalled that the decoder can correspond to the n-sample enlargement 434 when clipping the motion vector 154 so that the BIO tool deactivation for tile boundary recognition is not required.
[0140] For example, when the BIO is deactivated, each sample 442, Sample(x,y), is derived as a simple weighted sum of the corresponding samples 4401 and 4402 of the hypothesis blocks 4021 and 4022, predSamplesL0[x][y] and predSamplesL1[x][y], as follows. Sample(x,y) = round(0.5 * predSamplesL0[x][y] + 0.5 * predSamplesL1[x][y]) For all (x, y) of the current prediction block
[0141] When the BIO tool is activated, this changes as follows Sample(x,y) = round(0.5 * predSamplesL0[x][y] + 0.5 * predSamplesL1[x][y] + bioEnh(x,y)) For all (x, y) of the current prediction block 104 bioEnh(x,y) is the offset calculated with the gradients of the corresponding reference samples 4401 and 4402 of each of the two references 4021 and 4022
[0142] Alternatively, the decoder uses boundary padding to fill the surroundings 432 of the footprints 402 and 402', thereby expanding them into patches 430 and 430' that extend beyond the boundary 102 to neighboring tiles
[0143] In other words, in usage scenarios with independently encoded spatial segments such as MCTS, the above BIO procedure needs to be restricted to avoid dependencies between MCTS
[0144] As part of the above embodiments, if an initial unrefined reference block is placed at the position of the boundary of the spatial segment in each image, BIO is deactivated to include samples outside the spatial segment in the necessary gradient calculation procedure
[0145] Alternatively, in this case, the outside of the spatial segment is extended by boundary padding procedures such as repetition or mirroring of the sample values at the spatial segment boundary of the spatial segment, or more advanced padding modes. This padding enables the gradient calculation to be performed without using samples from neighboring spatial segments
[0146] The following final description is made with respect to the above embodiments and concepts. As has been stated several times throughout the description of the various embodiments and concepts, those embodiments and concepts can be used and implemented individually or simultaneously in a particular video codec. Further, in many of the figures, the fact that the motion information shown therein is shown as including motion vectors is to be construed as such only as a possibility and does not limit embodiments that do not specifically utilize the fact that the motion information includes motion vectors. For example, the tile boundary-dependent motion information derivation on the decoder side described with respect to FIG. 8 is not limited to motion vector-based motion information. With respect to FIG. 9, it should be noted that the application of the decoder's tile independence constraint enforcement to the motion information shown therein is not limited to motion vector-based types of motion information nor to the predictive coding of the motion information presented in FIG. 10. When referring to FIG. 11, an example is presented where a patch can be made larger than the footprint of an inter-prediction block. However, the embodiments of the present application described with respect to the various figures, except for the embodiments specifically mentioned in that context, can be relevant to codecs where such an expansion, for example, does not occur. For example, in the embodiments of FIGS. 12a and 12b, this context was specifically mentioned. With respect to FIG. 13, a concept is presented of applying the tile independence constraint enforcement to motion information prediction or a motion information predictor. However, in FIG. 13, this concept was illustrated with respect to the use / establishment of a motion information predictor candidate list, but it should be noted that this concept can also be applied to codecs that do not establish a motion information predictor candidate list for inter-prediction blocks. That is, simply, one motion information prediction / predictor may be derived for block 104 in a manner that uses the tile independence constraint enforcement. However, in the above, FIGS. 13 and 14 represent embodiments that can be modified to refer to only one motion information predictor. Further, all the details regarding the enforcement described with respect to the foregoing figures, namely, in particular, FIGS. 11, 12a, and 12b, can be used to more specifically specify and implement the embodiments of FIGS. 13 and 14 as has already been explained above.With respect to FIG. 15, embodiments / concepts are presented according to which the availability of motion information predictor candidates is given depending on whether they cause a contradiction in the tile independence constraint. This concept can be mixed with the embodiments of FIGS. 13 and 14, for example, in that certain motion information predictor candidates (primitives) are determined using the enforcement of the tile independence constraint and some have their availability derived depending on whether each candidate causes a contradiction in the tile independence constraint. Further, the embodiment of FIG. 15 can be combined, for example, with the application of the enforcement of the tile independence constraint on the decoder side as described above with respect to FIG. 9. In all embodiments that refer to or utilize a list of motion information predictor candidates, according to which the determination of the temporal motion information predictor candidates is made in a way that depends on whether the current block is adjacent to a particular side of the current tile, and the reading order is changed depending on whether the current block is adjacent to a particular side of the current tile with respect to this temporal motion information predictor candidate, the concept of FIG. 16 can be used. FIG. 17 provides details. Again, this concept of FIGS. 16 and 17 can be combined with the application of the enforcement of the tile independence constraint on the decoder side for the final state motion information and / or the (one or more) predictors, as well as the availability control according to FIGS. 12 and 14. The latter is also the concept of deriving temporal motion information predictor candidates according to FIG. 18. This concept is combinable with all the concepts stated to be combinable with the concepts of FIGS. 16 and 17, and it is even possible to combine on the one hand with the concepts of FIGS. 16 and 17 and on the other hand with the concept of FIG. 18. Since FIG. 19 represents an alternative view of FIG. 18, the explanations made with respect to FIG. 18 are also relevant to FIG. 19. FIGS. 18 and 19 can be used especially when only one motion information predictor for a block is determined, i.e., without the establishment of a candidate list. With respect to FIG. 20, note that the mode underlying this concept, i.e., the affine motion model mode, can represent, for example, one mode of a video codec other than the modes described for all the other embodiments described herein, or can represent the only inter-prediction mode of a video codec.In other words, the motion information represented by vectors 306a and 306b may be input into the candidate list mentioned with respect to any other embodiment, or may represent the only predictor of block 104. The variability increase concept for increasing the variability of the motion information included in the candidate list presented with respect to FIG. 21 can, as described above, be combined with any of the other embodiments, and in particular is not limited to history-based motion information predictor candidates. Similar descriptions have been made with respect to the embodiment of the BIO tool described with respect to FIG. 21 and with respect to the embodiment of the affine motion model of FIG. 20. For all embodiments described herein, when multiple motion information predictor candidates are derived for block 104, as shown with respect to FIG. 21, they do not necessarily refer to the same reference picture 12b, and it should be noted that this is not only applicable to the B prediction block outlined in FIG. 22. Further, in the above description, mainly, the boundary between neighboring tiles 100a and 100b, i.e., the boundary through the image, was focused on. However, it should be noted that the above details can also be transferred to the tile boundary that coincides with the image boundary, i.e., the boundary of the tile adjacent to the outer periphery of the image. For these boundaries as well, for example, the enforcement of the tile independence constraint on the decoder side described above can be established, or the availability constraint can be applied, and the same applies hereinafter. Still further, in all of the above embodiments, the video codec may include a flag indicating whether or not that type of tile independence constraint is used in the encoding of the video in the data stream. That is, the video codec can signal to the decoder whether the decoder will apply a particular constraint and whether the encoder has applied the monitoring of its encoder-side constraints.
[0147] Although some aspects are described in the context of an apparatus, it will be apparent that these aspects also represent corresponding methods where a block or device corresponds to a method step or a feature of a method step. Similarly, aspects described in the context of a method step also represent corresponding blocks or items or features of a corresponding apparatus. Some or all of the method steps may be performed by (or using) a hardware device such as, for example, a microprocessor, a programmable computer, or an electronic circuit. In some embodiments, one or more of the most important method steps may be performed by such a device.
[0148] The encoded video signal or data stream according to the present invention can be stored respectively in a digital storage medium or transmitted via a transmission medium such as a wireless transmission medium or a wired transmission medium like the Internet.
[0149] Depending on specific implementation requirements, embodiments of the present invention can be implemented in hardware or software. The embodiments can be executed using a digital storage medium storing electronically readable control signals that cooperate (or can cooperate) with a programmable computer system such that each method is performed, for example, a floppy disk, a DVD, a Blu-Ray, a CD, a ROM, a PROM, an EPROM, an EEPROM, or a flash memory. Thus, the digital storage medium may be computer-readable.
[0150] Some embodiments according to the present invention include a data carrier having electronically readable control signals that can cooperate with a programmable computer system such that one of the methods described herein is performed.
[0151] In general, embodiments of the present invention can be implemented as a computer program product having program code, which, when the computer program product operates on a computer, causes one of the methods to be performed. The program code can be stored, for example, on a machine-readable carrier.
[0152] Other embodiments include a computer program for performing one of the methods described herein, stored on a machine-readable carrier.
[0153] In other words, one embodiment of the method of the present invention is thus a computer program having program code for performing one of the methods described herein when the computer program operates on a computer.
[0154] Another embodiment of the method of the present invention is thus a data carrier (or digital storage medium, or computer-readable medium) on which a computer program for performing one of the methods described herein is recorded. The data carrier, digital storage medium, or recorded medium is typically tangible and / or non-transitory.
[0155] Another embodiment of the method of the present invention is thus a data stream or sequence of signals representing a computer program for performing one of the methods described herein. The data stream or sequence of signals can be configured to be transferred, for example, via a data communication connection, such as via the Internet.
[0156] Another embodiment includes processing means, such as a computer or a programmable logic device, configured or adapted to perform one of the methods described herein.
[0157] Another embodiment includes a computer on which a computer program for performing one of the methods described herein is installed.
[0158] Another embodiment according to the present invention includes an apparatus or system configured to transfer (e.g., electronically or optically) a computer program for performing one of the methods described herein to a receiver. The receiver may be, for example, a computer, a mobile device, a memory device, etc. The apparatus or system may include, for example, a file server for transferring the computer program to the receiver.
[0159] In some embodiments, a programmable logic device (e.g., a field programmable gate array) may be used to perform some or all of the functions of the methods described herein. In some embodiments, a field programmable gate array may cooperate with a microprocessor to perform one of the methods described herein. Generally, these methods are preferably performed by any hardware device.
[0160] The apparatuses described herein may be implemented using a hardware device, or using a computer, or using a combination of a hardware device and a computer.
[0161] The apparatuses described herein, or any component of the apparatuses described herein, may be implemented at least partially as hardware and / or as software.
[0162] The methods described herein may be performed using a hardware device, or using a computer, or using a combination of a hardware device and a computer.
[0163] The methods described herein, or any component of the apparatuses described herein, may be implemented at least partially by hardware and / or by software.
[0164] The above-described embodiments are merely illustrative of the principles of the present invention. It should be understood by those skilled in the art that modifications and variations of the configurations and details described herein will become apparent. Therefore, it is intended to be limited only by the imminent claims, rather than by the specific details presented as descriptions and explanations of the embodiments herein.
Claims
1. A video decoder comprising at least one processor and a memory, wherein when executed by the at least one processor, the memory causes the video decoder to perform the following operations: identifying the position of a current block of a current image, where the current image is one of a sequence of images presented in temporal order; deriving motion information for the current block by adding a temporal motion vector MVTemp from already decoded blocks of the current image to the position of the current block; clipping the motion information to identify collocated blocks within an independently encoded spatial region of a reference image that precedes the current image in the temporal presentation order; identifying the position of the collocated blocks within the reference image based on the clipped motion information; determining a motion vector for the collocated blocks in the reference image; determining a predicted motion vector for the current block based on the motion vector from the collocated blocks. The video decoder includes instructions for performing the above operations.
2. The video decoder according to claim 1, wherein the motion information is clipped such that it is contained within the corresponding CTU boundary of the current block.
3. The video decoder according to claim 1, wherein when executed by the at least one processor, the memory further causes the video decoder to scale the motion vector from the collocated blocks according to a temporal difference between the included images.
4. The video decoder according to claim 1, wherein when executed by the at least one processor, the memory further causes the video decoder to decode a reference image index from a data stream for the reference image.
5. The video decoder according to claim 1, wherein when executed by the at least one processor, the memory further causes the video decoder to decode a motion vector prediction residual from a data stream.
6. The memory further includes instructions that, when executed by the at least one processor, cause the video decoder to select the predicted motion vector as a candidate selected from a motion vector candidate list based on an index from the data stream. The video decoder according to claim 1.
7. The memory further includes instructions that, when executed by the at least one processor, cause the video decoder to add the motion vector prediction residual from the data stream to the selected candidate that generates the motion vector for the current block. The video decoder according to claim 6.
8. Identifying a position of a current block of a current image, where the current image is one of consecutive images in a temporal presentation order, the identifying, Deriving motion information for the current block by adding a temporal motion vector MVTemp from already decoded blocks of the current image to the position of the current block, Clipping the motion information to identify collocated blocks within an independently encoded spatial region of a reference image that precedes the current image in the temporal presentation order, Identifying a position of the collocated blocks in the reference image based on the clipped motion information, Determining a motion vector for the collocated blocks in the reference image, Determining a predicted motion vector for the current block based on the motion vector from the collocated blocks. A method for video decoding, including.
9. The method according to claim 8, wherein the motion information is clipped to be included within a corresponding CTU boundary of the current block.
10. The method according to claim 8, further including scaling the motion vector from the collocated blocks according to a temporal difference of the included images.
11. The method according to claim 8, further including decoding a reference image index from a data stream for the reference image.
12. The method according to claim 8, further including decoding a motion vector prediction residual from a data stream.
13. selecting, as a candidate selected from a motion vector candidate list, the predicted motion vector based on an index from a data stream, the method of claim 8, further comprising.
14. the method of claim 13, further comprising adding, for the current block, a motion vector prediction residual from the data stream to the selected candidate that generates a motion vector.
15. A non-transitory digital storage medium storing computer program instructions, which, when executed, cause at least one processor to identify the position of a current block of a current image, the current image being one of a sequence of images presented in a temporal presentation order, said identifying; deriving motion information for the current block by adding a temporal motion vector MVTemp from already decoded blocks of the current image to the position of the current block; clipping the motion information to identify collocated blocks within an independently encoded spatial region of a reference image that precedes the current image in the temporal presentation order; identifying the position of the collocated blocks in the reference image based on the clipped motion information; determining a motion vector for the collocated blocks in the reference image; determining a predicted motion vector for the current block based on the motion vector from the collocated blocks, a non-transitory digital storage medium.
16. The non-transitory digital storage medium of claim 15, wherein the motion information is clipped to be included within a corresponding CTU boundary of the current block.
17. The non-transitory digital storage medium of claim 15, wherein the computer program further comprises instructions to scale the motion vector from the collocated blocks according to a temporal difference of the included images.
18. The non-transitory digital storage medium of claim 15, wherein the computer program further comprises instructions to decode a motion vector prediction residual from a data stream.
19. The non - transitory digital storage medium according to claim 15, wherein the computer program further includes an instruction to select the predicted motion vector as a candidate selected from a motion vector candidate list based on an index from a data stream.
20. The non - transitory digital storage medium according to claim 19, wherein the computer program further includes an instruction to add a motion vector prediction residual from the data stream to the selected candidate that generates a motion vector for the current block.
Citation Information
Patent Citations
Method for inducing a temporally predictive motion vector and apparatus using such method
JP2014520477A
Video coding device and video decoding device
JP2020145484A
Sub-block motion derivation and decoder-side motion vector refinement for merge mode
JP2021502038A