Inter prediction with slice independent constraints
By using the slice boundary awareness mechanism of the decoder to process motion information, the problem of coding efficiency loss in slice-by-slice coding is solved, and a more efficient coding process is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2019-11-25
- Publication Date
- 2026-03-31
AI Technical Summary
In existing technologies for slice-based video coding, the independent coding of slices results in significant coding efficiency loss, and the codec behavior involves numerous modifications along the boundaries, further impacting coding efficiency.
The decoder uses a slice boundary awareness mechanism to identify and process motion information, ensuring that the motion information meets the slice independence requirements. The encoder uses the decoder's awareness mechanism to select the optimal motion information to reduce bit rate and avoid inter-slice dependency interruptions.
It effectively reduces the encoding efficiency loss caused by slice boundaries, improves encoding efficiency, and simplifies the encoding and decoding process.
Smart Images

Figure CN113366834B_ABST
Abstract
Description
Technical Field
[0001] This application relates to inter-coding concepts used in block-based codecs, such as hybrid video codecs, and more particularly to the concept of allowing slice-based coding, i.e., the independent coding of slices into the spatially subdivided video. Background Technology
[0002] Existing applications, such as 360° video services based on the MPEG OMAF standard, heavily rely on spatial video partitioning or segmentation techniques. In such applications, spatial video segments are transmitted to the client and jointly decoded in a manner adjusted to the client's current viewing orientation. Another related application that relies on spatial segmentation of the video plane is the parallelization of encoding and decoding operations, for example, to facilitate the multi-core capabilities of modern computing platforms.
[0003] One such spatial segmentation technique, implemented in HEVC, is called tiling, which divides the image plane into segments forming a rectangular grid. These spatial segments are encoded independently of entropy coding and intra-prediction. Furthermore, there are methods to indicate that spatial segments are also encoded independently of state-of-the-art inter-prediction. For some of the applications listed above, constraints on all three domains—entropy coding, intra-prediction, and inter-prediction—are crucial.
[0004] However, with the continuous development of video coding technology, new coding tools have emerged, many of which are related to the domain of inter-prediction. That is, many tools incorporate new dependencies into different regions within previously coded images or currently coded images. Appropriate attention must be paid to how to ensure independence in all of these mentioned domains.
[0005] Up to this point, the encoder is responsible for setting the encoding parameters using the available encoding tools mentioned earlier in a way that adheres to the independent encoding of video slices. The decoder "depends" on the corresponding guarantees signaled to it by the encoder via the bitstream.
[0006] There is a worthwhile concept at hand that enables encoding independent of slices in a way that results in less coding efficiency loss due to the disruption of coding dependencies caused by slice partitioning; however, it only leads to marginal modifications to the codec behavior along the boundaries.
[0007] Therefore, I want to have a concept at hand that allows video coding to be done in a way that enables slice-independent coding, however, by reducing the coding efficiency loss associated with slice-dependent interruptions, although all of this is done by only slightly modifying the codec behavior along the slice boundaries. Summary of the Invention
[0008] This objective is achieved through the subject matter of the independent claims of this application.
[0009] Generally speaking, the discovery of this application is that a more efficient way has been found whereby the obligation to encode video slices independently is partially inherited from the encoder to the decoder, or in other words, is partially shared by the decoder so that the encoder can utilize this shared attention to allow slice-independent encoding of the video material. More precisely, according to embodiments of this application, the decoder has slice boundary awareness. That is, the decoder operates in a manner that depends on the boundary positions between the slices into which the video is spatially divided. In particular, this slice boundary awareness is also related to the motion information derived by the decoder from the data stream. This "awareness" causes the decoder to identify motion information signaled in the data stream, which, if applied according to the signal notification, would result in a violation of the slice independence requirement, and thus causes the decoder to map such signaled motion information, which would violate slice independence, allowing motion information states corresponding to the motion information, when used for inter-prediction, not to violate slice independence. The encoder may rely on this behavior, i.e., be aware of the decoder's awareness, and especially due to the redundancy of signalable motion information states resulting from the decoder's adherence to or enforcement of slice independence constraints. Specifically, the encoder can leverage the slice-independent constraints enforced / obeyed by the decoder to select the less bit rate required for the signalable motion information states that result in the same motion information on the decoder side due to decoder behavior, such as the bit rate associated with, for example, the motion information prediction residual with zero signal. Therefore, on the encoder side, the motion information of a specific prediction block is determined to satisfy the constraint that the block predicted from the specific prediction block lies within and does not cross the boundary of the slice composed of the specific prediction blocks, i.e., the specific prediction block lies within it. However, when encoding the motion information of the specific prediction block into the data stream, the encoder utilizes the fact that it performs a derivation from the data stream based on the slice boundaries, i.e., the slice boundary awareness just outlined is required.
[0010] According to embodiments of this application, motion information includes motion vectors for specific inter-prediction blocks, and slice boundary awareness processed by the decoder is related to the motion vectors. Specifically, according to embodiments of this application, the decoder implements slice independence constraints relative to the predictive encoded motion vectors. That is, according to these embodiments, the decoder adheres to or enforces constraints on the motion vectors such that blocks from which inter-prediction blocks are to be predicted exceed the boundaries of the slices to which the inter-prediction blocks are included. When determining motion vectors, the decoder predicts based on both the motion vector predictor / prediction and the motion information from the residuals transmitted in the data stream of the inter-prediction blocks. In other words, the decoder performs the aforementioned adherence / enforcement by using an irreversible mapping: instead of mapping all combinations of motion information predictions and motion information prediction residuals to their sum to produce the final motion vectors, this mapping redirects all possible combinations of motion vector predictions and motion vector prediction residuals, the sum of which results in motion vectors associated with blocks exceeding the current slice boundary, i.e., the boundary of the slice containing the current inter-prediction block, towards the slices to which they are associated, and motion vectors that do not exceed the boundary of the current slice. In this way, the encoder can utilize the signal to notify ambiguities when a specific motion vector of a predetermined inter-prediction block is used, and can, for example, select the signal notification for the motion vector prediction residual that results in the lowest bit rate for this inter-prediction block. For example, this might be a motion vector difference of zero.
[0011] A variation of the above-described idea, which provides the decoder with at least partial slice boundary awareness, implements slice independence constraints. For this, the encoder is otherwise solely responsible, and the decoder applies the slice independence constraints to one or more motion information predictors for a specific inter-prediction block, rather than to the resulting motion information due to the combination of motion information prediction and motion information prediction residuals. Both the encoder and decoder perform slice independence on the motion information predictors so that both use the same motion information predictors. Signal ambiguity and the possibility of utilizing the latter to minimize bit rate are not issues here. However, preparing the motion information predictors for a specific inter-prediction block in advance—that is, customizing or “focusing” the available motion vector predictors for that specific inter-prediction block before using them for motion information prediction encoding / decoding—points only to block locations within the current slice, instead of wasting one or more motion information predictors pointing to conflicting block locations, i.e., block locations beyond the current slice boundary, and using them to predictively encode the motion information of the inter-prediction block, will require signaling of non-zero motion information prediction residuals in order to redirect conflicting motion information predictors to block locations within the current slice. Even here, the motion information can be motion vectors.
[0012] Related to, but still distinct from, the further embodiments of this application aim to avoid directly applying motion information prediction candidates for specific inter-prediction blocks, i.e., having zero motion information prediction residuals, which would compromise slice independence constraints. Specifically, in addition to the previous variants, such motion information prediction candidates will simply not be used to populate the motion information prediction candidate list for the currently predicted inter-prediction block. The encoder and decoder function identically. No redirection is performed. These predictors are simply ignored. The motion information prediction candidate list is built in the same inter-slice boundary-aware manner. In this way, all signalable members for the currently encoded inter-prediction block will be focused on non-conflicting motion information prediction candidates. Therefore, the complete list can be signaled at a point in the data stream, for example, in cases where no signalable motion information prediction candidate is "wasted" in the signalable state of such an indicator, where the slice independence constraint pointed to by the motion information prediction candidate would conflict with this constraint or any preceding motion information prediction candidate in the hierarchy, in order to comply with the slice independence constraint.
[0013] Similarly, further embodiments of this application aim to avoid filling the motion information prediction candidate list with candidates whose origin resides in blocks located outside the current block, i.e., slices including the current inter-prediction block. Therefore, and according to these embodiments, the decoder and encoder check whether the inter-prediction block is adjacent to a predetermined edge, such as the bottom and / or right-hand edge of the current slice. If so, a first block in the motion information reference image is identified, and the list is filled with motion information prediction candidates derived from the motion information of this first block. If not, a second block in the motion information reference image is identified, and the motion information prediction candidate list is filled with motion information prediction candidates derived from the motion information of the second block. For example, the first block may be a block inheriting the same position as a first alignment position within the current inter-prediction block, while the second block is a block containing the same position as a second predetermined position outside the inter-prediction block, i.e., an offset relative to the inter-prediction block in a direction perpendicular to the predetermined edge.
[0014] A further variation of the ideas outlined above in this application involves constructing / establishing the motion information prediction candidate list in a manner that moves motion information prediction candidates whose origins are prone to conflict with slice-independent constraints to the end of the list, where the origin is located outside the current slice. In this way, signaling indicators to the motion information prediction candidate list on the encoder side is not subject to many constraints. In other words, an indicator sent and signaled to a motion information prediction candidate for a specific inter-prediction block is actually used for the current inter-prediction block, and the indicator indicates the actual motion information candidate to be used by virtue of its rank position in the motion information prediction candidate list. By moving motion information prediction candidates that might not be available in the list because their origins are outside the current slice, so that they appear at the end of the list or at least later, i.e., at a higher rank, all motion information prediction candidates in the list preceding the latter can still be signaled by the encoder, and therefore, they are available for prediction. The motion information prediction candidate library for the current block remains large compared to moving such "problematic" motion information prediction candidates to the end of the list without moving them. According to some embodiments relating to the aspects just outlined, populating the motion prediction candidate list by moving "problematic" motion prediction candidates to the end of the list is performed in a slice boundary-aware manner at both the encoder and decoder. In this way, the slight coding efficiency loss associated with this potentially more efficient movement of motion prediction candidates to the end of the list is limited to regions of the video image along slice boundaries. However, according to an alternative, moving "problematic" motion prediction candidates to the end of the list is performed regardless of whether the current block is located along any slice boundary. While slightly reducing coding efficiency, the latter alternative may improve robustness and simplify the encoding / decoding process. The "problematic" motion prediction candidates may be those derived from blocks in a reference image or from motion information history management.
[0015] According to a further embodiment of this application, by applying slice-independent constraints to the predicted motion vectors, the decoder and encoder determine temporal motion information prediction candidates in a slice boundary-aware manner. The predicted motion vectors are sequentially used to point to blocks in the motion information reference image, and the motion information of the blocks is used to form a list of temporal motion information prediction candidates. The predicted motion vectors are cropped to remain within the current slice or point to a position within the current slice. Therefore, the availability of such candidates is guaranteed within the current slice, thereby maintaining conformity with the list of slice-independent constraints. According to an alternative concept, if the first motion vector points outside the current slice, then the second motion vector is used instead of the cropped motion vector. That is, if the first motion vector points outside the slice, the second motion vector is used to locate blocks based on the motion information forming the temporal motion information prediction candidates.
[0016] According to a further embodiment, the idea outlined above of providing slice boundary awareness to the decoder to help implement slice-independent constraints supports a decoder for motion-compensated prediction. This decoder, based on its motion information encoded in the data stream of specific prediction blocks, derives the motion vectors of each sub-block from which the prediction block is divided into sub-blocks. The sub-block motion vectors are derived, or, based on the boundary positions between slices, the derived motion vectors are used to predict each sub-block, or both. In this way, the number of cases where the encoder cannot use this efficient encoding mode due to conflicts with slice-independent constraints is greatly reduced.
[0017] According to a further aspect of this application, a codec that supports motion-compensated bidirectional prediction and includes bidirectional optical flow tools at the encoder and decoder to improve motion-compensated bidirectional prediction conforms to slice-independent coding. Slice-independent coding, in cases where the application of the tool would cause a conflict with slice-independent constraints, determines the region of the block of prediction blocks between specific bidirectional predictions using the bidirectional optical flow tools, located outside the current slice, by providing automatic deactivation of the bidirectional optical flow tools to the encoder and decoder, or by using boundary padding.
[0018] Another aspect of this application relates to populating a motion information predictor candidate list using a motion information history list that stores previously used motion information. This aspect can be used regardless of whether slice-based coding is used. This aspect seeks to provide a more compression-efficient video codec by presenting the selection of motion information predictor candidates from the motion information history list, by selecting the entries to be populated in the candidate list, depending on those motion information predictor candidates that have been populated so far. The aim of this dependency is to select motion information entries in the history list of motion information predictor candidates that are more likely to have motion information far removed from those that have been populated so far. Appropriate distance measurements can be defined, for example, based on motion vectors composed of motion information entries in their history lists and motion information predictor candidates in the candidate list, and / or reference image indices composed of them. In this way, populating the candidate list using history-based candidates results in a higher degree of "refreshing" of the resulting candidate list, making it more likely that the encoder will find good candidates in the candidate list for large distortion optimization, i.e., those that have recently been added to the motion information history list, compared to selecting history-based candidates purely based on their rank in the motion information history list. This concept can also be applied to any other motion information predictor candidate that needs to be selected from a set of motion information predictor candidates to populate the motion information predictor candidate list.
[0019] Advantages of the invention are the subject of the dependent claims. Attached Figure Description
[0020] The preferred embodiments of this application will now be described with reference to the accompanying drawings, wherein:
[0021] Figure 1 A block diagram of a block-based video encoder is shown as an example of an encoder in which an inter-prediction concept can be implemented according to embodiments of this application;
[0022] Figure 2 Showing suitable Figure 1 A block diagram of a block-based video decoder of an encoder, as an example of an encoder in which the inter-prediction concept according to embodiments of this application can be implemented;
[0023] Figure 3 A schematic diagram illustrating the relationship between the predicted residual signal, the predicted signal, and the reconstructed signal is shown to illustrate the possibility of setting subdivisions for coding mode selection, transform selection, and transform performance, respectively.
[0024] Figure 4 A schematic diagram illustrating an exemplary division of a video into slices and the overall goal of slice-by-slice encoding is shown.
[0025] Figure 5 A schematic diagram illustrating the affine motion model of JEM (Joint Exploratory Encoding / Decoding Mode) is shown.
[0026] Figure 6 A schematic diagram illustrating the ATMVP (Optional Temporal Motion Vector Prediction) procedure is shown;
[0027] Figure 7 A schematic diagram illustrating the optical flow trajectory used in the BIO (Bidirectional Optical Flow) tool is shown;
[0028] Figure 8 The diagram illustrates how a decoder, according to some embodiments of the present application, derives slice-aware or slice-aware motion information about a predetermined inter-prediction block and how an encoder utilizes the behavior of this decoder to make the slice-independent codec more efficient; and the diagram illustrates how a decoder's compliance / enforcement characteristics are shown according to a corresponding embodiment.
[0029] Figure 9 A schematic diagram illustrating the compliance / enforcement features of a decoder according to some embodiments of this application and the use of an irreversible mapping for implementing associated redirection of motion data is shown.
[0030] Figure 10 A schematic diagram illustrating the encoding / decoding of predicted motion information is shown to illustrate the possibility of applying the compliance / enforcement characteristics of the decoder to signal notification or the final reconstructed motion information state or motion information prediction;
[0031] Figure 11 A schematic diagram illustrates the potential for changes in block size and / or enlargement in block size due to the combination of mathematical samples to compute individual samples of inter-prediction blocks;
[0032] Figure 12a and Figure 12b This diagram illustrates different possibilities for achieving motion redirection or irreversible mapping, depending on the motion accuracy-dependent activation of the interpolation filter and the size of the motion accuracy-dependent block. The aim is to allow the blocks to... Figure 12a In the case of getting as close as possible to the boundary of the current slice, and in Figure 12b To achieve a safe boundary for motion vectors under certain conditions;
[0033] Figure 13 A schematic diagram illustrating the possibility of applying the decoder's compliance / enforcement function to the resulting motion information prediction is shown, along with an optional construction of the motion information predictor candidate list.
[0034] Figure 14 A schematic diagram illustrating the irreversible motion vector mapping involved in the decoder's compliance / enforcement function and an exemplary set of motion information predictor candidates obtained when the motion information predictor candidate list is optionally interpreted in addition;
[0035] Figure 15 A schematic diagram illustrates the possibility of presenting candidate availability in a list of populated motion information predictors based on whether a particular motion information predictor candidate conflicts with slice independence constraints.
[0036] Figure 16 This diagram illustrates the possibility of changing the reference block identification based on temporal motion information predictor candidates derived from whether the inter-predicted block is adjacent to a specific slice edge.
[0037] Figure 17 A schematic diagram illustrating the filling order in the list of filling motion information predictors based on whether the inter-prediction block is adjacent to certain slice edges is shown, here relating to the juxtaposition of temporal motion information predictor candidates and one or more spatial motion information predictor candidates;
[0038] Figure 18 A schematic diagram illustrating the possibility of applying the decoder's compliance / enforcement function to predict motion vectors used to derive candidates for the temporal motion information predictor is shown.
[0039] Figure 19 Explanation is shown Figure 18 A schematic diagram of an alternative to the concept, in which an alternative predicted motion vector is used when the first predicted motion vector conflicts with the slice-independent constraint;
[0040] Figure 20 A schematic diagram illustrating the motion information derived from the sub-blocks of the current inter-predicted block based on an affine motion model is shown to illustrate the different concepts of the decoder that assist in achieving slice independence constraints.
[0041] Figure 21 A schematic diagram illustrating the selection of historical motion information candidates to improve the resulting candidate list by considering the motion information of candidates that have been used to fill the list so far; and
[0042] Figure 22 A schematic diagram illustrating the functionality of BIO (Bidirectional Optical Flow) tools embedded in the encoder and decoder is shown to illustrate different concepts of the decoder that assist in achieving slice independence constraints. Detailed Implementation
[0043] The following description of the accompanying figures begins with the presentation of an encoder and decoder for encoding images in a block-based predictive codec for video, in order to form an example of an encoding framework for an embodiment of an inter-predictive codec. The preceding encoder and decoder are relative to... Figures 1 to 3 The following is a description of embodiments of the inter-prediction concepts of this application. These concepts can be combined or used in combination. In particular, all concepts described later can be individually incorporated into... Figure 1 and 2 In the encoder and decoder, although with subsequent... Figure 4 The embodiments have been described, and encoders and decoders can then be used to form non-based Figure 1 and Figure 2 The encoder and decoder operate within an encoder-decoder-based encoding framework.
[0044] Figure 1 An apparatus is shown for predictively encoding a video 11 consisting of an image sequence 12 into a data stream 14 using exemplary transform-based residual coding. The apparatus or encoder is indicated by reference numeral 10. Figure 2 The corresponding decoder 20 is shown, i.e., the device 20 is configured to also use transform residual decoding to predictively decode the video 11' consisting of the image sequence 12' from the data stream 14, wherein the apostrophe has been used to indicate that the video 11' and the image 12' reconstructed by the decoder 20 deviate from the image 12 originally encoded by the device 10 in terms of the coding loss introduced by the quantization of the predictive residual signal. Figure 1 and Figure 2 Transform-based prediction residual coding is used as an example, but embodiments of this application are not limited to this prediction residual coding. Regarding... Figure 1 and 2The same applies to other details described below.
[0045] Encoder 10 is configured to perform a spatial-to-spectral transformation on the prediction residual signal and encode the resulting prediction residual signal into data stream 14. Similarly, decoder 20 is configured to decode the prediction residual signal from data stream 14 and perform a spectral-to-spatial transformation on the resulting prediction residual signal.
[0046] Internally, encoder 10 may include a prediction residual signal former 22 that generates a prediction residual 24 to measure the deviation of the prediction signal 26 from the original signal, i.e., the current image 12. The prediction residual signal former 22 may, for example, be a subtractor that subtracts the prediction signal from the original signal, i.e., the current image 12. Encoder 10 then further includes a transformer 28 that subjects the prediction residual signal 24 to a space-to-spectrum transformation to obtain a spectral domain prediction residual signal 24', which is then quantized by a quantizer 32, which is composed of encoder 10. The quantized prediction residual signal 24' is thus encoded into bitstream 14. For this purpose, encoder 10 may optionally include an entropy encoder 34 that entropy-encodes and quantizes the prediction residual signal into datastream 14. Prediction residual 26 is generated by prediction stage 36 of encoder 10 based on the prediction residual signal 24'' decoded into and decodeable from datastream 14. For this purpose, prediction stage 36 may be internally configured, such as... Figure 1 As shown, the system includes a dequantizer 38 that dequantizes the prediction residual signal 24” to obtain a spectral domain prediction residual signal 24”', which corresponds to signal 24', except for quantization loss. This is followed by an inverse transformer 40 that performs an inverse transform on the latter prediction residual signal 24”', i.e., a spectral-to-spatial transformation, to obtain a prediction residual signal 24”', which corresponds to the original prediction residual signal 24, except for quantization loss. Then, a combiner 42 of prediction stage 36 recombines the prediction signal 26 and the prediction residual signal 24”', such as by addition, to obtain a reconstructed signal 46, i.e., a reconstruction of the original signal 12. The reconstructed signal 46 may correspond to signal 12'. Then, prediction module 44 of prediction stage 36 generates the prediction signal 26 based on signal 46 using, for example, spatial prediction (i.e., intra-prediction) and / or temporal prediction (i.e., inter-prediction).
[0047] Similarly, decoder 20 can internally consist of components corresponding to prediction stage 36 and interconnected in a manner corresponding to prediction stage 36. Specifically, the entropy decoder 50 of decoder 20 can perform entropy decoding on the quantized spectral domain prediction residual signal 24” from the data stream. Therefore, the dequantizer 52, inverse transformer 54, combiner 56, and prediction module 58, interconnected and cooperating in the manner described above with respect to prediction stage 36, recover the reconstructed signal based on the prediction residual signal 24”, thereby achieving... Figure 2As shown, the output of combiner 56 produces the reconstructed signal, i.e., image 12'.
[0048] Although not specifically described above, it is clear that encoder 10 can set some coding parameters, including prediction modes and motion parameters, according to some optimization schemes, such as optimizing some rate and distortion-related standards, i.e., coding costs. For example, encoder 10, decoder 20, and corresponding modules 44 and 58 can support different prediction modes, such as intra-coding modes and inter-coding modes. The granularity of switching between these prediction mode types by encoder and decoder can correspond to subdividing images 12 and 12' into coding segments or coding blocks, respectively. For example, based on these coding segments, the image can be subdivided into intra-coded blocks and inter-coded blocks. Intra-coded blocks are predicted based on the spatial and encoded / decoded neighborhoods of the corresponding blocks. Multiple intra-coding modes can exist and be selected for each intra-coding segment including directional or angular intra-coding modes. The corresponding segments are filled into the corresponding intra-coding segments by extrapolating sample values of the neighborhood along a specific direction according to these modes. For example, the intra-coding mode may also include one or more further modes, such as a DC coding mode, according to which the prediction of the corresponding intra-coded block assigns DC values to all samples within the corresponding intra-coded segment, and / or an in-plane coding mode, according to which the prediction of the corresponding block is approximated or determined as a spatial distribution of sample values described by a two-dimensional linear function at the sample locations of the corresponding intra-coded block, based on neighborhood samples, resulting in the tilt and offset of a plane defined by a two-dimensional linear function. In contrast, inter-coded blocks can be predicted, for example, temporally. For inter-coded blocks, motion information can be signaled within the data stream: the motion information may include a vector indicating the spatial displacement of a portion of the previously encoded image of the video to which image 12 belongs, where the previously encoded / decoded image was sampled to obtain the prediction signal for the corresponding inter-coded block. More complex motion models may also be used. This means that, in addition to the residual signal encoding of data stream 14, such as the entropy-coded transform coefficient level representing the quantized spectral domain prediction residual signal 24", data stream 14 may have already encoded therein coding mode parameters for assigning coding modes to various blocks, prediction parameters for some blocks, motion parameters for inter-coded blocks, and optional further parameters, such as parameters for controlling and signaling the subdivision of images 12 and 12' into blocks, respectively. Decoder 20 uses these parameters to subdivide the image in the same way as encoder, assigning the same prediction modes to blocks, and performing the same predictions to produce the same prediction signal.
[0049] Figure 3This illustrates the relationship between the reconstructed signal, i.e., the reconstructed image 12', and the combination of the prediction residual signal 24"" and the prediction signal 26' signaled in the data stream on the other hand. As mentioned above, this combination can be additive. The prediction signal 26' is... Figure 3 The image region is shown as being subdivided into intra-coded blocks, illustratively represented by shading, and inter-coded blocks, illustratively represented without shading. The subdivision can be any subdivision, such as dividing the image region into blocks or rows and columns of blocks, or subdividing the multi-tree structure of image 12 into leaf blocks of different sizes, such as quadtree subdivisions, etc., where their mixture is as follows: Figure 3 As shown, the image region is first subdivided into rows and columns of root blocks, and then further subdivided according to recursive multi-tree subdivision. Similarly, for inner coding blocks 80, data stream 14 may have inner coding modes encoded therein, which assign one of several supported inner coding modes to the corresponding inner coding block 80. For inter coding blocks 82, data stream 14 may have motion information encoded therein, such as motion information including one or more motion vectors. Details are set forth below. In general, inter coding blocks 82 are not limited to being temporally encoded. Alternatively, inter coding blocks 82 may be any block predicted from previously encoded portions beyond the current image 12 itself, such as previously encoded images of the video to which image 12 belongs, or, in the case that the image encoder and decoder are scalable encoders and decoders respectively, another view or a lower layer of hierarchy. Figure 3 The prediction residual signal 24” is also shown as subdividing the image region into blocks 84. These blocks can be referred to as transform blocks to distinguish them from coded blocks 80 and 82. In fact, Figure 3 The encoder 10 and decoder 20 are shown to be able to use two different subdivisions of images 12 and 12' into blocks, namely, one subdivision into coding blocks 80 and 82, and another subdivision into block 84. The subdivisions may be the same, i.e., each coding block 80 and 82 can simultaneously form transform block 84, but... Figure 3 This illustrates a situation where, for example, subdividing into transform blocks 84 forms an extension of subdividing into coding blocks 80 / 82 such that any boundary between the two blocks 80 and 82 overlaps the boundary between the two blocks 84, or alternatively, each block 80 / 82 either coincides with one of the transform blocks 84 or with a set of transform blocks 84. However, subdivisions can also be determined or selected independently of each other, such that transform blocks 84 can alternatively span the block boundaries between blocks 80 / 82. Therefore, with regard to subdivision into transform blocks 84, similar statements as those made regarding subdivision into blocks 80 / 82 are true, i.e., block 84 can be the result of a recursive multi-tree subdivision of an image region, regularly subdividing it into blocks / blocks arranged in rows and columns, or a combination thereof, or any other type of block. Incidentally, it should be noted that blocks 80, 82, and 84 are not limited to square, rectangular, or any other shape.
[0050] In the embodiments described below, specific details of various embodiments are described using inter-prediction block 104 as an example. This block 104 may be one of the inter-prediction blocks 82. Other blocks mentioned in subsequent figures may be any of blocks 80 and 82.
[0051] Figure 3 The combination of prediction signal 26 and prediction residual signal 24”” is shown to directly generate reconstructed signal 12'. However, it should be noted that, according to alternative embodiments, more than one prediction signal 26 can be combined with prediction residual signal 24”” to form image 12'.
[0052] exist Figure 3 In this context, transform segment 84 should have the following meaning. Transformer 28 and inverse transformer 54 perform their transforms in units of these transform segments 84. For example, many codecs use some kind of DST or DCT for all transform blocks 84. Some codecs allow skipping transforms, such that for some transform segments 84, the prediction residual signal is directly encoded in the spatial domain. However, according to the embodiments described below, encoder 10 and decoder 20 are configured in a way that they support multiple transforms. For example, the transforms supported by encoder 10 and decoder 20 may include:
[0053] o DCT-II (or DCT-III), where DCT stands for Discrete Cosine Transform
[0054] o DST-IV, where DST represents Discrete Sine Transform
[0055] o DCT-IV
[0056] o DST-VII
[0057] oIdentity Transformation (IT)
[0058] Naturally, while transformer 28 will support all forward transform versions of these transforms, decoder 20 or inverse transformer 54 will support their respective backward or inverse versions:
[0059] o Inverse DCT-II (or Inverse DCT-III)
[0060] oInverse DST-IV
[0061] oReverse DCT-IV
[0062] oInverse DST-VII
[0063] oIdentity Transformation (IT)
[0064] The following description provides further details on how inter-prediction is implemented in encoder 10 and decoder 20. All other modes described above, such as the intra-prediction mode, can be supported additionally, individually, or in full. Residual coding can be performed in different ways, such as in the spatial domain.
[0065] As mentioned above, Figure 1-3 An example of an encoder that performs block-based video coding using one or more of the concepts outlined in more detail below has been presented. The details elaborated below can be found on [the following page]. Figure 1 An encoder or another block-based video encoder, making the latter different from... Figure 1 The encoder, for example, does not support intraprediction, or is subdivided into blocks 80 and / or 82 to differ from... Figure 3 The example in the document describes this approach, or even because this encoder does not use transform prediction residual coding to encode the prediction residual, for example, it encodes directly in the spatial domain. Similarly, the decoder according to embodiments of this application can perform block-based video decoding from data stream 14 using any inter-predictive coding concept further outlined below, but can be combined with, for example... Figure 2 The difference between the decoder 20 and the previous one is that it does not support intraprediction, or in the same case, it differs from the previous one. Figure 3 The described method subdivides the image 12' into blocks and / or derives prediction residuals not from the data stream 14 in the transform domain, but, for example, in the spatial domain.
[0066] However, before focusing on the description of the embodiments, the ability to encode video in segments and the way segments are encoded / decoded independently of each other are explained. Furthermore, a detailed description of additional encoding tools, which have not been discussed so far, is described and can be used selectively according to the embodiments outlined below, while still maintaining the independence of the encoding of each segment.
[0067] Therefore, after describing the potential implementations of block-based video encoders and video decoders, Figure 4 This is used to illustrate the concept of slice-independent coding. Slice-independent coding may only represent one of several coding settings for the encoder. That is, the encoder can be configured to follow slice-independent coding or it can be configured not to follow slice-independent coding. Naturally, video encoders may inevitably apply slice-independent coding.
[0068] Based on slice-by-slice encoding, the images in video 11 are divided into slices 100. Figure 4In this illustration, only three images 12a, 12b, and 12c from video 11 are shown, and they are exemplarily shown as being divided into six slices 100 each, although dividing into any other number of slices, such as two or more slices, is also possible. Furthermore, although... Figure 4 The method shown is to divide the slices into slices 100 such that the slices 100 within an image are arranged regularly in rows and columns, where each slice 100 is a rectangular portion of the corresponding image, and the boundaries between the slices 100 form a straight line 102 passing through images 12a to 12c. However, it should be noted that alternative solutions exist when the slices have a different shape. It is possible that, in order to encode and decode the corresponding images separately, each image 12 is divided into blocks, such as... Figure 3 As explained, the slice division is aligned with slice boundary 102, and none of the encoded blocks, prediction blocks and / or transform blocks, crosses and slice boundaries 102 are specifically located in one of the slices 100 of the corresponding image. For example, slice division can be accomplished in a manner aligned with the blocks that divide image 12 into root blocks, i.e., blocks 80 / 82 individually subdivided by the encoder through recursive multi-tree partitioning, to which block 104 is also representatively shown, and in which the individual subdivision in the data stream is signaled to the decoder.
[0069] Slice-independent encoding means the following: encoding the corresponding portions of a slice in a manner independent of any other slice within the unpositioned block 104, such as... Figure 4 Block 104 is illustrated exemplarily. This encoding independence relates not only to other slices of the same image, i.e., block 104 is a part of other slices in the same image, Figure 4 The example arrow 106 illustrates this intra-image independence: this arrow 106 is interrupted to show that the encoding of block 14 does not depend on any other slice in the same image. Encoding independence also involves encoding dependencies relative to other images: all images 12a to 12c of video 11 are divided into slices 100 in the same manner. Therefore, each slice 100 of a particular image 12a has a corresponding, co-located slice in all other images. Thus, a "slice" describes not only a spatial segment of a particular image, such as the segment containing block 104, but also... Figure 4 The image 12a is a fragment, and it also describes a spatiotemporal segment of video 11 composed of all co-positional slices of all images of the video, which is co-positioned with this slice in a particular image. Figure 4 In this context, the spatiotemporal segment is indicated by a dashed line relative to the slice containing block 104. When encoding block 104, the encoder will correspondingly restrict the encoding dependencies to images other than image 12a containing block 104, in order to avoid generating encoding dependencies on slices other than spatiotemporal segment 108. Figure 4In the image, interrupted arrows 110 pointing from the slice containing block 104 to different slices in reference image 12b are shown.
[0070] The concepts and embodiments further outlined below present the possibility of how to guarantee between the encoder and decoder how to avoid coding dependencies between different slices of different images, which would otherwise be caused, for example, by motion vectors pointing to distant locations away from co-located slices of the reference image and / or by motion information predictions derived from reference images from regions outside the boundaries of the slice to which block 104 belongs.
[0071] However, before focusing on the description of embodiments of this application, a particular encoding tool is described as an example, which can be implemented in the block-based video encoder and block-based video decoder according to embodiments of this application, and which is prone to causing slice interdependencies between different slices of different images, and thus is easily caused by using interrupt arrow 110 to... Figure 4 The constraint conflict is shown.
[0072] Existing video coding technologies heavily utilize inter-frame prediction to improve coding efficiency. For inter-prediction constrained coding, the inter-prediction process typically requires adapting spatial segment boundaries, which is achieved by coding bit-rate significant motion information prediction residuals or MV differences (relative to available motion information predictors or MV predictors).
[0073] Typical video coding techniques in the present technology rely heavily on the concept of collecting motion vector predictors in a so-called candidate list at different stages of the coding process, such as merging candidate lists, which also include hybrid variants of MV predictors.
[0074] Typical candidates added to this list are the MVs of co-occurring blocks from Temporal Motion Vector Prediction (TMVP), which are the lower-right blocks next to the currently encoded block in the same spatial location, but for most current blocks in the reference frame. The only exceptions are blocks located at image boundaries and blocks located at the bottom boundary of CTU lines (at least in HEVC), where actual co-occurring blocks are used to derive so-called center co-occurring candidates because there are no lower-right blocks within the image.
[0075] There are use cases where using such candidates can cause problems. For example, their availability can change when using slices and performing partial decoding of those slices. This happens because the TMVP used at the encoder might belong to another slice that is not currently being decoded. This process also affects all candidates after the index of the corresponding candidate. Therefore, decoding such a bitstream can lead to an encoder / decoder mismatch.
[0076] Video coding standards such as AVC or HEVC rely on high-precision motion vectors with a resolution higher than integer pixels. This means that motion vectors can reference samples from previous images, which are located at subsample locations rather than integer sample locations. Therefore, video coding standards define an interpolation process to derive the values at the subsample sampling locations through an interpolation filter. For example, in the case of HEVC, the luma component uses an 8-tap interpolation filter (while the chroma component uses a 4-tap filter).
[0077] Because many video coding standards employ a sub-pixel interpolation process, when using a motion vector MV2 pointing to non-integer sample locations, it may occur that its length (horizontal or vertical component) is smaller than that of a motion vector MV1 pointing to integer sample locations in the same direction. Due to the sub-pixel interpolation process, samples farther from the origin (e.g., from adjacent slices) are required, while using the MV1 motion vector will not require samples from adjacent slices.
[0078] Traditionally, video coding employs a translation-only motion model, where rectangular blocks are displaced according to a two-dimensional motion vector to form the motion compensation predictor for the current block. Such a model cannot represent rotational or scaling motions common in many video sequences. Therefore, efforts have been made to slightly extend the traditional translational motion model, for example, to what is called affine motion in JEM. Figure 5 As shown, in this affine motion model, two (or three) so-called control point motion vectors (v0, v1) for each block 104 are used to simulate rotational and scaling motions. Control point motion vectors v0, v1 are selected from MV candidates of neighboring blocks. The resulting motion compensation predictor is obtained by computing the sub-block-based motion vector field and relying on conventional translational motion compensation of the rectangular sub-block 112. Figure 1 An exemplary example is shown of the resulting predictor or block 114 (dashed line), where the upper right sub-block (solid line) makes predictions based on the motion vector of the corresponding sub-block.
[0079] Motion vector prediction is a technique for deriving motion information from temporal or space-time correlated blocks. In one such process, known as Alternative TMVP (ATMVP), such as... Figure 6 As shown, the default MV candidate (e.g., the first merge candidate in the merge candidate list) from the list of blocks 104 in the current image 12a is called the time vector 116 and is used to determine the relevant block 118 in the reference frame 12b.
[0080] In other words, for the current predicted block 104 within the CTU, a co-predicted block 118 is determined in the reference image 12b. To locate the co-predicted block 118, a motion vector MVTemp 116 is selected from the spatial candidates within the current image 12a. The co-predicted block is determined by adding the motion vector MVTemp 116 to the current predicted block position 12a.
[0081] Then, for each sub-block 122 of the current block 104, the motion information of the corresponding sub-block 124 of the related block 118 is used for inter-prediction of the current sub-block 122, which is also combined with additional procedures, such as MV scaling for the time difference between related frames.
[0082] Bidirectional optical flow (BIO) uses per-sample motion refinement performed on top of block-by-block motion compensation for bidirectional prediction. Figure 7 The process is explained. Refinement is based on minimizing the difference between image blocks (A, B) in two reference frames 12b and 12c (Ref0, Ref1) using the optical flow theorem. This theorem is based on the horizontal and vertical gradients of the corresponding image blocks. The gradient image, i.e., a measurement of the direction of brightness intensity change in the image, is used to calculate the enhancement offset of each sample in the prediction block in the horizontal and vertical directions of each image block in A and B.
[0083] A new concept has been introduced in state-of-the-art video codecs, such as the test model for VVC, called History-Based Motion Vector Prediction (HMVP). This involves maintaining a non-repeating FiFo buffer for the last used motion vector (MV) to populate the MV candidate list with more promising candidates than zero motion vectors. State-of-the-art codecs add HMVP candidates to the motion vector candidate list following (in a sub-block or block-based manner) co-occurring or spatial candidates.
[0084] Based on the above... Figures 1 to 3 Having described the general functionality of block-based video encoders and decoders through examples, and having generally explained the purpose of slice-independent coding, as well as some coding tools whose use, individually, may shorten slice-independent coding, as it may become very inefficient without special consideration, or if considered in an unfavorable manner. Below, embodiments are described that can maintain slice independence while still using coding tools such as those exemplified above very effectively.
[0085] Figure 8 A block-based video encoder 10 and a block-based video decoder 20 according to embodiments of this application are shown, wherein, as described above, the blocks are reused. Figure 1 and Figure 2 The reference symbol does not mean Figure 8 The encoder 10 and decoder 20 will be required as described above. Figure 1and Figure 2 The explanation is as outlined, although it may be interpreted as such, or perhaps only partially so. Encoder 10 encodes video 11 into data stream 14, while decoder 20 decodes the reconstructed video 11 from data stream 14. Images 12 of video 11 are divided into slices 100, as described above. Figure 4 Explained. Encoding independence in slice 100 is a task, according to Figure 8 Not only does encoder 10 take video 11 into account when encoding it into data stream 14, but decoder 20 also knows the slice boundaries 102 that pass through the interior of image 12 of video 11 and the adjacent slices 100 that are separated from each other. Specifically, as mentioned above regarding... Figure 4 As already outlined, slice 100 also extends in the temporal dimension because each image 12 is divided into slice 100 in the same way, such that each slice of an image has the same slices as any other image in video 11. Therefore, the motion information used to predict inter-prediction blocks within an image 12 should remain within the current slice, i.e., within the same slice as the corresponding inter-prediction block. In other words, such motion information for inter-prediction blocks in the current image, which determines the location of the corresponding inter-prediction block in the reference image, should not lead to inter-slice dependencies. Figure 8 The decoder at least partially alleviates the situation for the encoder 10 and derives motion information for a predetermined inter-prediction block 104 of the current image 12a of video 11, depending on the position of the boundary 102 within the image 12a of video 11, from which the block 130 is positioned in the reference image 12b of the video, from which the predetermined inter-prediction block 104 is predicted. Thus, for example, the motion information may define or include motion vector 132. The encoder 10 may again rely on this decoder behavior and may operate as follows: again, the encoder 10 determines the motion information 132 for the predetermined inter-prediction block of the current image 12a in a certain way, such that the block 130 remains within the boundary 102 of the slice in which the inter-prediction block 104 is located and does not cross the boundary 102. More precisely, the encoder 10 determines the motion information 132 such that the block is within the boundary of a slice in the same position as the slice in the reference image 12b, i.e., a slice in the same position as the slice in which block 104 is located. The encoder 10 determines this motion information 132 in this way and uses this motion information for inter-prediction of block 104. However, when encoding motion information into data stream 14, the encoder utilizes the functionality of decoder 40, or more precisely, decoder 20's understanding of boundary 102. Therefore, encoder 10 encodes motion information 132 into data stream 14 such that its derivation from data stream 14 depends on, i.e., on the position of boundary 102 between slices 100.
[0086] According to the embodiment described below, decoder 20 uses the aforementioned slice boundary awareness, i.e., the dependence on the slice boundary position derived from motion information in the data stream, to perform slice dependency analysis relative to the signal state of motion information of pre-determined inter-prediction blocks 104. See also Figure 9 Data stream 14 includes a portion or syntax portion 140 that signals motion information of block 104. Therefore, this portion 140 determines the state of the motion information of block 104 signaled. As will be described later, portion 140 may include, for example, motion vector differences, i.e., motion vector prediction residuals, and optionally, motion predictor indices pointing to a list of motion predictors. Reference image indices identifying reference images 12b associated with motion vectors may also optionally be comprised of portion 140. Using any precise description of the motion information portion 140, as long as the signaling state of the motion information of block 140 results in the corresponding block 130 to be predicted for block 104 being within the same slice 100a to which block 104 belongs, this signaling state does not conflict with slice independence constraints. Therefore, decoder 20 can leave this state and use it as the final state of the motion information, i.e., as the state for inter-prediction from block 130 for block 104. Figure 9 The diagram indicates two exemplary signal notification states: the state indicated by surrounding 1 is associated with motion information 132a causing block 130a to move, which is clearly within slice 100a. The other state indicated by surrounding 2 is associated with motion information 132b and causes block 130b to cross boundary 102 from the current slice 100a to an adjacent slice 100b. The overlapping portion of block 130b that overlaps with the adjacent slice 100b is defined by... Figure 9 The shading in the image represents this. Predicting block 104 based on block 130b will introduce inter-slice dependencies because samples from adjacent slices 100b will be used to predict block 104. Therefore, based on... Figure 9 For example, decoder 20 will check whether the signal notification state of section 140 is similar to Figure 9 In state 2, where the block is at least partially outside slice 100a, if so, the motion information is redirected from the signaling state to the redirection state, which causes the block 130 to remain within the boundary 102 of the current slice 100a. Figure 9 The dotted line in the text indicates that such a block 130'b is located by the redirection state 2' of the redirection motion information 130b'.
[0087] There are different ways to perform this redirection. These are described in detail below. In particular, as the description below reveals, blocks 130a, 130b, and 130'b may be of different sizes. At least some of them may have different sizes compared to block 104. This could be due, for example, to the need to predict block 104 from the reference image using interpolation filters at the corresponding blocks, for example, in cases involving corresponding motion information of sub-pixel motion vectors. Figure 9 As can be clearly seen from the description, the redirection 142 just described results in the encoder having two options to signal the motion information 130b' within the data stream 14: the encoder can use signal state 2 or state 2'. This "freedom" may result in a lower bit rate. Imagine, for example, using predictive coding to signal the motion information. If the predictor of the motion information in block 104 indicates state 2, the encoder will be able to simply use this predictor and signal the zero motion information prediction residual in part 140 of the data stream 14, thus signaling state 2, which will be redirected to state 2' by the decoder.
[0088] Figure 10 It shows about Figure 9 As already mentioned, the decoder can use the motion information derived from the data stream to determine the dependence of the slice boundary positions on the motion information in order to comply with or enforce constraints on the slice independence of the motion information signal notification state, which in turn is the result of the prediction of the motion information and its correction using the motion information prediction residual signal in portion 140 of the data stream 14 of block 104. That is, for block 104, the encoder and decoder will provide motion information prediction or motion information predictor 150, such as a vector, by using previously encoded blocks, such as motion information based on spatially adjacent blocks used to encode / decode block 104 in the current image or blocks located near block 104 in a specific reference image. The latter could be a reference image 12b containing block 130 or another image used for motion information prediction. A motion information prediction residual 152, such as a difference in motion vectors, is signaled in portion 140 of block 104. Therefore, the decoder decodes the motion information prediction residual 152 from the data stream 14 and determines the motion information (such as motion vectors) for block 104 based on the motion information prediction 150 and the motion information prediction residual 152. For example, in the case of motion vectors, the sum of motion vectors 150 and 152 indicates a motion vector 154 signaled by a signal notification state, such as the motion information of block 104, pointing to a footprint or block 130, wherein block 104 will be predicted from the footprint or block 130 by placing the bottom of motion vector 154 to a predetermined position 156, such as, for example, the position in a reference image located at a predetermined alignment position. Figure 10The position of block 104, such as the top left corner 158 shown, is like a corner. The top of vector 154 will then point, for example, to the top left corner of the footprint of block 104. As previously mentioned, due to filtering or other techniques, such as interpolation filtering used to derive the prediction of block 104 from block 130, block 130 can be amplified or otherwise distorted compared to the pure footprint. According to the embodiment described below, it is precisely this predictively decoded motion information 154 that is constrained to comply with / enforce.
[0089] That is, the decoder will adhere to or enforce the constraints imposed on motion information 154, such that block 130 does not exceed the boundary 102 of the slice 108 to which block 104 belongs. (As per...) Figure 9 As explained, the decoder uses an irreversible mapping to achieve this compliance / enforcement. The mapping will signal the state, as shown here... Figure 10 In the process, possible combinations of motion information prediction 150 and motion information prediction residual 152 are mapped to non-conflicting states of motion information 154. Irreversible mapping maps each possible combination to the final motion information, such that the block 130 placed according to this final motion information remains within the boundary 102 of the current block 100a. Therefore, based on the position 104 relative to the slice boundary of the current slice and / or the orientation and magnitude of motion information prediction 150, there exist combinations of motion information prediction 150 and motion prediction residual 152, i.e., possible signal notification states, which are mapped to the same final motion information 154, i.e., the same final state, even though these different combinations share, for example, the same motion information prediction 150 and differ only in the residual 152. Therefore, the encoder has the freedom to choose the motion information prediction residual 152. For example, Figure 10 This explains that the motion information prediction residual 152' may result in effectively the same motion information 154 due to the constraint compliance / enforcement on the decoder side, i.e., due to the redirection 142 of the aforementioned irreversible mapping.
[0090] In this case, where the encoder is free to choose any of the motion information prediction residuals that effectively result in the same motion information being applied to block 104 to perform inter-prediction, the encoder 10 can use one that is zero, as this setting may be the setting that results in the lowest bit rate.
[0091] Figure 11As mentioned above, block 104 can be indirectly predicted from block 130, which may have a different shape than block 104, where block 130 is located based on specific motion information such as a specific motion vector 132. Typically, block 130 is larger than the simple footprint 160 of block 104 within the reference image, displaced relative to its position in the current image according to motion information 132. Specifically, when performing the actual prediction of block 104 from block 130, the decoder and encoder can predict each sample 162 of block 104 using a mathematical combination of samples 164 within the corresponding portion of block 160. Typically, portions of block 130 help the mathematical combination for a particular block sample 162 to be located at and around the position of the corresponding sample 162, displaced at or around the motion information 132, i.e., its displaced sample position 162' within the reference image 12b. For example, the mathematical combination could be a weighted sum of samples corresponding to the corresponding samples in block 104. This portion of each sample 162 can be determined, for example, by the kernel of the interpolation filter, if the motion vector 132 is a subpixel motion vector, and thus determine the subpixel displacement of block 104 to the subpixel sample position within the reference image 12b. Interestingly, for example, if the motion information 132 includes a full-pixel motion vector, an interpolation filter may not be needed, in which case block 130 will be as large as block 104, and each sample 162 will be predicted to be set to equal one of the corresponding samples 164 in block 130. Another type, origin, or contribution to expanding footprint 160 to block 130 may alternatively or additionally derive from other inter-prediction tools, such as the reference BIO tool as described in detail below. In particular, the predictor derived from block 130 according to BIO may only be one of two assumptions about block 104, in which case bidirectional prediction will be performed. However, to obtain this predictor, the initial version of the predictor derived by interpolation in the case of sub-pixel motion vectors and the initial version of the predictor without interpolation filters in the case of full-pixel motion vectors are determined as follows: the initial version is magnified in all directions by n samples (n>0, such as 1) relative to the actual block size of 104. This is done to determine the local brightness gradients at multiple locations on the block 104 region distributed in this initial predictor, and to combine the two hypotheses—the initial predictor obtained from reference image 12b and the same other initial predictor obtained from another reference image—in a linear combination within the block region according to how the local gradients vary across the block region, to obtain the final dual-prediction predictor. Regarding Figure 22This explains why the contribution of the latter BIO to block expansion can be ignored, and whether the contribution of the n-sample-width BIO expansion (n>0) to block expansion might cause block 130 to cross boundary 102 and thus disable the BIO tool. That is, the existing redirection process will initially execute without considering the BIO tool, and the BIO tool will be disabled and replaced by the encoder in a boundary-aware manner. Furthermore, another possible mode besides the BIO tool is an affine motion model tool, according to which the motion information contains two motion vectors of the current block 104, defining the motion vectors of the motion vector field within block 104 at two different corners of block 104, and calculating the sub-block motion vectors of the sub-blocks of block 104 using the affine motion model based on it. The resulting sub-blocks can be considered as forming a combined block, for example, in... Figure 20 In this context, the slices are composed of sub-blocks, and again, slice-independent execution can execute combined slices as a whole without crossing boundary 102, meaning there are no individual sub-blocks in the slices. For explanations regarding redirection and execution, please refer to [link to relevant documentation]. Figure 9 and Figure 10 The function explained above means that the change in block size of block 130 is taken into account in two aspects: the actual block size of the motion information transmitted by the signal is subject to conditional redirection, and the motion information settings regarding the redirection or remapping of any "problematic" motion information. That is, although there is no specific... Figure 9 It is explicitly mentioned, but the redirection of specific motion information states may lead to changes in the chunk size: states that have not yet been redirected, such as... Figure 9 State 2 in the code may have a different block size, i.e., 2', compared to the result of redirection 142.
[0092] There are different possibilities regarding how to perform the aforementioned irreversible mapping or redirection. This applies to both sides: on the one hand, redirection occurs when the input of the irreversible mapping is formed; on the other hand, redirection occurs when the output of the irreversible mapping is redirected. Figure 12a This demonstrates a possibility. Figure 12a Four different settings 1, 2, 3, and 4 are shown or illustrated by displaying the corresponding blocks relative to boundary 102. These settings have not yet been redirected or undergone irreversible mapping. Boundary 102 separates the current slice 100a on one hand and the adjacent slice 100b on the other. For the first three states 1, 2, and 3, slice 130 is smaller than in state 4. In particular, Figure 12a This shows that in states 1, 2, and 3, block 130 coincides with the footprint 160 of the inner predicted block. In state 4, block 130 is magnified relative to footprint 160, which... Figure 12aThis setting 4 is shown with a dashed line. For example, the magnification is due to the fact that the corresponding motion vector in state 4 is a sub-pixel motion vector, while those in states 1 to 3 are full-pixel motion vectors. Due to this magnification, although blocks 130 in states 1 to 3 are located within the current block 100a and do not cross the boundary 102, block 130 in state 4, despite its corresponding motion vector being located between the full-pixel motion vectors of states 2 and 3, extends beyond the boundary 102.
[0093] In other words, constraint compliance / enforcement must be redirected to state 4 under any circumstances. According to some embodiments described later, states 1, 2, and 3 remain unchanged because the corresponding chunks do not cross boundary 102. This effectively means that the common domain of the irreversible mapping, i.e., the library or domain of the redirected states, allows chunk 130 to be closer to the boundary 102 of the current slice 100a than the width 170 of the extended edge portion 172, and chunk 130 is widened by the edge portion 172 compared to the footprint 160 of the inter-predicted block 104. In other words, for example, in the redirection process 142, the full-pixel motion vector of chunk 130 that does not cross boundary 102 remains unchanged. In different embodiments, even if those states of motion information are redirected, they do not extend beyond boundary 102 into adjacent chunks, but do not extend relative to the footprint if they would cross boundary 102: Figure 12a In the example, these are states 1, 2, and 3, and even those will be redirected to maintain a distance from boundary 102 in a way that accommodates at least a width of 170. See also Figure 12b In state 3, block 130 is shown to be smaller than the extension reaches 170 at a distance from boundary 102. Therefore, it is redirected to full-pixel state 5, where the associated block 130 or footprint 160 has a distance 171 at least as large as the extension reaches 171. That is, state 3 is redirected to state 5 where the corresponding block 130 does not include the boundary region 172 and is at least 171 large as the width 170 from boundary 102.
[0094] In other words, as a result of the subpixel interpolation process used in many video coding standards, it may be necessary to sample multiple integer sampling locations that are farthest from the nearest full-pixel reference block location. Assume motion vector MV1 is the full-pixel motion vector, and MV2 is the subpixel motion vector. More specifically, assume a size of B... W xB H Block 104 at position Pos x Pos y Located in slice 100a. Focusing on the horizontal component, to avoid using samples from adjacent slice 100b or the slice to the right, MV1 can be at most (assuming MV1...). x >0)Pos x +B W-1+Mv1 x Equal to the rightmost sample of slice 100a to which block 104 belongs, or (assuming MV1) x <0, meaning the vector points to the left) Pos x +Mv1 x This equals the leftmost sample of slice 100a to which the block belongs. However, for MV2, we must merge the interpolation from the filter kernel size to avoid reaching adjacent slices. For example, if we use Mv2... x If we consider the integer part of MV2, then MV2 can be at most as large as (assuming MV2). x >0)Pos x +B W -1+Mv2 x It equals the sample size of the slice to which the rightmost block belongs – 4, and (assuming MV2) x <0)Pos x +Mv2 x It equals the leftmost sample of the slice to which the block belongs + 3.
[0095] The “3” and “-4” here are examples of a width of 170 caused by the interpolation filter. However, other kernel sizes may also apply.
[0096] However, there is a desire to develop a means to effectively constrain motion compensation prediction at slice boundaries. Several solutions to this problem are described below.
[0097] One possibility is to clip the motion vectors relative to the boundaries 102 of slice 100, which are independent spatial regions within the encoded image. Furthermore, they can be clipped in a manner adapted to the resolution of the motion vectors.
[0098] Given block 104:
[0099] – Size equals B W xB H
[0100] – Position equals Pos x Pos y
[0101] – Belongs to those with Tile Left Tile Top Tile Right and Tile Bottom Slice 100a of defined boundary 102
[0102] definition:
[0103] –MVX Int For MVX >> precision
[0104] –MVYInt For MVY >> precision
[0105] –MVX Frac For MVX&(2 precision -1)
[0106] –MVY Frac For MVY&(2 precision -1)
[0107] Where MVX is the horizontal component of the motion vector, MVY is the vertical component of the motion vector, and precision represents the accuracy of the motion vector. For example, in HEVC, the accuracy of MV is 1 / 4 of the sample, therefore precision = 2.
[0108] We clip the full-pixel portion of the motion vector MV so that when the sub-pixel portion is set to zero, the resulting clipped full-pixel vector produces a footprint or block that is strictly within slice 100a:
[0109] MVX Int =Clip3(Tile Left -Pos x Tile Right -Pos x -(B W -1),MVX Int )
[0110] MVY Int =Clip3(Tile Top -Pos y Tile Bottom -Pos y -(B H -1),MVY Int )
[0111] When MVX Int <=Tile Left -Pos x +3 or MVX Int >=Tile Right -Pos x -(B W When the value is -1) to -4 (assuming an 8-tap filter in HEVC), it means that in the horizontal direction, the full-pixel portion of the motion vector (after cropping) is closer to the boundary 102, and then expands to 170.
[0112] MVX Frac Set to 0
[0113] When MVY Int <=Tile Top -Pos y+3 or MVY Int >=Tile Bottom -Pos y -(B H When the threshold is -1)-4 (assuming an 8-tap filter in HEVC), this means that in the horizontal direction, the full-pixel portion of the motion vector (after cropping) is closer to the boundary 102, then expands to 170.
[0114] MVY Frac Set to 0
[0115] This corresponds to about Figure 12a In the described scenario, the full-pixel motion vectors of states 1, 2, and 3 are allowed to remain unchanged because the corresponding blocks do not extend into the adjacent slice 100b. Any full-pixel motion vector that causes a block 130 to extend beyond boundary 102 will be clipped to the nearest full-pixel motion vector, whose corresponding block remains within boundary 102 without crossing it. For sub-pixel motion vectors of the expanded block 130 that cross boundary 102, the following procedure is applied: the sub-pixel motion vector is mapped to the nearest full-pixel motion vector, whose block remains within slice 100a, but is smaller than the sub-pixel motion vector. That is, the full-pixel portion of the motion vector is clipped accordingly, and the sub-pixel portion is set to zero. This means that if the full-pixel portion of the sub-pixel motion vector, i.e., the rounded-down version of the sub-pixel motion vector, does not cause footprint 160 to leave slice 100a, then the sub-pixel portion of the sub-pixel motion vector is only set to zero. Therefore, state 4 is redirected to state 2.
[0116] As an alternative to setting the decimal component to zero, the integer pixel portion of the motion vector may not be rounded to the next smaller integer pixel position as specified above, but it can also be rounded to the spatially nearest neighboring integer pixel position, as shown below:
[0117] The full-pixel portion of the motion vector MV is clipped so that when the sub-pixel portion is set to zero, the resulting clipped full-pixel vector will produce a footprint or block that is strictly within slice 100a, just as we did before.
[0118] MVX Int =Clip3(Tile Left -Pos x Tile Right -Pos x -(B W -1),MVX Int )
[0119] MVY Int =Clip3(Tile Top -Posy Tile Bottom -Pos y -(B H -1),MVY Int )
[0120] When MVX Int <=Tile Left -Pos x +3 or Tile Right -Pos x -(B W -1)>MVX Int >=Tile Right -Pos x -(B W -1)-4 (Assuming an 8-tap filter in HEVC)
[0121] MVX Int =MVX Int +(MVX Frac +(1<<(precision-1))>>precision)
[0122] When MVX Int <=Tile Left -Pos x +3 or MVX Int >=Tile Right -Pos x -(B W -1)-4 (Assuming an 8-tap filter in HEVC)
[0123] MVX Frac Set to 0
[0124] When MVY Int <=Tile Top -Pos y +3 or Tile Bottom -Pos y -(B H -1)>MVY Int >=Tile Bottom -Pos y -(B H -1)-4 (Assuming an 8-tap filter in HEVC)
[0125] MVY Int =MVY Int +(MVY Frac +(1<<(precision-1))>>precision)
[0126] When MVYInt <=Tile Top -Pos y +3 or MVY Int >=Tile Bottom -Pos y -(B H -1)-4 (Assuming an 8-tap filter in HEVC)
[0127] MVY Frac Set to 0
[0128] That is, the previously discussed process will change as follows: the motion vector is rounded to the nearest full-pixel motion vector that does not deviate from slice 100a. That is, the full-pixel portion can be cropped if necessary. However, if the motion vector is a sub-pixel motion vector, the sub-pixel portion is not simply set to zero. Instead, rounding to the nearest full-pixel motion vector is performed, i.e., by conditionally cropping the full-pixel portion to the nearest full-pixel motion vector that differs from the initial sub-pixel motion vector. For example, this will result in mapping... Figure 12a State 4, or redirect it to state 3 instead of state 2.
[0129] In other words, according to the first alternative just discussed, the motion vector extending between boundaries 102 in block 130 is redirected to a full-pixel motion vector by cropping the full-pixel portion of the motion vector, so that footprint 160 remains within slice 100a, and then the subpixels are set to zero. As a second option, after cropping the full-pixel portion, rounding is performed to the nearest full-pixel motion vector.
[0130] Another option is to be more stringent and to prune both the integer and fractional parts of the motion vector together in such a way that subsample interpolation of samples from another slice is avoided. See [link to relevant documentation] for more details on this possibility. Figure 12b .
[0131] MVX = Clip3((Tile Left -Pos x +3)< <precision,(Tile Right -Pos x -B W -1-4)< <precision,MVX)
[0132] MVY = Clip3((Tile Top -Pos y +3)< <precision,(Tile Bottom -Pos y -B H-1-4)< <precision,MVY)
[0133] As the description above reveals, MV clipping depends on the block size to which the MV is applied.
[0134] This description ignores the juxtaposition of color components that may have different spatial resolutions. Therefore, the cropping process shown above does not consider chroma formats, or only a 4:4:4 mode where each luminance sample has one chroma sample. However, there are two additional chroma formats where the relationship between chroma and luminance samples differs:
[0135] –4:2:0, the chromaticity has half the luminance sample in both the horizontal and vertical directions (i.e., one chromaticity sample for every two luminance samples in each direction).
[0136] –4:2:2, the chromaticity samples are half in the horizontal direction and half in the vertical direction (i.e., there is one chromaticity sample for each luminance sample in the vertical direction, and one chromaticity sample for every two luminance samples in the horizontal direction).
[0137] The chroma subpixel interpolation process can use a 4-tap filter, while the luminance filter uses an 8-tap filter, as described above. As mentioned above, the exact numbers are not important. The chroma interpolation filter kernel size can be half that of the luminance filter, but in alternative embodiments, even this can be changed.
[0138] For the 4:2:0 case, the integer and fractional parts of the motion vector are obtained as follows:
[0139] –MVX Int As MVCX >> (precision +1)
[0140] –MVY Int As MVCY >> (precision + 1)
[0141] –MVX Frac As MVCX&(2) precision+1 -1)
[0142] –MVY Frac As MVCY&(2 precision+1 -1)
[0143] This means that the integer part is half of the corresponding luminance part, while the fractional part has finer-grained signaling. For example, in HEVC, with precision = 2, 4 subpixels can be inserted between samples in the luminance case, while 8 can be inserted for chrominance.
[0144] This leads to the following fact: when xPos+MVX Int =Tile Left +1 and MVX IntWhen defined for luminance (not the chrominance mentioned above) in a 4:2:0 format, it falls within integer luminance samples but within fractional chrominance samples. Such samples will require more than one tile for subpixel interpolation. Left The chromaticity samples will prevent slice independence. This problem occurs when:
[0145] –4:2:0 (Color type in HEVC = 1)
[0146] o xPos+MVX Int =Tile Left +1 and MVX Int Defined as brightness
[0147] o xPos+(B W -1)+MVX Int =Tile Right -2 and MVX Int Defined as brightness
[0148] o yPos+MVY Int =Tile Top +1 and MVY Int Defined as brightness
[0149] o yPos+(B H -1)+MVY Int =Tile Bottom -2 and MVY Int Defined as brightness
[0150] –4:2:2 (Color type in HEVC = 2)
[0151] o xPos+MVX Int =Tile Left +1 and MVX Int Defined as brightness
[0152] o xPos+(B W -1)+MVX Int =Tile Right -2 and MVX Int Defined as brightness
[0153] There are two possible solutions.
[0154] Alternatively, cropping can be done in a restrictive manner based on chroma type (Ctype):
[0155] ChromaOffsetHor=2*(Ctype==1||Ctype==2)
[0156] ChromaOffsetVer=2*(Ctype==1)
[0157] MVX Int =Clip3(Tile Left -Pos x +ChromaOffsetHor,Tile Right -Pos x -(B W -1)-ChromaOffsetHor,MVX Int )
[0158] MVY Int =Clip3(Tile Top -Pos y +ChromaOffsetVer,Tile Bottom -Pos y -(B H -1)-ChromaOffsetVer,MVY Int )
[0159] Alternatively, crop as previously outlined without making additional changes due to chroma, but check if:
[0160] –4:2:0 (Color type in HEVC = 1)
[0161] o xPos+MVX Int =Tile Left +1 and MVX Int Defined as brightness
[0162] o xPos+(B W -1)+MVX Int =Tile Right -2 and MVX Int Defined as brightness
[0163] o yPos+MVY Int =Tile Top +1 and MVY Int Defined as brightness
[0164] o yPos+(B H -1)+MVY Int =Tile Bottom -2 and MVY Int Defined as brightness
[0165] –4:2:2 (Color type in HEVC = 2)
[0166] o xPos+MVXInt =Tile Left +1 and MVX Int Defined as brightness
[0167] o xPos+(B W -1)+MVX Int =Tile Right -2 and MVX Int Defined as brightness
[0168] And change MVX Int or MVY Int (For example, use +1 or the nearest direction that rounds to +1 or -1 based on the fractional part) so that the prohibition condition does not occur.
[0169] In other words, while the above description focuses on only one color component and ignores the possibility that different color components of a video image may have different spatial resolutions, this statement also applies to the corrections described below, such as the implementation of the slice independence constraint that considers the juxtaposition of different color components with different spatial resolutions. For example, the implementation applies to the MI predictor rather than the final MI. Frankly, when any color component requires an interpolation filter and correspondingly enlarges the block, the motion vector can be considered as a sub-pixel motion vector. Similarly, the motion vector for retargeting is chosen in a way that avoids any interpolation filtering, which may become necessary for any color component crossing the boundary 102 of the current slice 100a.
[0170] According to an alternative embodiment, where constraint compliance / enforcement has been applied to the motion vector 154 ultimately notified by a signal (compare) Figure 10 The above embodiments are modified to apply the same concept to the decoder, namely constraint compliance / enforcement, i.e., cropping to predictor 150 rather than the final motion vector or motion information 154 in the case of motion vectors. Therefore, if blocks of block 104 are to be placed according to the corresponding motion information predictor 150, i.e., if the motion information prediction residual 152 is zero, both the encoder and decoder act identically to enforce that all motion information predictors 150 comply with constraints, then the blocks are within the constraints of the current slice 100a. All the details outlined above can be transferred to these alternative embodiments.
[0171] However, with respect to encoder 10, the alternative embodiment just outlined results in a different situation for the encoder: by signaling one of the different motion information prediction residuals, the encoder is no longer ambiguous or "free" to allow the decoder to use the same motion information for block 104. Instead, once the motion information to be used for block 104 is selected, the signaling within portion 140 is uniquely determined, at least relative to its motion information predictor 150. The improvements are as follows: because motion information predictor 150 is prevented from conflicting with slice-independent constraints, there is no need to "redirect such motion information predictor 150 via the corresponding non-zero motion information prediction residual 152," whose signaling is generally more expensive in terms of bit rate than the signaling cost of zero motion information prediction residuals. Furthermore, in the case of establishing a list of motion information predictors for block 104, the automatic and synchronous execution of all motion information predictors upon which such a list of motion information predictors for block 104 is based does not conflict with slice-independent constraints, thereby increasing the likelihood that any of these available motion information predictors will be very close to the optimal motion information in terms of rate distortion optimization, since the encoder must select motion information in a manner that conforms to slice-independent constraints anyway.
[0172] In other words, while current video coding standards perform any MV cropping on the final MV—that is, if any—after adding the motion vector difference to the predictor, this is done additionally relative to the prediction and correction done using the residuals. When the predictor is pointed out of the image (potentially with boundary extension), if it is pointed really far from the boundary, the motion vector difference may not contain any component in the direction in which it is cropped, at least not after the cropped block is located at the image boundary. The motion vector difference is only meaningful if the resulting motion vector points to a location in the reference image that is located within the image boundary. However, adding such a large motion vector difference to reference a block within the image may be too costly compared to having it cropped to the image boundary.
[0173] Therefore, the embodiments will include trimming the predictors according to the block position, such that all MV predictors used from adjacent blocks or time candidate blocks always point to the slice containing the block, thus reducing the remaining motion vector difference that signals a good predictor and enabling more efficient signaling.
[0174] exist Figure 13 An example of implementation utilizing the slice independence associated with the aforementioned motion vector predictor is described. Specifically, Figure 13The diagram illustrates one possible scenario frequently mentioned above: the creation of a list 190 of motion information predictor candidates 192 for block 104 according to specific rules and in a specific order. One of these motion information predictor candidates 192 is then selected by the encoder for block 104, and is signaled in the syntax section 140 by a corresponding indicator 193 indicating the position of the selected motion information predictor candidate 192 within the list 190. The creation of the list 190 is performed in the same manner on both the encoder and decoder sides, and, according to the current embodiment, involves slice independence with respect to each motion information predictor candidate 192, through which the list 190 is filled. Therefore, all motion information predictors 192 in the list 190 are not associated with any block non-proprietarily within the boundary of the current slice to which block 104 belongs. Then, in cases where the motion information predictor candidate 192 is associated with a motion vector, a further syntax element 194 in section 140 indicates the motion information prediction residual, i.e., the motion vector difference 152.
[0175] Figure 14 The effect of subjecting each motion information prediction candidate 192 in list 190 to independent slice execution is illustrated: For block 104, blocks 130a to 130c of the three motion information predictor candidates 192a to 192c within list 190 are shown. However, block 130c is associated with motion information predictor candidate 192c, which is actually the result of redirection 142. That is, the encoder and decoder have actually derived the motion information predictor candidate for block 104, indicated by the dashed line and 192c'. However, this predictor 192c' results in block 130c' not being within the boundary 102 of the current slice 100a. Therefore, this motion information predictor candidate 192c' is redirected to become motion information predictor candidate 192c. The effect is as follows: the motion information predictor candidate 192c' has become, in the sense of large distortion, the most likely most effective motion information predictor candidate, which is quite low, because the encoder needs to put it together with the non-zero motion information prediction residual. Therefore, including such a candidate 192c' in list 190 would likely only result in a waste of candidate positions within list 190. Instead, any of the predictors 192a-c can be selected via indicator 193, and it is likely that only a small residual 194 needs to be included in the residual vector 152 (compared to...). Figure 10 ) and the selected predictor, such as 192c, is combined and transmitted to the final motion vector 154 (compare) Figure 10 The result of ).
[0176] Figure 13 and 14 The concept can also be applied to a single prediction 150 in the current block 104 without building a list 190 at all.
[0177] The alternative concept regarding the construction / establishment of the candidate list for motion information predictors is now regarding... Figure 15 The subject of the described embodiments. As mentioned above regarding Figure 13 The possibilities already outlined can be filled by using motion information predictor candidates to populate list 190, which are derived sequentially from previously encoded / decoded blocks according to predetermined rules. For example, these rules could relate to the positioning of the corresponding previously encoded / decoded blocks. For instance, it could be the spatially adjacent block to the left or top of the current block 104. The derivation of the corresponding motion information predictor candidates can then simply employ the corresponding motion information used for each block in the encoding / decoding process; that is, the corresponding motion information predictor candidate for block 104 can be set equal to the corresponding motion information used for the corresponding block. Another motion information predictor candidate can originate from a block in a reference image. Again, this could be one of the candidates in 192 or another related reference image. Figure 15 This illustrates motion information predictor candidate primitive 200. Primitives 200 form a motion information predictor candidate pool, filling entries in list 190 according to a certain order 204 called the filling order. Based on this order 204, the encoder and decoder derive primitives 200 and check their availability at 202. The filling order is indicated at 204. For example, availability can be rejected because the corresponding block of primitive 200, i.e., the block that has already been used to predict encoded / decoded motion information, is not within the same slice 100a as the current block 104. However, according to… Figure 15 In some embodiments, availability check 202 is accompanied by or includes the following additional or alternative checks: checking whether the corresponding primitive 200 conflicts with slice-independent constraints. Specifically, the availability of the origin of the corresponding primitive 200—that is, the block to the left, the block at the top of the current block 104, or a co-located block in the reference image—is not checked, but whether the motion vector predictor represented by such primitive 200 is checked; that is, whether the motion information derived from the motion information of previously encoded / decoded blocks, associated with the primitive and to be used to form primitive 200, is checked, and whether the primitive is associated with a block 130 located within the current slice 100a without exceeding its boundary. If not, the corresponding primitive is marked as unavailable and the process moves to the next primitive. For example, Figure 15This shows that the second candidate primitive 200 will be unavailable, and the first two motion information predictor candidates 192 in list 190 will therefore be formed by the first and third candidate primitives 200 in sequence 204. In other words, another solution to the above problem is to change the derivation of the availability flag in the motion vector list construction process. If constraintFlag is set, this can be done by setting motion vectors of samples outside the target independent encoding region (e.g., the currently encoded motion constraint block) to be unavailable.
[0178] For example, the following is the process of constructing the candidate list of motion vector predictors, mvpListLX.
[0179]
[0180] It can be seen that when mvLXA and mvLXB are available but point outside a given MCTS, potentially more promising co-located MV candidates or zero motion vector candidates will not be added to the list because it is already full. Therefore, it is advantageous to have a constraintFlag signaling within the bitstream, which controls the derivation of MV candidate availability in conjunction with the availability of reference samples with respect to spatial segment boundaries, such as slices.
[0181] In another embodiment, availability in the context of bidirectional prediction can allow for a more granular description of availability status, and thus further allow for the population of the motion vector candidate list with a hybrid version of partially available candidates.
[0182] For example, if the current block is a bidirectionally predicted region similar to its bidirectionally predicted spatial neighborhood, the above concept (availability labeling depends on the location of the reference sample in the slice) will result in the fact that spatial candidates whose MV0 points outside the current slice are marked as unavailable. Therefore, the entire candidate with MV0 and MV1 will not enter the candidate list. However, MV1 is a valid reference within the slice. To make this MV1 accessible through the motion vector candidate list, MV1 is added to a temporary list of partially available candidates. Combinations of partial candidates in the temporary list can then be added to the final MV candidate list. For example, MV0 of spatial candidate A can be mixed with MV1 or zero motion vector or HMVP candidate components of spatial candidate B.
[0183] The latter hint clearly indicates that, in the case of block 104 being a bidirectional prediction type, Figure 15The process described can be performed by a slightly more complex encoder and decoder. In this case, each candidate primitive 200 passes two motion vectors for two different reference images. One possibility is to mark such candidate primitive 200 as unavailable if one of the motion vectors causes a slice independence conflict. Therefore, this pair of motion vectors would be skipped in group order and not added to list 190. However, according to an alternative, slice independence checks are performed on assumptions separately. That is, for a given candidate primitive 200, one motion vector might cause a slice independence conflict, while the other might not. In that case, the non-conflicting motion vector hypothesis can be added to a specific alternative list of alternative motion information predictors. Where needed, i.e., if no further candidate primitives 200 are available in group order 204, the encoder and decoder can then form one or more other motion information predictor candidates 192 for list 190 from a single hypothetical motion information predictor and a list of alternatives, such as combining their pairs or a single entry and the list of alternatives with a default motion information predictor, such as a zero motion vector.
[0184] In order to complete the Figure 15 Note that additional details may be added regarding the description in the previous figures. For example, such details may relate to the indicator 193 and residual 194, as well as details about the block size and the actual use of the motion information for predicting block 104, which is ultimately signaled.
[0185] The embodiments described below process motion information prediction candidates originating from a specific reference image. Again, this reference image need not be the one from which the actual inter-image prediction is performed. However, it is possible. To distinguish the reference image used for MI (motion information) prediction from the reference image containing block 130, the former will be referred to with an apostrophe in the following text.
[0186] Figure 16 The current inter-prediction block 104 in the current image 12a and its co-located portion, namely the undisplaced footprint 104' within the reference image 12b', are shown. If temporal motion information predictor candidates (primitives) are formed using motion information of blocks in the reference image 12b' determined in the following manner, the results demonstrate cost-effectiveness in terms of RD performance: preferably, the blocks in the reference image 12b' are blocks including predetermined positions 204' in the reference image 12b, i.e., sample positions co-located with a predetermined positional relationship to block 104 in the current image 12a at a certain aligned position 204. Specifically, it is a sample position diagonally adjacent to the lower right sample of block 104 but located outside of block 104. Therefore, in Figure 16 In the example, position 204 is offset relative to block 104 along both the horizontal and vertical directions. The block in reference image 12b', which includes position 204', is... Figure 16 The location is indicated by reference symbol 206. Its motion information is used to derive or immediately represent, where available, the temporal motion information predictor candidate (primitive). However, availability may not be applicable, for example, to locations 204' and 204' outside the image or block 206, which are not inter-prediction blocks but, for example, intra-prediction blocks. In this case, i.e., when unavailable, another alignment location 208 is used to identify the source block or source block within the reference image 12b. This time, alignment location 208 is located inside block 104, such as centered. The corresponding location in the reference image is indicated by reference numeral 210. The block including location 210 is indicated by reference numeral 212, and this block is instead used to form the temporal motion information predictor candidate (primitive), i.e., by using the motion information of this block 212 that has already been predicted.
[0187] To avoid problems that would arise when applying the aforementioned concepts to every block 104 within the current image 104, the following alternative concepts are applied. Specifically, the encoder and decoder check whether block 104 is adjacent to a predetermined edge of the current slice 100a. This predetermined edge is, for example, the right boundary of the current slice 100a and / or, as in this example, the bottom edge of this piece 100a, offset horizontally and vertically relative to block 104 by alignment position 204. Therefore, if block 104 is adjacent to one of these edges, the encoder and decoder skip using alignment position 204 to identify the temporal motion information predictor candidate (primitive) and instead use only alignment position 208. The latter alignment position only “hits” blocks located within the reference image 12b within the same slice 100a as the current block 104. For all other blocks not adjacent to the right or bottom edge of the current slice 100a, temporal motion information predictor candidates (primitives) can be derived, including using alignment position 204.
[0188] However, it should be noted that regarding Figure 16 Many variations are possible with respect to the embodiments described. Figure 16 This possible correction in the proposed description involves, for example, aligning the exact positions of positions 204 and 208.
[0189] In other words, typical state-of-the-art video coding specifications also heavily rely on the concept of collecting motion vector predictors from a so-called candidate list at different stages of the coding process. The following describes the concept within the context of constructing the MV candidate list, which allows for more efficient coding of inter-predictive-constrained coding.
[0190] The list of candidate mergers can be interpreted as follows:
[0191] I = 0
[0192] if(availableFlagA1)
[0193] mergeCandList[i++] = A1
[0194] if(availableFlagB1)
[0195] mergeCandList[i++] = B1
[0196] if(availableFlagB0)
[0197] mergeCandList[i++] = B0
[0198] if(availableFlagA0)
[0199] mergeCandList[i++] = A0
[0200] if(availableFlagB2)
[0201] mergeCandList[i++] = B2
[0202] if(availableFlagCol)
[0203] mergeCandList[i++] = Col
[0204] When slice_type equals B, if there are not enough candidates in many video codec specifications, the process of combining bidirectional predictions to merge MV candidates is performed to populate the candidate list.
[0205] – Candidates are combined by using different combinations of the L0 component of one MV candidate and the L1 component of another MV candidate.
[0206] - If there are not enough candidates in the list, add zero motion vectors to merge candidates.
[0207] When considering collinear MV candidates (Col), issues arise related to the merging list in the context of inter-prediction constraints. If collinear blocks (not the bottom right but the center collinear) are unavailable, it's impossible to know if a Col candidate exists without resolving adjacent slices in the reference image. Therefore, the merging list may differ when decoding all slices, or only one or more slices may be decoded in a different arrangement than during encoding. Consequently, the MV candidate lists at the encoder and decoder ends may mismatch, and candidates from a certain index (col index) cannot be safely used.
[0208] The above Figure 16The implementation avoids the problem of mismatch in the MV candidate list mentioned above. For the rightmost and bottommost blocks in a slice, the block with the same center position is used as a Col candidate belonging to the current slice, instead of the bottom right block that does not belong to the current slice.
[0209] An alternative solution would be to modify the list construction based on group order 204. Specifically, when using about Figure 16 When describing concepts, time-motion information predictor candidates (primitives) can be used at a relatively early time in the list construction, such as the first primitive 200 to check availability (comparison). Figure 15 Slice boundary awareness at the encoder and decoder ends will ensure that no list mismatch occurs. However, another possibility is to move such temporal motion information predictor candidates (primitives) to the end of list 190 and combine spatial motion information candidate displacements, such as combining spatial motion vectors, before candidates that are temporally co-located, preferably based on block 206 and only determined in an auxiliary manner based on block 212. Relating to the pseudocode above, this would mean that candidate Col, determined primarily based on block 206 MI and only based on block 212 MI, would be moved to the end of the completed list 190 after the current state-of-the-art list construction. This change in filling order can be accomplished through slice boundary awareness: that is, the encoder and decoder will change the filling order by moving the Col candidate to the end of list 190, only for blocks 104 adjacent to the right or bottom of the current slice, and in another way for other slices, i.e., with Col preceding the combined spatial motion information candidates.
[0210] Naturally, for a more uniform creation of the motion vector candidate list, it would be feasible to perform the filling order variation just described on all blocks 104 within the slice, rather than only on blocks adjacent to a specific edge.
[0211] Figure 17 This illustrates the change in slice boundary-aware filling order. The case where the predicted block 104 is adjacent to one side of the current slice in question, i.e., the right or bottom edge, is shown. Figure 17 It is depicted in the left half, and Figure 17 The right-hand side depicts the situation of block 104, its distance from the current slice 100a at some point, or more precisely, its distance from a specific edge of the current slice 100a.
[0212] Using 200a, a motion information predictor candidate primitive is indicated, which is derived in the manner described above. According to this manner, motion information from block 206 is preferably used, with motion information from block 212 used as an alternative only in the case of, for example, internally encoded type 206. The block of the underlying primitive 200a is also called the alignment block, which may be 206 or 212. Another motion information predictor candidate primitive 200b is shown. This motion information predictor candidate primitive 200b is derived from the motion information of spatially adjacent blocks 220, which are spatially adjacent to the current block 104 on the side facing away from a specific edge of slice 100a, i.e., the top and left-hand sides of the current block 104. For example, the average or median of the motion information of adjacent blocks 220 is used to form the motion information predictor candidate primitive 200b. However, in the filling direction 204, the two cases are different: when block 104 is adjacent to one of the specific edges of the current slice 100a, the combined spatial candidate primitive 200b precedes the temporally co-located candidate primitive 200a; in the other case, for block 104 that is not adjacent to any specific edge of the current slice 100a, the order is changed so that the temporally co-located candidate primitive 200a is used to fill list 190 earlier in the hierarchical order 195, where the indicator 193 points to list 190 along the hierarchical order instead of the combined spatial candidate primitive 200b.
[0213] It should be noted that this was not targeted at Figure 17 Many details discussed in detail can be taken from any previous embodiment, such as regarding Figure 16 The embodiment concerning primitive 200a. Primitive 200b can also be obtained in different ways.
[0214] Figure 18 Further possibilities for obtaining slice-independent encoding in an efficient manner are shown. Here, a technique for deriving temporal motion information prediction candidates (primitives) is applied, according to which the encoder and decoder use additionally predicted motion vectors 240 to locate or identify predetermined blocks 242 in the reference image 12b, whose motion information, i.e., the motion information used to encode / decode block 242, is then used to derive motion information prediction candidates (primitives). For example, motion vector 240 can be predicted spatially. It may originate from one of the other candidates in list 190, rather than from... Figure 18 The target element. According to... Figure 18 In the described embodiment, the encoder and decoder check whether the position of motion vector 240 relative to block 104 points to a point outside the current slice 100a. Figure 18In this case, the predicted motion vector 240 is shown to point outside slice 100a. Therefore, this motion vector 240 is clipped by clipping operation 244 so that it remains within the boundary 102 of the current slice 100a, starting from block 104, or more precisely, from its predetermined alignment position 246, thus preventing it from pointing outside slice 100a. The thus clipped motion vector 250 relative to block 248, which block 104 points to, is then used as the basis for deriving temporal motion information prediction candidates (primitives), such as by simply using the motion information already predicted for block 248.
[0215] Figure 19 Another possibility is described. Here, the encoder and decoder first test the first predicted motion vector 240. By doing so, they check whether this first predicted motion vector points outside slice 100a. If so, they use another predicted motion vector 260 to derive temporal motion information prediction candidates (primitives), i.e., by taking the motion information used by the block 262 pointed to by the already predicted second predicted motion vector 260. For example, in Figure 19 As already described, the first predicted motion vector 240 points outside the current slice 100a, such that the block 242 to which this motion 240 points is located outside the current slice 100a in the reference image 12b. However, motion vector 260 does not point outside the current slice 100a. If the latter is true, then the corresponding temporal motion information prediction candidate (primitive) can be marked as unavailable.
[0216] In other words, in use cases involving independent coding space segments such as MCTS, the above ATMVP process needs to be restricted to avoid dependencies between MCTS.
[0217] In the construction of the subblock merge candidate list, subblockMergeCandList is constructed as follows, where the first candidate SbCol is the subblock temporal motion vector predictor:
[0218]
[0219] It is advantageous to ensure that the relevant block in the reference frame generated by the time vector does indeed belong to the spatial segment of the current block.
[0220] The location of the co-position prediction block is constrained to the following: the motion vector mvTemp, which is the motion information used to locate the co-position sub-block in the reference image, is cropped and contained within the co-position CTU boundary. This ensures that the co-position prediction block is located in the same MCTS region as the current prediction block.
[0221] Alternatively, for the above clipping, when the mvTemp of a spatial candidate does not point to a sample location within the same spatial segment (slice), the next available MV candidate is selected from the MV candidate list of the current block until a candidate is found that leads to a reference block located within the spatial segment.
[0222] The embodiments described below relate to another efficient encoding tool for reducing bitrate or efficiently encoding video. According to this concept, the motion information transmitted within the data stream of block 104 allows the decoder and encoder to define more complex motion fields within block 104, i.e., not only constant motion fields but also varying motion fields. For example, this motion information determines two motion vectors for the motion fields at two different angles of block 104, such as... Figure 5 As illustrated in the example. Therefore, the decoder and encoder are able to derive from the motion information each sub-block into which block 104 is divided. Figure 20 The example depicts a regular division into 4×4 sub-blocks 300, but the division into sub-blocks 300 does not need to be regular, nor is the number of sub-blocks limited to 16. The division into sub-blocks 300 can be signaled within the data stream or can be set by default. Therefore, each sub-block generates a motion vector indicating the translational displacement between this sub-block 300 and the corresponding block 302 in the reference image 12b, which will be predicted from the reference image 12b.
[0223] To avoid conflicts regarding slice independence, the encoder and decoder operate as follows according to this embodiment. Based on the... Figure 20 The two motion vectors 306a and 306b exemplarily represent the motion information transmitted in the data stream, resulting in a sub-block motion vector 304, which is executed by the encoder and decoder in the same manner. The resulting sub-block motion vector 304 presents an initial state. According to a first embodiment, for example, the decoder subjects each initial sub-block motion vector 304 to the above-mentioned... Figure 9 The slice independence described above is implemented. Therefore, for each sub-block, the decoder will test whether the corresponding sub-block motion vector 304 causes the corresponding block 302 of the corresponding sub-block to extend beyond the boundary of the current slice 100a. If so, the corresponding sub-block motion vector 304 is processed accordingly. Each sub-block 300 will then be predicted using its sub-block motion vector 304, which may have been redirected. The encoder does this in the same way, ensuring that all predictions for sub-block 300 are identical on both the encoder and decoder sides.
[0224] Alternatively, the two motion vectors 306a and 306b are redirected so that none of the sub-block motion vectors 304 results in slice dependency. Therefore, decoder-side pruning of vectors 306a and 306b can be used for this purpose. That is, as described above, the decoder treats motion vector pairs 306a and 306b as motion information, treats the combination of blocks 302 of sub-block 3000 as blocks of block 104, and performing combined pruning does not conflict with slice independence. Similarly, in the case where the combined pruning extends to another slice, the motion information corresponding to motion vector pairs 306a and 306b can be removed from the candidate list as taught above. Alternatively, a correction with residuals can be performed by the encoder, adding the corresponding MV difference 152 to vectors 306a and 306b.
[0225] According to alternative approaches, in different ways, such as, for example, using intra-prediction, those sub-blocks 300 corresponding to block 302 that exceed the slice boundary of the current slice 100a are predicted. Of course, this intra-prediction can be performed after inter-prediction of other sub-blocks of block 104 for which the corresponding sub-block motion vector 304 does not cause any slice independence conflicts. Furthermore, the prediction residuals passed even in the data stream of these non-conflicting sub-blocks may already be used on the encoder and decoder sides to reconstruct the interior of these non-conflicting sub-blocks before performing intra-prediction of these conflicting sub-blocks.
[0226] In use cases involving independent coding space segments such as MCTS, it is necessary to restrict the aforementioned affine motion process to avoid dependencies between MCTS.
[0227] When the MV of the sub-block motion vector field causes the sample position of the predictor to be outside the spatial segment boundary, the sub-block MV is clipped to the sample position inside the spatial segment boundary.
[0228] Additionally, when the MV of the sub-block motion vector field causes the predictor's sample position to be outside the spatial segment boundary, the resulting predictor sample is discarded and a new predictor sample is obtained from an intra-prediction technique, such as using an angular prediction pattern derived from adjacent and already decoded sample regions. The samples used in this scenario may belong to neighboring blocks of the current block as well as surrounding sub-blocks.
[0229] Alternatively, the motion vector candidates are checked against the predicted location and whether the predicted sample belongs to a spatial segment. If not, the resulting sub-block motion vectors are more likely not pointing to sampling locations within the spatial segment. Therefore, it is advantageous to prune motion vector candidates pointing to sampling locations outside the spatial segment boundary to locations within the spatial segment boundary. Alternatively, the mode using affine motion models can be disabled in this scenario.
[0230] Another concept involves constructing motion information prediction candidates (primitives) based on historical information. These can also be placed at the end of the padding order 204 above to avoid mismatches. In other words, adding HMVP candidates to the motion vector candidate list introduces the same problem described earlier: when the availability of temporally co-located MVs after encoding changes due to slice layout changes during decoding, the decoder-side candidate list may mismatch the encoder-side and cannot use indices after Col if such a use case is envisioned. The concept here is similar in spirit to the one above, where Col candidates are moved to the end of the list after HMVP when necessary.
[0231] The above concept can be applied to all blocks within a slice to create the MV candidate list.
[0232] It should be noted that the placement of historical motion information predictor candidates in the filling order can be accomplished alternatively, depending on the examination of the historical list of motion information currently included in the most recently predicted block. For example, some central trends of the motion information included in the historical list, such as the median, average, etc., can be used by the encoder and decoder to detect how likely it is that historical motion information predictor candidates will actually cause slice independence conflicts for the current block. For example, if, according to a certain average, the motion information included in the historical list points to the center of the current slice, or more generally, to a position sufficiently far from the current position of block 104 to the boundary 102 of the current slice 100a, it can be assumed that the probability of historical motion information predictor candidates causing slice independence conflicts is very small. Therefore, in this case, historical motion information predictor candidates (primitives) can be placed in the filling order earlier than if the average historical list points to a position close to or even beyond the boundary 102 of the current slice 100a. In addition to central trend measurements, dispersion measurements of the motion information included in the historical list can also be used. The greater the dispersion, such as variants, the higher the likelihood that historical motion information predictor candidates will cause slice independence conflicts, and they should be placed further closer to the end of the candidate list 190.
[0233] Figure 21The concept of improving the filling of the motion information predictor candidate list 190 is illustrated, wherein 500 motion information predictor candidates are selected from the motion information history list 502, which buffers a collection of recently used motion information, i.e., motion information of most recently encoded / decoded inter-prediction blocks. Specifically, each entry 504 in list 502 stores motion information that has been used for one of the most recently encoded / decoded inter-prediction blocks. The buffering in the history list 502 can be performed by the encoder and decoder in a first-in, first-out manner. Maintaining updates to the contents or entries of the history list 502 presents different possibilities. Due to the relative... Figure 21 The outlined concepts are useful both with and without slice-based independent coding. The encoder and decoder can limit the clustered regions of inter-predictive blocks, and the motion information used to fill the history list 502 with entries 504 can be limited to a single slice 100, i.e., the current slice 100a when examining the current coded block 104, or, in the case of not using slice-based independent coding, the entire image. Further, regarding... Figure 21 The concepts described are not necessarily used in the context of the concepts mentioned above. According to this aspect, historical motion information predictor candidates move towards the end of the motion information predictor candidate list 190, at least after spatial candidates, such as combined spatial candidates. Conversely, the following regarding... Figure 21 The outlined concepts assume only the existence of motion information predictor candidate 192, which has been used to populate motion information predictor candidate list 190 when selecting 500 history-based motion information predictor candidates from history list 502. As a final note, even... Figure 21 In cases where slice-based independent encoding is used, then Figure 21 The concept described above can or cannot be combined with any concept, wherein the decoder performs slice execution independently of the motion information signaled to the final block 104 or the motion information predictor candidates input into list 190. Figure 21 In the following description, it is assumed that slice-based independent encoding applies, but the decoder does not perform slice independence with respect to motion information predictor candidate 192.
[0234] like Figure 21 As shown, during the prediction block 104 between encoding / decoding, the encoder and decoder establish a candidate list 190 for motion information predictors. Figure 21 For example, assuming that the motion information predictor candidate list 190 has already been filled with two motion information predictor candidates 192, in Figure 21The terms A and B are used hereafter. "Two" is naturally just an example. In particular, it should be remembered that when selecting 500 history-based candidates for filling list 190 using another history-based predictor, the number of candidates 192 that have already filled block 104 in list 190 depends on the availability of inter-predictive blocks with respect to some co-location in other reference images and their surrounding space. However, Figure 21 The fact that two candidates 192 already exist in list 190 when selecting a history-based candidate from the 500-history list 502 does not imply that... Figure 21 The concept is limited to the following cases: at least two candidate primitives precede the history-based candidate primitives in the filling order, and it will not occur when there are already fewer than two candidates 192 in list 190 when selecting 500 history-based candidates for list 190.
[0235] Figure 21 The text describes the process of deriving candidate primitives for a history-based motion information predictor, where the selection of the next free entry 506 from the candidate list 190 will populate the history-based motion information predictor candidates. Specifically, Figure 21 The diagram shows motion information A and B defined by candidates 192 that have filled list 190 so far, as well as motion information in entry 504 of historical list 502. Figure 21 The motion vectors are represented by circles 1 to 5. For ease of understanding... Figure 21 The concept is that all motion vectors A, B and 1-5 are exemplary indicated as referring to the same reference image 12b, but those skilled in the art should understand that this is not necessarily the case, and the motion information of each candidate 192 and entry 504 may actually include the reference image index of the reference image to which the motion information corresponding to the index is involved, and this index may vary between candidate 190 and entry 504 respectively.
[0236] Instead of simply selecting the most recently input motion information from the history list 502, the encoder and decoder use motion information predictor candidates 192 that already exist in list 190 when selection 500 is performed. Figure 21In this case, these are motion information predictor candidates A and B. Specifically, to perform selection 500, the encoder and decoder determine the dissimilarity of each entry 504 in the history list 502 relative to the motion information predictor candidates 192 that were already filled with list 190 before selection 500. This dissimilarity can, for example, depend on the difference in motion vectors between the corresponding motion information 504 in the history list 502 and the motion information predictor candidates 192 in list 190. In the case where there is more than one candidate 192 in list 190, the minimum distance between the corresponding motion information 504 in list 502 and the candidate 192 in list 190 can be used, for example, to determine the dissimilarity. By using a double-headed arrow, in Figure 21 The difference in motion information 1 in list 502 is illustrated exemplarily. However, this dissimilarity may also depend on the difference in the reference image index. There are different possibilities regarding how to use this dissimilarity in performing selection 500. Generally, the aim is to perform selection 500 in a way that makes the motion information 504 selected from list 502 completely different from the motion information predictor candidate 192, which has so far filled list 190. Therefore, the dependency is designed in such a way that the higher the probability that a particular motion information 504 in the historical list 502 is selected, the higher its dissimilarity to the motion information predictor candidate 192 that already exists in list 190. For example, selection 500 can be performed in such a way that the selection depends on the dissimilarity and the rank at which the corresponding motion information 504 has been entered into the historical list 502. For example, Figure 21Arrow 508 indicates the order in which motion information 504 is input into list 502, i.e., the entry at the top of list 502 was input more recently than the entry at the bottom. According to a particular embodiment, for example, the encoder and decoder exclude motion information entries 504 from the historical list 502 whose motion information is less than a predetermined threshold and dissimilar to the motion information predictor candidates 192 currently used in list 190. For this purpose, the encoder and decoder can use the above-described example for distance measurement based on, for example, motion vector differences and / or reference image index differences, and threshold the resulting dissimilarity for a specific predetermined threshold parameter. The result of this process is the exclusion of specific motion information entries 504 from list 502, thus retaining only a subset of motion information entries, which can then be used to complete selection 500 to form motion information predictor candidates to be input into list 190 at position 506. This process can be repeated if list 190 has another candidate position in hierarchical order that must be filled using the historical list 502. In this case, lowering the aforementioned threshold parameter might result in a less stringent exclusion of motion information entries from list 502. That is, allowing more similar motion information to lead to a subset of motion information entries outside list 502, where the next motion information predictor candidate in list 190 is selected as the one most recently entered into historical list 502. Naturally, other possibilities exist for performing selection 500 in a manner dependent on the dissimilarity of motion information entries 504 in list 502 on motion information predictor candidate 192, which has so far populated candidate list 190 with motion information predictor candidate 192.
[0237] In fact, Figure 21 The concept leads to the collection of motion information predictor candidates 192 into the motion information predictor candidate list 190 of prediction block 104. Due to their increased dissimilarity, the encoder is more likely to find the best candidate outside this list 190 in terms of large distortion optimization. The encoder can then signal the selected candidate from list 190 for block 104 in data stream 14 via a corresponding indicator 193 along with a specific prediction residual 194.
[0238] Naturally, the possibilities described above are not limited to historical motion information predictor candidates selected from the historical list 502. Instead, this process can also be used to select from another pool of possible motion information predictor candidates. For example, it can be used to select from... Figure 21 The described method selects a subset of motion information predictor candidates from a larger set of motion information predictor candidates.
[0239] In other words, the state-of-the-art technique involves adding an available HMVP candidate only if none of the existing candidate lists matches the HMVP candidate to be added in terms of both horizontal and vertical components and the reference index. To improve the quality of HMVP candidates added to a motion vector candidate list, such as a merged candidate list or a sub-block motion vector candidate list, the insertion of each HMVP candidate into the list is based on a self-adjusting threshold. For example, assuming the motion vector candidate list is already filled with spatial domain A and further candidates B and Col are unavailable, the difference between the qualified HMVP candidate and the existing list entry (A in the given example) is measured for the next unoccupied list entry before adding the first qualified HMVP candidate. Only when the threshold is met is the corresponding HMVP candidate added; otherwise, the next qualified HMVP is tested, and so on. The measurement of the aforementioned difference can combine the horizontal and vertical components of the MV with the reference index. In one embodiment, the threshold is applied to each HMVP candidate, for example, the threshold is lowered.
[0240] at last, Figure 22 This involves the use of bidirectional optical flow tools in codecs to improve motion-compensated bidirectional prediction. Figure 22 The bidirectional prediction coding block 104 is shown. In the data stream of block 104, two motion vectors 1541 and 1542 are signaled, one associated with reference image 12b1 and the other with reference image 12b2. The footprints of block 104 within reference images 12b1 and 12b2' are respectively positioned in... Figure 22 The BIO tool determines the brightness gradient direction across the local region of block 104 to combine two hypotheses derived from the corresponding blocks 1301 and 1302 in reference images 12b1 and 12b2. Using these brightness gradient directions, the BIO tool alters the way blocks 104 are predicted from the two reference images 12b1 and 12b2, respectively, in a manner that varies across block 104. The BIO tool causes blocks 1301 and 1302 to be widened relative to the size of block 104, not only when the corresponding motion vectors 1541 or 1542 are subpixel motion vectors and / or require interpolation filters, but also when they are full-pixel vectors. In both cases, the additional widening, performed by an n-sample-width extension, widens the blocks, where n is an integer greater than 0, such as 1. This is the result of the BIO tool's behavior: In order to determine the local brightness gradient, the BIO tool derives respective hypothesis blocks 4021 and 4022 from each of blocks 1301 and 1302, which, relative to the size of block 104, widen the aforementioned n-sample-width extended edge portion, by... Figure 22Reference symbol 432 is used in the diagram. To generate the widened hypothetical blocks 4021 and 4022, the extended edge portions 1721 and 1722 of blocks 1301 and 1302 at least accommodate a corresponding n-sample widening 434 compared to the width of block 104, and an additional widening 436 corresponding to the width of the kernel range of the interpolation filter if the corresponding motion vectors 1541 and 1542 are sub-pixel motion vectors. Therefore, n is the quasi-sample width or widening associated with the bidirectional optical flow tool. Figure 22 In this example, motion vector 1542 is assumed to be a sub-pixel vector, such that the BIO tool derives hypothesis block 4022 from block 1302 via interpolation 4352, whereas in the case where motion vector 1541 is a full-pixel motion vector, the BIO tool can derive block 4021 simply by sample copying 4351. The BIO tool then determines the local brightness gradient across hypothesis blocks 4021 and 4022, so that the gradients determined in blocks 4021 and 4022 are used respectively to combine the two hypotheses, namely hypothesis predictor 4021 obtained from reference image 12b1 and predictor 4022 obtained from reference image 12b2, to obtain the final bidirectional prediction predictor 438 of block 104 in regions 1601 and 1602 of block 104 by linear combination 436 in a manner that depends on the variation of local gradients in the regions of the block. Specifically, for each sample 442 in block 104, the corresponding samples 4401 and 4402 in blocks 4021 and 4022 are summed, weighted by their respective hypothetical weights, such as half of their respective weights, or by other weights. The sum of these weights can optionally be summed until one is produced, and then a constant is added. This constant depends on the brightness gradient determined for the corresponding samples 4401 and 4402 in blocks 4021 and 4022 to derive the corresponding sample 442. This is done by the encoder and decoder, both of which have BIO tools.
[0241] It is possible that, with regard to widening 436, i.e., widening due to interpolation, the encoder notices that neither of blocks 1301 nor 1302 crosses boundary 102, regardless of whether the decoder performs slice independence constraint enforcement on the motion vector predictor used to encode motion vectors 1541 and 1542 into the data stream as described above, or alternatively, the decoder performs slice independence constraint enforcement on the final motion vectors 1541 and 1542 themselves. However, it is still possible for the expansion of blocks 1301 and 1302 beyond boundary 102 to be less than or equal to the sample width n associated with the bidirectional optical flow tool. However, both the encoder and decoder check whether the additional n-sample-width expansion 434 results in either block 1031 or 1302 still crossing boundary 102. If this is the case, the video encoder and video decoder deactivate the BIO tool. Otherwise, the BIO tool is not deactivated. Therefore, signaling in the data stream is not required to control the BIO tool. Figure 22 The diagram shows that block 1302 crosses the boundary 102 of the current slice 100a towards the adjacent slice 100b. Therefore, the BIO tool will be disabled here because the decoder and encoder recognize that block 1302 crosses boundary 102. As on the other hand, recalling the above regarding... Figure 11 The decoder can also consider widening the n samples by 434, so that when clipping the motion vector by 154, it will not be necessary to disable the slice boundary-aware BIO tool.
[0242] For example, with BIO disabled, each sample 442Sample(x, y) is derived as a simple weighted sum of the corresponding samples 4401 and 4402 of the hypothetical blocks 4021 and 4022, predSamplesL0[x][y] and predSamplesL1[x][y]:
[0243] Sample(x,y)=round(0.5*predSamplesL0[x][y]+0.5*predSamplesL1[x][y])
[0244] For all (x, y) in the current prediction block.
[0245] With the activation of BIO tools, this will change.
[0246] Sample(x, y)=round(0.5*predSamplesL0[x][y]+0.5*predSamplesL1[x][y]+bioEnh(x, y))
[0247] For all (x, y) in the current prediction block 104
[0248] bioEnh(x, y) is the offset calculated using the gradients of each of the two references 4021 and 4022 for each of the corresponding reference samples 4401 and 4402.
[0249] Alternatively, the decoder uses boundary padding to fill the area around footprints 402 and 402' at 432, extending it to blocks 430 and 430', which in turn extend beyond boundary 102 into adjacent slices.
[0250] In other words, in use cases involving independently encoded spatial segments such as MCTS, the above BIO process needs to be restricted to avoid dependencies between MCTS.
[0251] As part of the embodiment just described, in the case where the initial unrefined reference block is located at the boundary of the corresponding spatial segment in the image, so that samples outside the spatial segment are involved in the required gradient calculation process, BIO has been deactivated.
[0252] Alternatively, in this case, the exterior of the spatial segment is extended by a boundary padding procedure, such as repeating or mirroring sample values at the segment boundaries or using more advanced padding patterns. This padding would allow gradient calculations to be performed without using samples from adjacent spatial segments.
[0253] The following final notes will be made regarding the above embodiments and concepts. As noted repeatedly throughout the description of the various embodiments and concepts, the same content can be used and implemented individually or simultaneously in a particular video codec. Furthermore, the fact that the motion information described in many figures has been illustrated as containing motion vectors should be interpreted only as a possibility and should not limit embodiments that do not specifically utilize the fact that motion information contains motion vectors. For example, regarding Figure 8 The slice boundaries interpreted on the decoder side depend on motion information to derive motion information that is not limited to motion vectors. Regarding... Figure 9 It should be noted that the slice-independent constraint execution of the decoder applies the motion information described therein not only to the type of motion information based on motion vectors, but also to the predictive coding of motion information. Figure 10 As shown in the diagram. Through the analysis of... Figure 11 Feedback has provided examples that could result in a footprint larger than the inter-predicted block. However, embodiments of this application described with respect to the various figures may also relate to codecs in which, for example, such amplification does not occur, except for those embodiments specifically relating to situations. For example, Figure 12a and 12b This situation is specifically mentioned in the embodiments described herein. Regarding Figure 13 The concept of applying slice-independent constraint enforcement to motion information prediction or motion information predictors has been proposed. However, it should be noted that, although... Figure 13 This explains the concept of using / building a candidate list of motion information predictors, but this concept can also be applied to codecs that do not build a candidate list of motion information predictors for inter-prediction blocks. That is, a motion information predictor / predictor can be simply derived for block 104 using slice-independent constraint execution. However, for now, Figure 13 and 14 This can be modified to refer to only one embodiment of the motion information predictor. Furthermore, as noted above, all the details regarding the execution described in the previous diagrams, particularly... Figure 11 , 12a In 12b, it can be used to specify and implement in more detail. Figure 13and 14 Implementation examples. Regarding... Figure 15 Implementations / concepts have been proposed, according to which the availability of motion information predictor candidates depends on whether they cause slice-independent constraint conflicts. It should be noted that this concept may be related to... Figure 13 and Figure 14 The obfuscation of the embodiments lies in, for example, the characteristic motion information predictor candidates (primitives) are determined using slice-independent constraints, and the availability of certain motion information predictor candidates (primitives) depends on whether the corresponding candidate causes a slice-independent constraint conflict. Further, for example, Figure 15 The embodiments can be compared with those previously mentioned. Figure 9 The described approach applies slice-independent constraints to perform phase combination on the decoder side. In all embodiments that reference or utilize a candidate list of motion information predictors, the following methods can be used: Figure 16 The concept is used to determine the temporal motion information predictor candidate based on whether the current block is adjacent to a specific edge of the current slice, and to change the filling order based on whether the current block is adjacent to a specific edge of the current slice relative to this temporal motion information predictor candidate. Figure 17 Detailed information was provided. Again, Figure 16 and 17 Decoder-side slice independent constraints can be applied by combining motion information of the final state and / or the predictor with availability control according to Figures 12 and 14. The latter statement is also correct, according to... Figure 18 A concept for a candidate time-motion information predictor is derived. This concept can be combined with all the mentioned concepts because it can be combined with... Figure 16 and 17 The concept combination, and even on the one hand with Figure 16 and 17 Concept combination and on the other hand Figure 18 It is feasible. Because Figure 19 express Figure 18 The alternative diagram, therefore regarding Figure 18 The statements made also involve Figure 19 In particular, if only one motion information predictor is determined for a block, i.e., no candidate list is built, it can also be used. Figure 18 and Figure 19 .about Figure 20 It should be noted that the pattern upon which this concept is based, namely the affine motion model pattern, can represent a pattern of the video codec in addition to the pattern described herein, for example, with respect to all other embodiments described herein, or can represent a unique prediction pattern of the video codec. In other words, the motion information represented by vectors 306a and 306b can be information input to the candidate list mentioned with respect to any other embodiment, or represent the unique predictor of block 104. This is used to add information about... Figure 21 The concept of increasing variability in the motion information contained in the candidate list, as described above, can be combined with any other embodiment, and is particularly not limited to historical motion information-based predictor candidates. Given the... Figure 21 The described biological tool embodiments, regarding Figure 20 Similar statements regarding the affine motion model embodiments are true. With regard to all embodiments described herein, note that, as with regard to… Figure 21 As indicated, if more than one motion information predictor candidate is derived for block 104, then "the same" does not necessarily refer to the same reference image 12b, not only for... Figure 22 The B-prediction blocks outlined above are true. Furthermore, the above description primarily focuses on the boundary between adjacent slices 100a and 100b, i.e., the boundary guiding through the image. However, it should be noted that the details outlined above can also be transferred to slice boundaries consistent with the image boundary, i.e., the boundary of slices adjacent to the image circumference. Even for these boundaries, for example, the aforementioned decoder-side slice-independent constraints can be established, or availability constraints can be applied, etc. And further still, in all the embodiments presented above, the video codec may include flags in the data stream indicating whether slice-independent constraints are used to encode the video. That is, the video codec may signal to the decoder whether a specific constraint should be applied and whether the encoder has applied encoder-side constraint monitoring.
[0254] Although some aspects have been described in the context of the apparatus, it is clear that these aspects also represent a description of the corresponding method, where a block or device corresponds to a method step or feature of a method step. Similarly, aspects described in the context of a method step also represent a description of a corresponding block or item or feature of the corresponding apparatus. Some or all of the method steps can be performed by (or using) hardware devices, such as microprocessors, programmable computers, or electronic circuits. In some embodiments, one or more of the most important method steps can be performed by such devices.
[0255] The encoded video signal or data stream of the present invention can be stored on a digital storage medium or transmitted on a transmission medium such as a wireless transmission medium or a wired transmission medium such as the Internet.
[0256] Depending on specific implementation requirements, embodiments of the present invention can be implemented in hardware or software. This implementation can be executed using digital storage media, such as floppy disks, DVDs, Blu-ray discs, CDs, ROMs, PROMs, EPROMs, EEPROMs, or flash memory, having electronically readable control signals stored thereon that cooperate (or are capable of cooperating with) a programmable computer system to execute the corresponding methods. Therefore, the digital storage media can be computer-readable.
[0257] Some embodiments of the invention include a data carrier having electronically readable control signals that are capable of cooperating with a programmable computer system to perform one of the methods described herein.
[0258] Typically, embodiments of the present invention can be implemented as a computer program product having program code that, when run on a computer, is operable to perform one of the methods. The program code may, for example, be stored on a machine-readable medium.
[0259] Other embodiments include a computer program stored on a machine-readable medium for performing one of the methods described herein.
[0260] In other words, embodiments of the method of the present invention are therefore computer programs having program code that, when run on a computer, performs one of the methods described herein.
[0261] Therefore, a further embodiment of the method of the present invention is a data carrier (or digital storage medium, or computer-readable medium) having a computer program recorded thereon for performing one of the methods described herein. The data carrier, digital storage medium, or recording medium is typically tangible and / or non-transitional.
[0262] Therefore, a further embodiment of the method of the present invention is a data stream or signal sequence, which represents a computer program for performing one of the methods described herein. The data stream or signal sequence may, for example, be configured to be transmitted via a data communication connection, such as via the Internet.
[0263] Further embodiments include processing means, such as a computer or programmable logic device, configured or adapted to perform one of the methods described herein.
[0264] Further embodiments include a computer on which a computer program for performing one of the methods described herein is installed.
[0265] Further embodiments of the invention include means or systems configured to transmit (e.g., electronically or optically) to a receiver a computer program for performing one of the methods described herein. For example, the receiver may be a computer, mobile device, storage device, etc. For example, the means or system may include a file server for transmitting the computer program to the receiver.
[0266] In some embodiments, a programmable logic device (e.g., a field-programmable gate array) may be used to perform some or all of the functions of the methods described herein. In some embodiments, the field-programmable gate array may cooperate with a microprocessor to perform one of the methods described herein. Generally, these methods are preferably performed by any hardware device.
[0267] The apparatus described herein can be implemented using hardware devices, computers, or a combination of hardware devices and computers.
[0268] The apparatus described herein, or any component thereof, may be implemented, at least in part, in hardware and / or software.
[0269] The methods described herein can be performed using hardware devices, computers, or a combination of hardware devices and computers.
[0270] The methods or any components of the apparatus described herein may be performed, at least in part, by hardware and / or software.
[0271] The above embodiments are merely illustrative of the principles of the present invention. It should be understood that modifications and variations to the arrangements and details described herein will be readily apparent to those skilled in the art. Therefore, it is intended to be limited only by the scope of the forthcoming patent claims, and not by the specific details presented through the description and interpretation of the embodiments herein.
Claims
1. A video decoder comprising at least one processor configured to: identify a position of a current block in a current picture, wherein the current picture is one of a sequence of pictures in a temporal presentation order; derive motion information for the current block by adding a temporal motion vector MVTemp from a decoded block of the current picture to the current block position, wherein the temporal motion vector MVTemp indicates a translational displacement due to a temporal disparity between the decoded block of the current picture and a corresponding block in a reference picture that precedes the current picture in the temporal presentation order; clip the motion information to identify a collocated block within an independently coded spatial region of the reference picture, wherein the motion information is clipped to be within collocated CTU boundaries of the current block, or wherein when the motion information does not point to a sample position within a same spatial segment, a next available motion vector candidate is selected from a motion vector candidate list for the current block until a candidate is found that results in a reference block being located within the spatial segment; identify a position of the collocated block within the reference picture based on the clipped motion information; determine a motion vector for the collocated block in the reference picture; and determine a predicted motion vector for the current block based on the motion vector from the collocated block.
2. The video decoder of claim 1, the at least one processor further configured to scale the motion vector from the collocated block according to a temporal difference of the associated pictures.
3. The video decoder of claim 1, the at least one processor further configured to decode a reference picture index for the reference picture from a data stream.
4. The video decoder of claim 1, the at least one processor further configured to decode a motion vector prediction residual from a data stream. the independently coded spatial region is a slice.
5. The video decoder of claim 1, wherein, 6. The video decoder of claim 1, the at least one processor further configured to select a predicted motion vector as a selected candidate from a motion vector candidate list according to an index from a data stream.
7. The video decoder of claim 6, the at least one processor further configured to add a motion vector prediction residual from a data stream to the selected candidate to generate a motion vector for the current block.
8. A method of video decoding, the method comprising: identifying a position of a current block (118) in a current picture (12b), wherein the current picture is one of a sequence of pictures in a temporal presentation order; deriving motion information for the current block by adding a temporal motion vector MVTemp (116) from a decoded block of the current picture to the current block position, wherein the temporal motion vector MVTemp indicates a translational displacement due to a temporal disparity between the decoded block of the current picture and a corresponding block in a reference picture that precedes the current picture in the temporal presentation order; clipping the motion information to identify a collocated block within an independently coded spatial region of the reference picture, wherein the motion information is clipped to be within the collocated CTU boundaries of the current block, or wherein when the motion information does not point to a sample position within the same spatial segment, selecting the next available MV candidate from the MV candidate list of the current block until a candidate is found that results in a reference block being located within the spatial segment; identifying a location of the collocated block within the reference picture based on the clipped motion information; determining a motion vector of the collocated block in the reference picture; and determining a predicted motion vector for the current block based on the motion vector from the collocated block.
9. The method of claim 8, further comprising scaling the motion vector from the collocated block according to a temporal difference of the related picture.
10. The method of claim 8, further comprising decoding a reference picture index of the reference picture from the data stream.
11. The method of claim 8, further comprising decoding a motion vector prediction residual from the data stream. the independently coded spatial region is a slice.
12. The method according to claim 8, wherein, 13. The method of claim 8, further comprising selecting a predicted motion vector from a list of motion vector candidates as a selected candidate according to an index from the data stream.
14. The method of claim 13, further comprising adding a motion vector prediction residual from the data stream to the selected candidate to generate the motion vector for the current block.
15. A non-transitory digital storage medium having stored thereon computer program instructions which, when executed by at least one processor, cause the at least one processor to perform: identifying a location of a current block (118) in a current picture (12b), wherein the current picture is one of a sequence of pictures in a temporal presentation order; deriving motion information for the current block by adding a temporal motion vector MVTemp (116) from a decoded block of the current picture to the current block location, wherein the temporal motion vector MVTemp indicates a translational displacement due to a temporal difference between the decoded block of the current picture and a corresponding block in a reference picture that precedes the current picture in the temporal presentation order; clipping the motion information to identify a collocated block within an independently coded spatial region of the reference picture, wherein the motion information is clipped to be within the collocated CTU boundaries of the current block, or wherein when the motion information does not point to a sample position within the same spatial segment, selecting the next available MV candidate from the MV candidate list of the current block until a candidate is found that results in a reference block being located within the spatial segment; identifying a location of the collocated block within the reference picture based on the clipped motion information; determining a motion vector of the collocated block in the reference picture; and determining a predicted motion vector for the current block based on the motion vector from the collocated block.
16. The non-transitory digital storage medium of claim 15, the computer program further comprising instructions to scale the motion vector from the collocated block according to a temporal difference of the related picture. 17. The non-transitory digital storage medium according to claim 15, the computer program further comprising instructions to decode a reference picture index of a reference picture from the data stream.
Citation Information
Patent Citations
Methods and apparatuses for encoding, extracting and decoding video using tiles coding scheme
US20140119671A1
Coding concept allowing efficient multi-view / layer coding
US20160057441A1