Motion vector processing for video encoding and decoding
By introducing flexible motion vector resolution control and affine motion model in video coding, the motion vector prediction and coding process is optimized, which solves the problems of low efficiency and high complexity in the existing technology and achieves more efficient video compression and decoding.
Patent Information
- Application Number
- CN202080054531.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-06-20
- Filing Date
- 2020-06-18
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2040-06-18
AI Technical Summary
Existing video coding technologies suffer from low efficiency and high complexity in the motion vector prediction and encoding processes. Especially in high-efficiency video coding standards such as VVC Draft 5, existing tools such as AMVR and IBC lack flexibility, resulting in poor compression and decoding efficiency.
By introducing flexible motion vector resolution control in the video encoding process, allowing the resolution differentiation of motion vector differences (MVd) and block vector components with different accuracy levels, combined with the affine motion model and adaptive motion vector resolution (AMVR) tool, the motion vector prediction and encoding process is optimized.
It improves the compression efficiency and decoding efficiency of video coding, reduces complexity, enhances the coding flexibility of motion information, and is suitable for more complex video content.
Smart Images

Figure CN114270844B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to video encoding and decoding. Background Art
[0002] To achieve high compression efficiency, image and video coding schemes typically employ prediction and transforms to exploit spatial and temporal redundancies in video content. Typically, intra-frame or inter-frame prediction is used to exploit intra-frame or inter-frame correlations. The difference between the original and predicted picture blocks, typically represented as prediction errors or prediction residuals, is then transformed, quantized, and entropy-coded. One aspect of the prediction process can involve motion-compensated temporal prediction based on motion vectors, as explained further below. To reconstruct the video, the compressed data is decoded using the inverse processes of prediction, transform, quantization, and entropy coding. Summary of the Invention
[0003] Generally speaking, an example of an embodiment may relate to a method comprising: determining a decoding mode associated with a current decoding unit including picture information; determining, based on the decoding mode, a first precision level associated with a first portion of motion vector information and a second precision level associated with a second portion of the motion vector information, wherein the motion vector information is associated with the current decoding unit; obtaining a motion vector associated with the current decoding unit based on the motion vector information and the first precision level and the second precision level; and decoding at least a portion of the picture information included in the current decoding unit based on the decoding mode and the motion vector.
[0004] In general, another example of an embodiment may relate to a method comprising: determining a decoding mode associated with a current decoding unit comprising picture information; determining, based on the decoding mode, a first precision level associated with a first portion of motion vector information and a second precision level associated with a second portion of the motion vector information, wherein the motion vector information is associated with the current decoding unit; obtaining a motion vector associated with the current decoding unit based on the motion vector information and the first precision level and the second precision level; and encoding at least a portion of the picture information based on the decoding mode and the motion vector.
[0005] In general, another example of an embodiment may involve a device comprising one or more processors configured to determine a decoding mode associated with a current decoding unit comprising picture information; based on the decoding mode, determine a first precision level associated with a first portion of motion vector information and a second precision level associated with a second portion of the motion vector information, wherein the motion vector information is associated with the current decoding unit; based on the motion vector information, and the first precision level and the second precision level, obtain a motion vector associated with the current decoding unit; and decode at least a portion of the picture information based on the decoding mode and the motion vector.
[0006] In general, another example of an embodiment may involve a device comprising one or more processors configured to determine a decoding mode associated with a current decoding unit comprising picture information; based on the decoding mode, determine a first precision level associated with a first portion of motion vector information and a second precision level associated with a second portion of the motion vector information, wherein the motion vector information is associated with the current decoding unit; based on the motion vector information, and the first precision level and the second precision level, obtain a motion vector associated with the decoding unit of picture information; and encode at least a portion of the picture information based on the decoding mode and the motion vector.
[0007] In general, another example of an embodiment may relate to a method comprising: determining a decoding mode associated with a current decoding unit including picture information; determining, based on the decoding mode, a first precision level associated with a first component of a motion vector residual and a second precision level associated with a second component of the motion vector residual, the motion vector residual being associated with the current decoding unit; obtaining a motion vector associated with the current decoding unit based on the motion vector residual and the first precision level and the second precision level; and decoding at least a portion of the picture information included in the current decoding unit based on the decoding mode and the motion vector.
[0008] In general, another example of an embodiment may relate to a method comprising: determining a decoding mode associated with a current decoding unit comprising picture information; determining, based on the decoding mode, a first precision level associated with a first component of a motion vector residual and a second precision level associated with a second component of the motion vector residual, the motion vector residual being associated with the current decoding unit; obtaining a motion vector associated with the current decoding unit based on the motion vector residual and the first precision level and the second precision level; and encoding at least a portion of the picture information based on the decoding mode and the motion vector.
[0009] In general, another example of an embodiment may involve a device comprising one or more processors configured to determine a decoding mode associated with a current decoding unit comprising picture information; determine, based on the decoding mode, a first precision level associated with a first component of a motion vector residual and a second precision level associated with a second component of the motion vector residual, the motion vector residual being associated with the current decoding unit; obtain a motion vector associated with the current decoding unit based on the motion vector residual and the first precision level and the second precision level; and decode at least a portion of the picture information based on the decoding mode and the motion vector.
[0010] In general, another example of an embodiment may involve a device comprising one or more processors configured to determine a decoding mode associated with a current decoding unit comprising picture information; determine, based on the decoding mode, a first precision level associated with a first component of a motion vector residual and a second precision level associated with a second component of the motion vector residual, the motion vector residual being associated with the current decoding unit; obtain a motion vector associated with the decoding unit of picture information based on the motion vector residual and the first precision level and the second precision level; and encode at least a portion of the picture information based on the decoding mode and the motion vector.
[0011] In general, another example of an embodiment may relate to a method comprising: determining a decoding mode associated with a current decoding unit comprising picture information; determining, based on the decoding mode, a first precision level associated with a first motion vector residual corresponding to a first reference picture list, and a second precision level associated with a second motion vector residual corresponding to a second reference picture list; obtaining a first motion vector based on the first motion vector residual and the first precision level, and obtaining a second motion vector based on the second motion vector residual and the second precision level; and decoding at least a portion of the picture information based on the decoding mode, and the first motion vector and the second motion vector.
[0012] In general, another example of an embodiment may relate to a method comprising: determining a decoding mode associated with a current decoding unit comprising picture information; determining a first precision level associated with a first motion vector residual corresponding to a first reference picture list and a second precision level associated with a second motion vector residual corresponding to a second reference picture list based on the decoding mode; obtaining a first motion vector based on the first motion vector residual and the first precision level, and obtaining a second motion vector based on the second motion vector residual and the second precision level; and encoding at least a portion of the picture information based on the decoding mode and the first and second motion vectors.
[0013] In general, another example of an embodiment may involve a device comprising one or more processors configured to determine a decoding mode associated with a current decoding unit comprising picture information; determine, based on the decoding mode, a first precision level associated with a first motion vector residual corresponding to a first reference picture list, and a second precision level associated with a second motion vector residual corresponding to a second reference picture list; obtain a first motion vector based on the first motion vector residual and the first precision level, and obtain a second motion vector based on the second motion vector residual and the second precision level; and decode at least a portion of the picture information based on the decoding mode, and the first motion vector and the second motion vector.
[0014] In general, another example of an embodiment may involve a device comprising one or more processors configured to determine a decoding mode associated with a current decoding unit comprising picture information; determine, based on the decoding mode, a first precision level associated with a first motion vector residual corresponding to a first reference picture list, and a second precision level associated with a second motion vector residual corresponding to a second reference picture list; obtain a first motion vector based on the first motion vector residual and the first precision level, and obtain a second motion vector based on the second motion vector residual and the second precision level; and encode at least a portion of the picture information based on the decoding mode, and the first motion vector and the second motion vector.
[0015] In general, another example of an embodiment may involve a bitstream formatted to include encoded picture information, wherein the encoded picture information is encoded by processing the picture information based on any one or more of the examples of embodiments of the method according to the present invention.
[0016] In general, one or more other examples of the embodiments may further provide a computer-readable storage medium, such as a non-volatile computer-readable storage medium, on which instructions for encoding or decoding picture information such as video data according to the method or apparatus described herein are stored. One or more embodiments may further provide a computer-readable storage medium on which a bitstream generated according to the method or apparatus described herein is stored. One or more embodiments may further provide a method and apparatus for transmitting or receiving a bitstream generated according to the method or apparatus described herein.
[0017] As explained below, various modifications and embodiments are contemplated that may provide improvements to video encoding and / or decoding systems, including but not limited to one or more of increased compression efficiency and / or coding efficiency and / or processing efficiency and / or reduced complexity.
[0018] The above provides a simplified overview of the present subject matter in order to provide a basic understanding of some aspects of the present disclosure. This overview is not an exhaustive review of the present subject matter. It is not intended to identify key / important elements of various embodiments or to delineate the scope of the present subject matter. Its sole purpose is to present some concepts of the present subject matter in a simplified form as a prelude to the more detailed description provided below. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The present disclosure may be better understood by considering the following detailed description in conjunction with the accompanying drawings, in which:
[0020] Figure 1 A block diagram is provided depicting an example of an embodiment of a video encoder;
[0021] Figure 2 A block diagram depicting an example of an embodiment of a video decoder is provided;
[0022] Figure 3 shows the coding tree unit (CTU) and coding tree concepts that can be used to represent compressed pictures;
[0023] Figure 4 shows the coding tree unit (CTU) and the partitioning of the CTU into coding units (CU), prediction units (PU), and transform units (TU);
[0024] Figure 5 Examples of spatial and temporal motion vector prediction candidates in AMVP mode such as VVC draft 5 are shown;
[0025] Figure 6 shows the calculation of motion vector prediction candidates in AMVP mode such as VVC draft 5;
[0026] Figure 7 An example of an affine motion model is shown;
[0027] Figure 8 An example of an affine motion vector field based on a 4×4 sub-CU is shown;
[0028] Figure 9 shows examples of spatial positions (A0, A1, B0, B1, and B2) from which CPVM prediction candidates are retrieved to predict the affine model of the current CU in an affine AMVP mode (such as that of VTM draft 3);
[0029] Figure 10 shows the motion vector prediction process for an affine AMVP CU such as a VTM;
[0030] Figure 11 An example of constructing a CPVMP candidate list in affine AMVP mode is illustrated;
[0031] Figures 12 to 27 Examples of various aspects, embodiments, and features according to the present disclosure are shown;
[0032] Figure 28 A block diagram illustrating an example of an embodiment of an apparatus according to various aspects and embodiments described herein is provided;
[0033] Figure 29 An example of an embodiment according to the present disclosure is shown; and
[0034] Figure 30 Another example according to an embodiment of the present disclosure is shown.
[0035] It should be understood that the drawings are for purposes of illustrating examples of various aspects and embodiments and are not necessarily the only possible configurations. Throughout the drawings, like reference numerals represent like or similar features. DETAILED DESCRIPTION
[0036] In the accompanying drawings, Figure 1 An example of a video encoder 100 is shown, such as a HEVC encoder. HEVC is a Video Encoding Joint Collaboration Team (JCT-VC) A compression standard being developed (see, for example, "ITU-T H.265 ITU Telecommunication Standards Section (10 / 2014), Series H: Audiovisual and Multimedia Systems, Infrastructure of Audiovisual Services - Coding of Motion Video, High Efficiency Video Coding, Recommendation ITU-T H.265"). Figure 1 Encoders in which improvements have been made to the HEVC standard or encoders employing techniques similar to HEVC may also be shown, such as encoders based on or improved upon the JEM (Joint Exploration Model) developed by the Joint Video Experts Group (JVET), e.g., associated with Versatile Video Coding (VVC) specified by the development work.
[0037] In this application, the terms "reconstruction" and "decoding" may be used interchangeably, the terms "pixel" and "sample" may be used interchangeably, and the terms "picture" and "frame" may be used interchangeably.
[0038] The HEVC specification distinguishes between "blocks," which address a specific area in the sample array (e.g., luma, Y), and "units," which include a collocated block of all coded color components (Y, Cb, Cr, or monochrome), syntax elements, and prediction data associated with the block (e.g., motion vectors).
[0039] For decoding, the picture is divided into coding tree blocks (CTBs) of a configurable square size, and a continuous set of coding tree blocks is grouped into slices. A coding tree unit (CTU) includes the CTB of the coded color component. The CTB is the root of the quadtree partition into coding blocks (CBs), and the coding block can be divided into one or more prediction blocks (PBs) and forms the root of the quadtree partition into transform blocks (TBs). Corresponding to the coding blocks, prediction blocks, and transform blocks, the coding unit (CU) includes a tree structure set of prediction units (PUs) and transform units (TUs), the PU includes prediction information for all color components, and the TU includes a residual coding syntax structure for each color component. The size of the CB, PB, and TB of the luma component applies to the corresponding CU, PU, and TU. In this application, the term "block" may be used to represent any one of CTU, CU, PU, TU, CB, PB, and TB. Additionally, "block" may also be used to refer to macroblocks and partitions as specified in H.264 / AVC or other video coding standards, and more generally to data arrays of various sizes.
[0040] exist Figure 1 In the encoder 100, as described below, the picture is encoded by the encoder elements. The picture information to be encoded is provided at the input and is mapped (101) and image partitioned (102). The mapping (101) is typically applied to each sample and a 1D function is applied to the input sample value to convert it to other sample values. For example, the 1D function can extend the sample value range and provide a better distribution of codewords over the codeword range. The image partitioning (102) divides the image into blocks of different sizes and shapes in order to optimize the rate-distortion tradeoff. The mapping and partitioning enable the picture information to be processed in units of CUs as described above. Each CU is encoded using intra or inter mode. When the CU is encoded in intra mode, it performs intra prediction (160). In inter mode, motion estimation (175) and compensation (170) are performed. The encoder decides (105) which of intra mode or inter mode to use for encoding the CU and indicates the intra / inter decision by a prediction mode flag. The intra or inter prediction decision (105) is followed by forward mapping (191) to produce a prediction block. In general, the forward mapping process (191) is similar to and can be complementary to mapping (101). The prediction residual is calculated by subtracting (110) the prediction block from the original image block.
[0041] The prediction residual is then transformed (125) and quantized (130). The quantized transform coefficients, along with motion vectors and other syntax elements, are entropy decoded (145) to output a bitstream. The encoder can also skip the transform and apply quantization directly to the untransformed residual signal on a 4×4 TU basis. The encoder can also bypass the transform and quantization, i.e., directly decode the residual without applying the transform or quantization process. In direct PCM decoding, no prediction is applied, and the coding unit samples are decoded directly into the bitstream.
[0042] The encoder decodes the coded block to provide a reference for further prediction. The quantized transform coefficients are dequantized (140) and inverse transformed (150) to decode the prediction residual. The decoded prediction residual and the prediction block are combined (155) to reconstruct the image block. An in-loop filter (165) is applied to the reconstructed image, for example, to perform deblocking / SAO (sample adaptive offset) filtering to reduce coding artifacts. The inverse mapping (190) is the inverse of the forward mapping (191). The filtered image is stored in a reference picture buffer (180).
[0043] Figure 2 A block diagram of an example of a video decoder 200, such as an HEVC decoder, is shown. In the example decoder 200, a signal or bitstream is decoded by decoder elements as described below. The video decoder 200 generally performs the same operations as described above. Figure 1 The encoding pass described in
[0065] is reciprocally compared to the decoding pass, which performs video decoding as part of encoding the video data. Figure 2 Decoders in which improvements are made to the HEVC standard or decoders employing techniques similar to HEVC may also be shown, such as decoders based on or improved upon JEM.
[0044] Specifically, the input to the decoder includes the Figure 1 A video signal or bitstream is generated by the video encoder 100 of the video encoder. The signal or bitstream is first entropy decoded (230) and then partitioned (210) to obtain transform coefficients, motion vectors and other decoding information. Partitioning (210) divides the image into blocks of different sizes and shapes based on the decoded data. The transform coefficients are dequantized (240) and inverse transformed (250) to decode the prediction residual. The decoded prediction residual is combined with the prediction block (255) and the image block is reconstructed. The prediction block can be obtained (270) from intra-frame prediction (260) or motion compensated prediction (i.e., inter-frame prediction) (275). Advanced motion vector prediction (AMVP) and merge mode techniques can be used to derive motion vectors for motion compensation, which can use interpolation filters to calculate interpolated values of sub-integer samples of a reference block. An in-loop filter (265) is applied to the reconstructed image. The filtered image is stored in a reference picture buffer (280).
[0045] In the HEVC video compression standard, motion compensated temporal prediction is used to exploit the redundancy that exists between consecutive pictures in a video. To do this, a motion vector is associated with each prediction unit (PU). Each coding tree unit (CTU) is represented by a coding tree (CT) in the compressed domain. This is a quadtree partitioning of the CTU, where each leaf is called a coding unit (CU), as shown in the following example. Figure 3 As shown in .
[0046] Each CU is then given some intra-frame or inter-frame prediction parameters or prediction information (prediction information). To this end, the CU is spatially divided into one or more prediction units (PUs), and each PU is assigned some prediction information. Figure 4 As shown in Figure 1, intra or inter coding modes are assigned at the CU level, which shows an example of dividing the coding tree unit into coding units, prediction units, and transform units. To decode a CU, a prediction block or prediction unit (PU) is constructed from adjacent reconstructed samples (intra prediction) or from a previously reconstructed picture stored in the decoded picture buffer (DPB) (inter prediction). Next, the residual samples calculated as the difference between the original samples and the PU samples are transformed and quantized.
[0047] Specifically, a motion vector (MV) is assigned to each PU in HEVC. This motion vector is used for motion compensated temporal prediction of the considered PU. Therefore, in HEVC, the motion model of the linked prediction block and its reference blocks simply consists of translation.
[0048] There are two modes used in HEVC to encode motion data. They are called AMVP (Adaptive Motion Vector Prediction) and Merge. AMVP basically involves signaling the reference picture(s) used to predict the current PU, the motion vector predictor index (taken in a list of two predictors), and the motion vector difference. Merge mode involves signaling and decoding the index of some motion data collected in a list of motion data predictors. The list consists of six candidates and is constructed in the same way on the decoder and encoder side. Therefore, merge mode aims to derive some motion information taken from the merge list. The merge list usually contains motion information associated with some spatial and temporal surrounding blocks that are available in their decoding state when the current PU is being processed.
[0049] Codecs and video compression tools other than HEVC, such as the Joint Exploration Model (JEM) and the codecs and video compression tools developed by the JVET (Joint Video Exploration Team) group in the generic video decoding (VVC) reference software known as the VVC Test Model (VTM), may provide various modes and features that are different from or in addition to the modes and features of HEVC. For example, one or more aspects of the present disclosure will be explained with respect to VVC Draft 5 in a non-limiting manner. That is, the explanation in the context of VVC Draft 5 is for ease of explanation only. The aspects and embodiments described herein may be applied to other environments, such as other standards, applications, hardware, software, etc.
[0050] For HEVC, the AMVP mode of VVC draft 5 involves two motion vector prediction candidates that can be used to predict the motion information of a CU in AMVP mode. The spatial and temporal motion vector prediction candidates available in the AMVP mode of VVC draft 5 are Figure 5 However, the process of constructing the set of AMVP candidates is slightly different. First, rounding of the calculated candidates to the motion vector resolution associated with the current CU is performed. This involves the Adaptive Motion Vector Resolution (AMVR) tool of VVC draft 5, which will be discussed further below. In addition, additional candidates that are history-based motion vector predictors (HMVPs) are considered after the TMVP candidates to populate the AMVP candidate list. Figure 6 The described process is shown in FIG, which shows an example of calculating motion vector prediction candidates in the AMVP mode of VVC draft 5. As can be seen, in the AMVP candidate list construction process, HMVP candidates are considered after the temporal motion vector prediction value (TMVP). It can also be noted that in Figure 6 In the final step of the processing, the motion vector prediction candidate constructed in AMVP is rounded to the precision level associated with the AMVR (Adaptive Motion Vector Resolution) associated with the currently processed decoding unit. VVC's AMVR decoding tool will be discussed further below.
[0051] As mentioned above, codecs and video compression tools other than HEVC can provide different or additional features. For example, VVC draft 5 supports additional motion models beyond HEVC's translation model to provide improved temporal prediction. To facilitate this additional model, PUs can be spatially partitioned into sub-PUs, and a more robust model can be used to assign dedicated motion vectors to each PU. For example, the motion model introduced in VVC is affine motion compensation, which involves using an affine model to represent the motion vectors in a CU.
[0052] Figure 7 The values for two control points ( Figure 7 ) or three control points ( Figure 7 The affine motion field for two control points (also called a 4-parameter affine model) involves the following motion vector component values for each position (x, y) within the considered block:
[0053]
[0054] Equation 1: 4-parameter affine model for generating motion fields within a CU for prediction
[0055] Where (v 0x , v 0y ) and (v 1x , v 1y ) are the so-called control point motion vectors used to generate the affine motion field, (v 0x , v 0y ) is the upper left corner control point of the motion vector, and (v 1x , v 1y ) is the upper right control point of the motion vector.
[0056] Equation 1 can be rewritten as Equation 2 based on the affine model parameters a and b as defined by Equation 3:
[0057]
[0058] Equation 2: 4-parameter affine model for representing the sub-block-based motion field of a CU decoded in 4-parameter affine mode
[0059]
[0060] Equation 3: Affine model parameters for a 4-parameter affine model
[0061] A model with three control points (called a 6-parameter affine motion model) can also be used to represent the sub-block-based motion field for a given coding unit. In the case of a 6-parameter affine model, the motion field is calculated as in Equations 4, 5, and 6:
[0062]
[0063] Equation 4: 6-parameter affine motion field, used to represent the sub-block-based motion field of a CU decoded in 6-parameter affine mode
[0064]
[0065] Equation 5: Alternative representation of the 6-parameter affine motion field
[0066]
[0067] Equation 6: Affine model parameters for a 6-parameter affine model
[0068] In practice, to keep the complexity reasonable, the same motion vector is calculated for each sample of a 4×4 sub-block (sub-CU) of the considered CU, e.g. Figure 8 As shown in , it shows the two control point model ( Figure 8 ) and three control point models ( Figure 8 The affine motion vector field for a 4×4 sub-CU (right side) is shown. At the center of each sub-block, an affine motion vector is calculated based on the control point motion vectors. The resulting motion vector (MV) is expressed with 1 / 16 pixel accuracy.
[0069] In VVC draft 5, CUs with a size greater than 8x8 can be predicted in affine AMVP mode. This is signaled via a flag in the bitstream for CU-level decoding. The generation of the affine motion field for the inter CU involves determining the control point motion vector (CPMV), which is obtained by the decoder by adding the motion vector difference plus the control point motion vector prediction (CPMVP). The CPMVP is a motion vector pair (for a 4-parameter affine model specified by 2 control point motion vectors) or a triple (for a 6-parameter affine model specified by 3 control point motion vectors) used as the predictor of the CPMV of the current CU.
[0070] The CPMVP of the current CU may be inherited from an affine neighboring CU (as in affine merge mode). Figure 9 shows the spatial location from which the inherited CPMVP is retrieved, Figure 9 Shows the spatial locations (A0, A1, B0, B1, and B2) from which CPMV prediction candidates are retrieved to predict the affine model of the current CU under affine AMVP, e.g., in VTM draft 3. The inherited CPMVP is considered valid if its reference picture is equal to the reference picture of the current CU.
[0071] Affine AMVP candidates can also be constructed from non-affine motion vectors. This is achieved by using motion vector prediction values in each set (A, B, C) and (D, E) and / or (F, G), respectively, as Figure 10 , which shows motion vector prediction for an affine AMVP CU such as in VTM. Figure 11 A more detailed example of constructing the CPVMP candidate list in affine AMVP is shown in .
[0072] Another example of a motion compensation tool is the Intra Block Copy (IBC) tool, included in VVC draft 5. It is a CU-level intra coding mode that involves assigning a so-called block vector to the CU in question. The block vector indicates the displacement between the current CU and a reference block, which is a reconstructed building block of samples within the current picture.
[0073] The Adaptive Motion Vector Resolution (AMVR) decoding tool of VVC Draft 5 provides a means for signaling the motion vector difference (MVd) between a CU's motion vector and its motion vector prediction value at various levels of accuracy (or precision, or resolution). In HEVC, the slice header flag (use_integer_mv_flag) indicates whether the MVd is coded with quarter-pixel accuracy or integer-pixel accuracy. This means that all CUs in the slice under consideration are coded / decoded with the same MVd accuracy level. In VVC Draft 5, some CU-level information can be signaled to indicate the resolution of the CU's MVd information. AMVR can be applied to CUs coded in normal AMVP or affine AMVP mode.
[0074] In normal AMVP mode, the supported MVd resolution levels are quarter luma samples, integer luma samples, or four luma samples. In affine AMVP mode, the supported MVd resolution levels are quarter luma samples, integer luma samples, or 1 / 16 luma samples. Signaling of AMVR information on the CU level involves a first flag that indicates the use of quarter luma sample accuracy for MVd information. If quarter luma sample accuracy is not used, the second flag indicates the use of integer luma samples or 4 luma sample accuracy levels in normal AMVP mode. In the case of affine AMVP mode, the second flag indicates the use of integer luma samples or 1 / 16 luma sample accuracy levels. Finally, as Figure 6 As shown in the example shown in , the motion vector prediction value is rounded to the same AMVR precision as that of MVd to ensure that the reconstructed motion vector has the desired level of precision.
[0075] With respect to IBC (Intra Block Copy) intra coding mode, the motion vector precision can be either 1 luma sample resolution or 4 luma sample resolution, depending on the amvr_precision_flag signaled for IBC coded CUs.
[0076] Figure 12 shows an example of reconstruction of motion information of a CU in AMVP and AMVR modes according to VVC draft 5, more specifically, Figure 12The example shown in shows a process for reconstructing motion vectors for an input CU to be decoded in translation AMVP mode (as opposed to affine AMVP mode), where AMVR is active for this CU. This process occurs at the decoder side. The input to the process is the parsed decoding unit.
[0077] Figure 12 The process shown in starts by testing whether the current CU uses temporal prediction from reference pictures contained in the L0 reference picture list. If so, the following four steps are applied.
[0078] - Apply the inverse of the AMVR process. This involves converting the resolved MVd (motion vector difference) associated with the L0 reference picture to the internal precision used by the decoder to represent motion data.
[0079] - The next step involves calculating the AMVP candidate list for the reference picture L0 of the current CU. This follows from the reference picture L0. Figure 6 Once this is done, the motion vector prediction value of the current CU in the current reference picture list is known and recorded as AMVPCand[mvpIdx L0 ], where mvpIdx L0 is the index of the MV prediction value candidate for the current CU, and it is generated from the parsing process.
[0080] - The next step reconstructs the motion vector of the current CU in the current reference picture list as the sum of the MV prediction value and the resolved motion vector difference.
[0081] - Next, the reconstructed motion vectors are pruned according to the internal representation of motion vectors in the considered codec.
[0082] The next step of the process tests whether the current block uses temporal prediction from a reference picture contained in the reference picture list L1. If so, the same process as above is applied, but in the context of the reference picture list L1. Figure 12 As a result of the process, two reconstructed motion vectors MV are obtained. L0 and MV L1 for the current inter CU and then for motion compensated temporal prediction of the current CU.
[0083] In general, one aspect of the present disclosure relates to improving the compression efficiency of methods such as the VVC Draft 5 described above, for example, the decoding gain provided by the AMVR tool. For example, the existing AMVR tool for VVC Draft 5 involves selecting the resolution level of MVd decoding, typically for x and y coordinates. In general, at least one example of the embodiments described herein provides increased flexibility in the AMVR tool. As an example, at least one embodiment provides different levels of accuracy between the x and y values of the decoded motion vector difference (MVd). As another example, in the case of bi-predictive decoding units, at least one embodiment provides different levels of accuracy between L0 and L1 motion vector difference decoding.
[0084] In general, at least one example of the embodiments described herein (referred to herein as Embodiment 1) involves distinguishing the resolution of two components of a block vector in the context of an IBC-coded coding unit.
[0085] Figure 13 An example of parsing motion data for a CU coded in IBC mode is shown in , which involves parsing the MVd (motion vector difference) of the current CU, and then parsing a flag associated with the current CU (e.g., a flag designated as mvp_l0_flag) indicating which MV prediction value to use to code the block vector of the current CU. If AMVR is enabled for the current CU, then parsing a flag associated with the current CU, e.g., a flag designated as amvr_precision_flag, to obtain an indication of whether the default 1 luma sample MV resolution is used for the current CU or whether a 4 luma pixel resolution is used.
[0086] Figure 14 An example of decoder reconstruction of motion data of a CU encoded in IBC mode according to VVC draft 5 is shown. Figure 14 An example in
[15] includes applying a decoder-side AMVR process by converting the parsed MVd vector of the current CU from the decoded AMVR precision level to the intra precision for motion vector representation. The block vector of the current CU is then reconstructed by adding the MV prediction value for the current CU to the intra precision converted motion vector difference (MVd). The reconstructed MV is then pruned to provide a valid MV representation, for example, according to VVC draft 5.
[0087] Figure 15An example of parsing motion data for a CU decoded in IBC mode according to the present embodiment is shown. This example involves parsing the MVd (motion vector difference) of the current CU, and then parsing a flag associated with the current CU, such as a flag designated as mvp_l0_flag, indicating which motion vector predictor was used to decode the block vector of the current CU. Next, if AMVR is enabled for the current CU, a flag associated with the current CU, such as a flag designated as amvr_precision_flag, is parsed to obtain an indication of whether the default 1 luma sample MV resolution is used for the current CU or whether 4 luma pixel resolution is used for the x or y component of the decoded MVd. Next, if this flag is true and the x component of the decoded MVd is non-zero, a flag, such as a flag designated as amvr_flag_x, is parsed to obtain an indication of whether the x component of the current MVd is signaled with the default 1 luma sample precision level or with the 4 luma sample precision level. Next, if both flags (e.g., amvr_precision_flag and amvr_flag_x) are true, then a flag (e.g., a flag designated as amvr_flag_y) is parsed to obtain an indication of whether the y component of the current MVd is signaled with the default luma sample precision level of 1 or with a luma sample precision level of 4. Otherwise, if amvr_precision_flag is true and amvr_flag_x is false, then amvr_flag_y is inferred to be true. Finally, if amvr_precision_flag is false, both amvr_flag_x and amvr_flag_y are inferred to be false.
[0088] Figure 16 An example of decoder reconstruction of motion data for a CU decoded in IBC mode according to this embodiment is shown. This example includes applying a decoder-side AMVR process by first converting the x component of the parsed MVd vector of the current CU from the decoded AMVR precision level of the x component to the internal precision of the motion vector representation. Then, a decoder-side AMVR process is applied to the y component by converting the y component of the parsed MVd vector of the current CU from the decoded AMVR precision level of the y component to the internal precision of the motion vector representation. The block vector of the current CU is then reconstructed by adding the MV prediction value for the current CU to the internal precision converted motion vector difference (MVd). The reconstructed MV is then pruned to provide a valid MV representation, for example, according to VVC draft 5.
[0089] An example of a coding_unit syntax structure suitable for implementing the present example of the embodiment (Embodiment 1) is shown in Syntax Structure Listing 1 attached to this document.
[0090] In general, at least one other instance of the embodiments described herein (referred to herein as Embodiment 2) involves decoupling resolution of x and y motion vector components in the case of translational AMVP motion compensation.
[0091] Figure 17 The parsing of AMVR data for a CU coded in AMVP mode (e.g., according to VVC draft 5) is shown, which involves parsing a first flag associated with the current CU (e.g., a flag designated as amvr_flag), and then, if the first flag is true, parsing a second flag associated with the current CU (e.g., a flag designated as amvr_precision_flag). The first flag, such as the amvr_flag syntax element, indicates whether the MV data of the current CU is coded using the default 1 / 4 luma sample precision level. Then, if the first flag is true, the second flag associated with the current CU, such as amvr_precision_flag, is parsed to indicate whether the default 1 luma sample or 4 luma sample MV resolution is used for the current CU.
[0092] Figure 18 An example of parsing AMVR data for a CU decoded in AMVP mode according to the present embodiment is shown in , and involves parsing a flag associated with the current CU, e.g., a flag designated as amvr_flag. Then, if amvr_flag is true and the decoded x component of the L0 or L1 list MVd is non-zero, another flag, e.g., designated as amvr_flag_x, is parsed to indicate whether the x component of the current MVd is signaled with a default 1 / 4 luma sample precision level. Next, if both the flags amvr_precision_flag and amvr_flag_x are true, another flag, e.g., a flag designated as amvr_flag_y, is parsed to indicate whether the y component of the current MVd is signaled with a default 1 / 4 luma sample precision level. Otherwise, if amvr_precision_flag is true and amvr_flag_x is false, amvr_flag_y is inferred to be true. Other cases exist as follows:
[0093] - if amvr_precision_flag is false, then both amvr_flag_x and amvr_flag_y are inferred to be false;
[0094] - If amvr_flag is true, then amvr_precision_flag of the current CU indicates whether the default 1 luma sample or 4 luma sample MV resolution is used for the current CU;
[0095] - If amvr_flag_x is true, the precision level used for the x component of MVd is the precision level indicated by amvr_flag and amvr_precision_flag;
[0096] - If amvr_flag_y is true, the precision level used for the y component of MVd is the precision level indicated by amvr_flag and amvr_precision_flag.
[0097] Figure 19 An example of reconstructing the motion vectors of an input CU to be decoded in the translating AMVP mode (as opposed to the affine AMVP mode) in this embodiment is shown, where AMVR is active for this CU. This process occurs at the decoder side. The input to this process is the parsed decoding unit.
[0098] Figure 19 The example in starts by testing whether the current CU uses temporal prediction from reference pictures contained in the L0 reference picture list. If so, the following four steps are applied.
[0099] - Apply the inverse of the AMVR process on the x-component of the decoded MVd. This involves converting the x-component of the resolved MVd (motion vector difference) associated with the L0 reference picture from the level of precision used for the decoding of the x-component of the L0 list MVd to the internal precision used by the decoder to represent motion data.
[0100] Next, the inverse of the AMVR process is applied to the y component of the decoded MVd. This involves converting the y component of the resolved MVd (motion vector difference) associated with the L0 reference picture from the level of precision used for decoding the y component of the L0 list MVd to the internal precision used by the decoder to represent motion data.
[0101] - Next, calculate the AMVP candidate list of the reference picture L0 of the current CU. This follows the reference Figure 6 Once this is done, the motion vector prediction value of the current CU in the current reference picture list is known and is recorded as AMVPCand[mvpIdx L0 ], where mvpIdx L0 is the index of the MV prediction value candidate for the current CU, and it is generated from the parsing process.
[0102] - Next, the motion vector of the current CU in the current reference picture list is reconstructed as the sum of the MV prediction value and the resolved and component-precision converted motion vector difference.
[0103] - Finally, the reconstructed motion vectors are pruned according to the internal representation of motion vectors in the considered codec.
[0104] Next, Figure 19 The example of tests whether the current block employs temporal prediction from a reference picture contained in reference picture list L1. If yes, the same process as above is applied, but in the context of reference picture list L1.
[0105] result, Figure 19 The example obtains two reconstructed motion vectors MV L0 and MV L1 for the current inter CU and then for motion compensated temporal prediction of the current CU.
[0106] An example of a coding_unit syntax structure suitable for implementing the present example of the embodiment (Embodiment 2) is shown in Syntax Structure Listing 2 attached to this document.
[0107] In general, at least one other example of the embodiments described herein (referred to herein as Embodiment 3) involves decoupling resolution of x and y motion vector components in the case of affine motion compensation. Figure 20 An example of parsing AMVR data of a CU coded in affine AMVP mode, for example according to VVC draft 5, is shown in FIG. Figure 20 An example involves parsing a flag for the current CU, such as the designated amvr_flag, if at least one control point motion vector has a non-zero associated MVd. Then, if amvr_flag is true, parsing another flag for the current CU, such as the designated amvr_precision_flag. The amvr_flag syntax element indicates whether the default 1 / 4 luma sample precision level is used to decode the MV data of the current CU. Then, if amvr_flag is true, parsing the flag amvr_precision_flag for the current CU. This indicates whether 1 luma sample or 4 luma sample MV resolution is used for the current CU.
[0108] Figure 21 An example of parsing AMVR data of a CU decoded in affine AMVP mode according to this embodiment is shown in FIG. Figure 21An example involves parsing a flag associated with the current CU, such as amvr_flag. Then, if amvr_flag is true and the x component of the decoded L0 or L1 list MVd is nonzero for at least one CPMV of the current CU, another flag, such as amvr_flag_x, is parsed. This flag indicates whether the x component of the MVd for each CPMV is signaled at the default 1 / 4 luma sample precision level. Next, if both flags, amvr_precision_flag and amvr_flag_x, are true, another flag, such as amvr_flag_y, is parsed. This flag indicates whether the y component of the current MVd is signaled at the default 1 / 4 luma sample precision level for each CPMV. Otherwise, if amvr_precision_flag is true and amvr_flag_x is false, amvr_flag_y is inferred to be true. Finally, in the case where amvr_precision_flag is false, both amvr_flag_x and amvr_flag_y are inferred to be false.
[0109] Other examples of combinations of flag values and related interpretations according to this embodiment include the following. If amvr_flag is true, the amvr_precision_flag of the current CU indicates whether the default 1-luma sample or 4-luma sample MV resolution is used for the current CU. If amvr_flag_x is true, the precision level of the x component of the MVd used for each CPMV is the precision level indicated by amvr_flag and amvr_precision_flag. If amvr_flag_y is true, the precision level of the y component of the MVd used for each CPMV is the precision level indicated by amvr_flag and amvr_precision_flag.
[0110] Figure 22 An example of reconstruction of motion vectors for an input CU to be coded in affine AMVP mode (as opposed to translational AMVP mode) according to VVC draft 5, for which AMVR is valid, is shown. Figure 22 The processing involved in the example of occurs at the decoder side.The input of the processing is the decoded unit that has been parsed.
[0111] exist Figure 22 In
[15] , the process starts by testing whether the current affine AMVP CU uses temporal prediction from reference pictures contained in the L0 reference picture list. If so, the following four steps are applied.
[0112] For each CPMV in the L0 list, the inverse of the AMVR process is applied to the MVd. This involves converting the parsed MVd (motion vector difference) of each CPMV associated with the L0 reference picture from the level of precision used for decoding the L0 list MVd to the internal precision used by the decoder to represent motion data.
[0113] - Next, calculate the affine AMVP list of the CPVM candidate prediction values of the reference picture L0 of the current CU. This follows the reference Figure 11 Once this is done, the motion vector prediction value of the current CU in the current reference picture list is known and is recorded as AMVPCand[mvpIdx L0 ], where mvpIdx L0 is the index of the CPMV prediction value candidate for the current CU, and it is generated from the parsing process.
[0114] - Next, reconstruct the control point motion vector of the current CU in the current reference picture list as the sum of its corresponding MV prediction value and the resolved motion vector difference.
[0115] - Finally, the reconstructed control point motion vectors are pruned according to the internal representation of motion vectors in the considered codec and the affine motion field associated with the current reference picture list contained in the current CU is constructed according to the VVC Draft 5 specification.
[0116] Figure 22 The next step of the process in the example of involves testing whether the current block employs temporal prediction from a reference picture contained in the reference picture list L1. If so, the same process as above is applied, but in the context of the reference picture list L1.
[0117] As Figure 22 As a result of the processing of the example, two reconstructed affine motion vector fields of the current CU are obtained for the reference picture list L0 and the reference picture list L1, respectively, and the two reconstructed affine motion vector fields are then used for motion compensated temporal prediction of the current CU.
[0118] Figure 23 The decoder side motion data reconstruction of the CU of the affine AMVP decoding according to this embodiment is shown. Figure 23 The reconstruction of the motion vectors of an input CU to be decoded in affine AMVP mode (as opposed to translational AMVP mode) in this embodiment is shown, where AMVR is active for this CU. This process occurs at the decoder side. The input to this process is the parsed decoding unit.
[0119] Figure 23The process in the example of begins by testing whether the current affine AMVP CU employs temporal prediction from reference pictures contained in the L0 reference picture list. If so, the following four steps are applied.
[0120] The inverse of the AMVR process is applied to the x-component of the decoded MVd of each CPMV of the L0 list. This involves converting the x-component of the parsed MVd (motion vector difference) of each CPMV associated with the L0 reference picture from the level of precision used for the decoding of the x-component of the L0 list MVd to the internal precision used by the decoder to represent motion data.
[0121] - Next, the inverse of the AMVR process is applied to the y component of the decoded MVd of each CPMV of the L0 reference list. This involves converting the y component of the parsed MVd (motion vector difference) of each CPMV associated with the L0 reference picture from the level of precision used for the decoding of the y component of the L0 list MVd to the internal precision used by the decoder to represent motion data.
[0122] - Next, calculate the affine AMVP candidate list of the reference picture L0 of the current CU. This follows the reference Figure 11 Once this is done, the motion vector prediction value of the current CU in the current reference picture list is known and is recorded as AMVPCand[mvpIdx L0 ], where mvpIdx L0 is the index of the MV prediction value candidate for the current CU, and it is generated from the parsing process.
[0123] - Next, the control motion vector of the current CU in the current reference picture list is reconstructed as the sum of its corresponding CPMV prediction value and its corresponding resolved motion vector difference for each CPMV of the current CU.
[0124] -Finally, the reconstructed CPMV is pruned and the affine motion field of the current CU is generated with respect to the L0 reference picture.
[0125] The next step of the process tests whether the current block employs temporal prediction from a reference picture contained in reference picture list L1. If so, the above process is applied, but in the context of reference picture list L1.
[0126] Figure 23 The result of the example processing is two reconstructed motion fields of list L0 and list L1 obtained for the current affine AMVP CU. These reconstructed motion fields are then used for motion compensated temporal prediction of the current CU.
[0127] An example of a decoding unit syntax structure suitable for implementing the present example (Embodiment 3) of the embodiment is shown in Syntax Structure List 3 attached to this document.
[0128] In general, at least one other example of the embodiments described herein (referred to herein as Embodiment 4) relates to decoupling motion vector resolutions in L0 and L1 reference picture lists in the context of AMVP translational motion compensation. One aspect relates to decoupling AMVR accuracy levels between L0 and L1 reference lists. Figure 24 Described is the parsing of AMVR data for a CU coded in AMVP mode according to this embodiment.
[0129] exist Figure 24 In
[15] , a flag, designated, for example, amvr_flag, is first parsed for the current CU. Then, if amvr_flag is true and the decoded MVd of list L0 is nonzero, another flag, designated, for example, amvr_flag_L0, is parsed. This flag indicates whether the current MVd of list L0 is signaled at the default 1 / 4 luma sample precision level. Next, if both flags, amvr_precision_flag and amvr_flag_L0, are true, another flag, designated, for example, amvr_flag_L1, is parsed. This flag indicates whether the current MVd associated with list L1 is signaled at the default 1 / 4 luma sample precision level. Otherwise, if amvr_precision_flag is true and amvr_flag_L0 is false, amvr_flag_L1 is inferred to be true. Finally, if amvr_precision_flag is false, both amvr_flag_L0 and amvr_flag_L1 are inferred to be false.
[0130] Other examples of flag value combinations and related interpretations according to this embodiment are as follows. If amvr_flag is true, the amvr_precision_flag of the current CU indicates whether the default 1-luma sample or 4-luma sample MV resolution is used for the current CU. If amvr_flag_L0 is true, the precision level of the MVd used for list L0 is the precision level indicated by amvr_flag and amvr_precision_flag. If amvr_flag_L1 is true, the precision level of the MVd used for list L1 is the precision level indicated by amvr_flag and amvr_precision_flag.
[0131] Figure 25An example of reconstruction of motion vectors for an input CU to be coded in translational AMVP mode (as opposed to affine AMVP mode) according to this example of an embodiment is shown, where AMVR is active for this CU. This reconstruction occurs at the decoder side. The input is the decoded unit that has been parsed. Figure 25 The example in starts by testing whether the current CU uses temporal prediction from the reference pictures contained in the L0 reference picture list. If so, the following four steps are applied.
[0132] The inverse of the AMVR process is applied to the coded MVd associated with the L0 list. This involves converting the resolved MVd (motion vector difference) associated with the L0 reference picture from the level of precision used for the coding of the L0 list MVd to the internal precision used by the decoder to represent motion data.
[0133] - Next, calculate the AMVP candidate list of the reference picture L0 of the current CU. This follows the reference Figure 6 Once this is done, the motion vector prediction value of the current CU in the current reference picture list is known and is recorded as AMVPCand[mvpIdx L0 ], where mvpIdx L0 is the index of the MV prediction value candidate for the current CU, and it is generated from the parsing process.
[0134] - Next, reconstruct the motion vector of the current CU in the current reference picture list as the sum of the MV prediction value and the resolved motion vector difference.
[0135] - Finally, the reconstructed motion vectors are pruned according to the internal representation of motion vectors in the considered codec.
[0136] Figure 25 The next step in the example of involves testing whether the current block employs temporal prediction from a reference picture contained in the reference picture list L1. If so, the same process as described above is applied, but in the context of the reference picture list L1. As a result, Figure 25 The example provides two reconstructed motion vectors MV L0 and MV L1 for the current inter CU and then for motion compensated temporal prediction of the current CU.
[0137] An example of a coding_unit syntax structure suitable for implementing the present example of the embodiment (Embodiment 4) is shown in Syntax Structure Listing 4 attached to this document.
[0138] In general, at least one other example of the embodiments described herein (referred to herein as Embodiment 5) relates to decoupling motion vector resolution in L0 and L1 reference picture lists in the case of affine motion compensation. One aspect relates to decoupling AMVR precision levels between the L0 reference list and the L1 reference list in the case of affine AMVP-coded CUs. Figure 26 , an example of parsing AMVR data of a CU decoded in affine AMVP mode according to this embodiment is shown.
[0139] Figure 26 An example involves first analyzing a flag associated with the current CU (e.g., a flag designated as amvr_flag). Then, if amvr_flag is true and the decoded L0 MVd is nonzero for at least one CPMV of the current CU, another flag, designated, for example, amvr_flag_L0, is parsed. This flag indicates whether the MVd for each CPMV for reference picture list L0 is signaled at the default 1 / 4 luma sample precision level. Next, if both the flags amvr_precision_flag and amvr_flag_L0 are true, the flag amvr_flag_L1 is parsed. This flag indicates whether the current MVd is signaled at the default 1 / 4 luma sample precision level for reference picture list L1. Otherwise, if amvr_precision_flag is true and amvr_flag_L0 is false, amvr_flag_L1 is inferred to be true. Finally, if amvr_precision_flag is false, then both amvr_flag_L0 and amvr_flag_L1 are inferred to be false.
[0140] The following are examples of other combinations of flag values and their associated interpretations. If amvr_flag is true, the amvr_precision_flag of the current CU is parsed. This indicates whether the default 1-luma sample or 4-luma sample MV resolution is used for the current CU. If amvr_flag_L0 is true, the precision level of the MVd for each CPMV associated with reference picture list L0 is the precision level indicated by amvr_flag and amvr_precision_flag. If amvr_flag_L1 is true, the precision level of the MVd for each CPMV associated with reference picture list L1 is the precision level indicated by amvr_flag and amvr_precision_flag.
[0141] Figure 27 An example of reconstruction of the motion vectors of an input CU to be decoded in affine AMVP mode (as opposed to translational AMVP mode) in this embodiment is shown, where AMVR is active for this CU. Figure 27 The processing associated with the example of occurs at the decoder side. The input to this process is the parsed decoding unit.
[0142] Figure 27 The example process of starts by testing whether the current affine AMVP CU employs temporal prediction from reference pictures contained in the L0 reference picture list. If so, the following four steps are applied.
[0143] For each CPMV of the L0 list, the inverse of the AMVR process is applied to the MVd. This involves converting the parsed MVd (motion vector difference) of each CPMV associated with the L0 reference picture from the level of precision used for decoding the L0 list MVd to the internal precision used by the decoder to represent motion data.
[0144] - Next, calculate the AMVP candidate list of the reference picture L0 of the current CU. This follows the reference Figure 11 Once this is done, the motion vector prediction value for the current CU in the current reference picture list is known and is designated as AMVPCand[mvpIdx L0 ], where mvpIdx L0 is the index of the CPMV prediction value candidate for the current CU, and it is generated from the parsing process.
[0145] - Next, the control point motion vector of the current CU in the current reference picture list is determined as the sum of the motion vector differences between the CPMV prediction value and the resolution of each CPMV of the current CU.
[0146] -Finally, the reconstructed CMVS is pruned and the reconstructed affine motion field of the current CU is generated relative to the L0 reference picture.
[0147] The next step of the process tests whether the current affine AMVP block uses temporal prediction from reference pictures contained in reference picture list L1. If so, the above process is applied, but in the context of reference picture list L1. Figure 27 As a result of the processing of the example, two reconstructed motion fields of list L0 and list L1 are obtained for the current affine AMVP CU, and are then used for motion compensated temporal prediction of the current CU.
[0148] An example of a coding_unit syntax structure suitable for implementing the present example of the embodiment (Embodiment 5) is shown in Syntax Structure Listing 5 attached to this document.
[0149] Systems involving video encoding and / or decoding according to one or more embodiments described herein may provide one or more of the following non-limiting feature examples, alone or in various combinations:
[0150] • Level of precision in the decoding of motion vector differences (or MV residuals) A distinction is made between the x-component and the y-component of the decoded motion vector differences.
[0151] • In case of bi-predicted blocks without using symmetric motion vector difference coding, the level of precision in motion vector difference (or MV residual) coding is differentiated between L0 and L1 reference picture lists.
[0152] • In case the considered block is coded by an intra block copy mode such as VVC, the level of precision in the coding of the motion vector difference (or MV residual) distinguishes between the x-component and the y-component of the decoded motion vector difference.
[0153] • In case the considered block is coded by an inter-translational AMVP mode such as VVC, the level of precision in the coding of the motion vector difference (or MV residual) distinguishes between the x-component and the y-component of the decoded motion vector difference.
[0154] • In case the considered block is coded by the inter-affine AMVP mode of VVC, the level of precision in the coding of the motion vector difference (or MV residual) distinguishes between the x-component and the y-component of the decoded motion vector difference.
[0155] • In case of bi-predicted blocks, the level of precision in the coding of motion vector differences (or MV residuals) is differentiated between the L0 and L1 reference picture lists, where symmetric motion vector difference coding is not used and the considered block is coded with the inter-translation AMVP mode of VVC.
[0156] • In the case of bi-predicted blocks, the level of precision in the coding of motion vector differences (or MV residuals) is differentiated between the L0 and L1 reference picture lists, where symmetric motion vector difference coding is not used, and where the considered block is coded by the inter-affine AMVP mode of VVC.
[0157] This document describes various examples of embodiments, features, models, methods, and the like. Many of these examples are described as being specific and, at least to illustrate individual characteristics, are often described in a manner that may appear restrictive. However, this is for clarity of description and does not limit the application or scope. In fact, the various examples of embodiments, features, and the like described herein can be combined and interchanged in various ways to provide additional examples of the embodiments.
[0158] In general, examples of the embodiments described and contemplated in this document may be implemented in many different forms. Figure 1 and 2 and the following Figure 28Some embodiments are provided, but other embodiments are contemplated and are not intended to be construed as limiting the scope of the invention. Figure 1 、 2 The discussion of 28 does not limit the breadth of implementation. At least one embodiment generally provides examples related to video encoding and / or decoding, and at least one other embodiment generally relates to transmitting a generated or encoded bitstream or signal. These and other embodiments can be implemented as methods, apparatus, computer-readable storage media having stored thereon instructions for encoding or decoding video data according to any of the described methods, and / or computer-readable storage media having stored thereon a bitstream or signal generated according to any of the described methods.
[0159] The terms HDR (high dynamic range) and SDR (standard dynamic range) are used in this disclosure. Those terms generally convey a specific value of dynamic range to those of ordinary skill in the art. However, additional embodiments are also intended in which references to HDR are understood to mean "higher dynamic range" and references to SDR are understood to mean "lower dynamic range." Such additional embodiments are not constrained by any specific value of dynamic range that may often be associated with the terms "high dynamic range" and "standard dynamic range."
[0160] Various methods are described herein, and each method includes one or more steps or actions for implementing the described method. Unless a specific order of steps or actions is required for proper operation of the method, the order and / or use of specific steps and / or actions may be modified or combined.
[0161] Various methods and other aspects described in this document may be used to modify modules of a video encoder and / or decoder, e.g. Figure 1 The motion compensation and / or motion estimation modules 170 and 175 of the encoder 100 shown in FIG. Figure 2 275 of the decoder 200 shown in FIG. 276 . Furthermore, aspects of this disclosure are not limited to VVC or HEVC and may be applied, for example, to other standards and recommendations (whether previously existing or developed in the future), as well as extensions of any such standards and recommendations (including VVC and HEVC). Unless otherwise indicated or technically excluded, the aspects described herein may be used alone or in combination.
[0162] For example, various numerical values are used in this document. The specific values are for illustrative purposes, and the described aspects are not limited to these specific values.
[0163] Figure 28A block diagram of an example of a system in which various aspects and embodiments can be implemented is shown. System 1000 can be implemented as a device including the various components described below, and is configured to perform one or more aspects described herein. Examples of such devices include, but are not limited to, various electronic devices such as personal computers, laptop computers, smart phones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected home appliances, and servers. The elements of system 1000 can be implemented individually or in combination in a single integrated circuit, multiple ICs, and / or discrete components. For example, in at least one embodiment, the processing and encoder / decoder elements of system 1000 are distributed over multiple ICs and / or discrete components. In various embodiments, system 1000 is communicatively coupled to other similar systems or other electronic devices via, for example, a communication bus or by dedicated input and / or output ports. In various embodiments, system 1000 is configured to implement one or more aspects described herein.
[0164] System 1000 includes at least one processor 1010, which is configured to execute instructions loaded therein, for implementing various aspects such as described herein. Processor 1010 may include embedded memory, input / output interface and various other circuits known in the art. System 1000 includes at least one memory 1020 (e.g., volatile memory device and / or non-volatile memory device). System 1000 includes storage device 1040, which may include non-volatile memory and / or volatile memory, including but not limited to EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash memory, magnetic disk drive and / or optical disk drive. As a non-limiting example, storage device 1040 may include internal storage device, attached storage device and / or network accessible storage device.
[0165] System 1000 includes an encoder / decoder module 1030, which is configured to process data to provide encoded video or decoded video, for example, and may include its own processor and memory. Encoder / decoder module 1030 represents a module (or modules) that may be included in a device to perform encoding and / or decoding functions. As is known, a device may include one or both encoding and decoding modules. In addition, encoder / decoder module 1030 may be implemented as a separate element of system 1000 or may be incorporated into processor 1010 as a combination of hardware and software as known to those skilled in the art.
[0166] Program code to be loaded onto the processor 1010 or the encoder / decoder 1030 to perform various aspects described in this document may be stored in the storage device 1040 and subsequently loaded onto the memory 1020 for execution by the processor 1010. According to various embodiments, one or more of the processor 1010, the memory 1020, the storage device 1040, and the encoder / decoder module 1030 may store one or more of various items during the execution of the processes described herein. These stored items may include, but are not limited to, input video, decoded video or portions of decoded video, bitstreams or signals, matrices, variables, and intermediate or final results from the processing of equations, formulas, operations, and operational logic.
[0167] In several embodiments, memory within the processor 1010 and / or encoder / decoder module 1030 is used to store instructions and provide working memory for processing required during encoding or decoding. However, in other embodiments, memory external to the processing device (e.g., the processing device may be the processor 1010 or the encoder / decoder module 1030) is used for one or more of these functions. The external memory may be memory 1020 and / or storage device 1040, such as dynamic volatile memory and / or non-volatile flash memory. In several embodiments, the external non-volatile flash memory is used to store the operating system of the television. In at least one embodiment, fast external dynamic volatile memory, such as RAM, is used as working memory for video encoding and decoding operations, such as MPEG-2, HEVC, or VVC (Versatile Video Coding).
[0168] As shown in block 1130, input to the elements of system 1000 may be provided through various input devices. Such input devices include, but are not limited to, (i) an RF section that receives an RF signal transmitted over the air, for example, by a broadcaster, (ii) a composite input terminal, (iii) a USB input terminal, and / or (iv) an HDMI input terminal.
[0169] In various embodiments, the input device of block 1130 has associated corresponding input processing elements known in the art. For example, the RF section may be associated with elements for: (i) selecting a desired frequency (also known as selecting a signal, or band-limiting a signal to a frequency band), (ii) down-converting the selected signal, (iii) again band-limiting to a narrower frequency band to select a signal frequency band (which may be referred to as a channel in some embodiments), (iv) demodulating the down-converted and band-limited signal, (v) performing error correction, and (vi) demultiplexing to select the desired data packet stream. The RF section of various embodiments includes one or more elements to perform these functions, such as a frequency selector, a signal selector, a band limiter, a channel selector, a filter, a down-converter, a demodulator, an error corrector, and a demultiplexer. The RF section may include a tuner that performs various of these functions, including, for example, down-converting a received signal to a lower frequency (e.g., an intermediate frequency or near-baseband frequency) or to baseband. In a set-top box embodiment, the RF part and its relevant input processing element receive the RF signal that sends by wired (for example, cable) medium, and by filtering, down-conversion and filtering to the frequency band of expectation again to perform frequency selection.Various embodiments rearrange the order of above-mentioned (and other) element, remove some in these elements, and / or add other element that performs similar or different functions.Adding element can be included in and inserts element between existing element, for example, inserts amplifier and analog to digital converter.In various embodiments, the RF part comprises antenna.
[0170] In addition, the USB and / or HDMI terminals may include corresponding interface processors for connecting the system 1000 to other electronic devices via USB and / or HDMI connections. It should be understood that various aspects of input processing, such as Reed-Solomon error correction, may be implemented, for example, within a separate input processing IC or within the processor 1010. Similarly, various aspects of USB or HDMI interface processing may be implemented within a separate interface IC or within the processor 1010. The demodulated, error-corrected, and demultiplexed streams are provided to various processing elements, including, for example, the processor 1010 and the encoder / decoder 1030, which operate in conjunction with memory and storage elements to process the data streams for presentation on an output device.
[0171] The various components of system 1000 may be disposed within an integrated housing. Within the integrated housing, the various components may be interconnected and transmit data therebetween using a suitable connection arrangement 1140, such as an internal bus known in the art, including an I2C bus, wiring, and a printed circuit board.
[0172] System 1000 includes a communication interface 1050 that enables communication with other devices via a communication channel 1060. Communication interface 1050 may include, but is not limited to, a transceiver configured to send and receive data over communication channel 1060. Communication interface 1050 may include, but is not limited to, a modem or a network card, and communication channel 1060 may be implemented, for example, within a wired and / or wireless medium.
[0173] In various embodiments, data is streamed to system 1000 using a Wi-Fi network such as IEEE 802.11. Wi-Fi signals for these embodiments are received via a communication channel 1060 and communication interface 1050 suitable for Wi-Fi communication. Communication channel 1060 for these embodiments is typically connected to an access point or router that provides access to external networks, including the Internet, to allow streaming applications and other over-the-top communications. Other embodiments use a set-top box that delivers data via an HDMI connection to input block 1130 to provide streaming data to system 1000. Still other embodiments use an RF connection to input block 1130 to provide streaming data to system 1000.
[0174] System 1000 can provide output signals to various output devices, including display 1100, speakers 1110, and other peripherals 1120. In various examples of various embodiments, other peripherals 1120 include one or more of the following: a stand-alone DVR, a disc player, a stereo system, a lighting system, and other devices that provide functionality based on the output of system 1000. In various embodiments, control signals are transmitted between system 1000 and display 1100, speakers 1110, or other peripherals 1120 using signaling such as AV Link, CEC, or other communication protocols that enable device-to-device control with or without user intervention. Output devices can be communicatively coupled to system 1000 via dedicated connections through respective interfaces 1070, 1080, and 1090. Alternatively, output devices can be connected to system 1000 via communication interface 1050 using communication channel 1060. Display 1100 and speakers 1110 can be integrated into a single unit in an electronic device (e.g., a television) along with the other components of system 1000. In various embodiments, the display interface 1070 includes a display driver, such as a timing controller (Tcon) chip.
[0175] For example, if the RF portion of input 1130 is part of a separate set-top box, the display 1100 and speaker 1110 may alternatively be separate from one or more of the other components. In various embodiments where the display 1100 and speaker 1110 are external components, the output signal may be provided via a dedicated output connection, such as an HDMI port, a USB port, or a COMP output.
[0176] These embodiments may be implemented by computer software implemented by the processor 1010 or hardware, or a combination of hardware and software. As a non-limiting example, embodiments may be implemented by one or more integrated circuits. The memory 1020 may be of any type suitable for the technical environment and may be implemented using any appropriate data storage technology, such as, as non-limiting examples, optical memory devices, magnetic memory devices, semiconductor-based memory devices, fixed memory, and removable memory. The processor 1010 may be of any type suitable for the technical environment and may include, as non-limiting examples, one or more of the following: a microprocessor, a general-purpose computer, a special-purpose computer, and a processor based on a multi-core architecture.
[0177] Figure 29 Another example of an embodiment is shown in Figure 29 In 2910, a coding mode associated with a current coding unit including picture information is determined. For example, the coding mode may include intra block coding (IBC) mode, translation adaptive motion vector prediction (AMVP) mode, affine AMVP mode, etc., as described above. Then, at 2920, a first level of precision and a second level of precision are determined based on the coding mode, for example, based on various flags and tests of the values of the flags, as described above. The first level of precision is associated with a first portion of motion vector information. For example, the motion vector information may include a residual (or motion vector difference (MVd)), and the first level of precision associated with the first portion may be the accuracy of the x-component of the residual or MVd. The second level of precision is associated with a second portion of the motion vector information. For example, the second level of precision associated with the second portion may be the accuracy of the y-component of the residual or MVd. As examples, the level of precision or accuracy may be values such as 1 / 16 luma samples, 1 / 4 luma samples, an integer number of luma samples, or four luma samples. In affine AMVP mode, the supported MVd resolution levels can be quarter luma samples or integer luma samples. At 2930, a motion vector associated with the current coding unit is obtained by combining a motion vector predictor and a motion vector residual based on the motion vector information and the first and second precision levels, for example, using the defined precision or accuracy as described above. At 2940, at least a portion of the picture information is decoded based on the decoding mode and the motion vector to generate decoded picture information.
[0178] Figure 30 Another example of an embodiment is shown, Figure 30 In 3010, a coding mode associated with a current coding unit including picture information is determined. For example, the coding modes may include intra block coding (IBC) mode, translation adaptive motion vector prediction (AMVP) mode, affine AMVP mode, and the like, as described above. Then, at 3020, a first level of precision and a second level of precision are determined based on the coding mode, for example, based on various flags and testing of the flag values as described above. The first level of precision is associated with a first portion of motion vector information. For example, the motion vector information may be a residual or motion vector difference (MVd), and the first level of precision associated with the first portion may be the accuracy of the x-component of the residual. The second level of precision is associated with the second portion of the motion vector information (e.g., the accuracy of the y-component of the residual). As examples, the level of precision or accuracy may be values such as 1 / 16 luma samples, 1 / 4 luma samples, an integer luma sample, or four luma samples. In affine AMVP mode, the supported MVd resolution levels are quarter luma samples or an integer luma sample. At 3030, a motion vector associated with the current coding unit is obtained based on the motion vector information and the first and second levels of precision (e.g., combining a motion vector predictor with a motion vector residual using defined precision or accuracy as described above). At 3040, at least a portion of the picture information is encoded based on the coding mode and the motion vector to generate encoded picture information.
[0179] In another example of an embodiment, the coding mode, and the first and second precision levels are determined as described above. However, in this example, the motion vector information may include a first motion vector residual corresponding to a first reference picture list (e.g., L0 MVd) and a second motion vector residual corresponding to a second reference picture list (e.g., L1 MVd). In this case, the first precision level is associated with the portion of the motion vector information including the first motion vector residual, and the second precision level is associated with the portion of the motion vector information including the second motion vector residual. A first motion vector is obtained based on the first motion vector residual and the first precision level, and a second motion vector is obtained based on the second motion vector residual and the second precision level. The image information is then decoded or encoded based on the first motion vector, the second motion vector, and the coding mode.
[0180] In the present disclosure, various implementations relate to decoding. As used in this application, "decoding" may include, for example, all or part of the processing performed on a received coded sequence to produce a final output suitable for display. In various embodiments, such processes include one or more of the processes typically performed by a decoder, such as entropy decoding, inverse quantization, inverse transform, and differential decoding. In various embodiments, such processes also or alternatively include processes performed by the decoder of the various embodiments described in this application, such as extracting a picture from a (packed) picture of a tile, determining an upsampling filter to use and then upsampling the picture, and flipping the picture back to its intended orientation.
[0181] As a further example, in one embodiment, "decoding" refers only to entropy decoding, in another embodiment, "decoding" refers only to differential decoding, and in another embodiment, "decoding" refers to a combination of entropy decoding and differential decoding. Whether the phrase "decoding process" is intended to refer specifically to a subset of operations or generally to a broader decoding process will be clear based on the context of the specific description and is believed to be well understood by those skilled in the art.
[0182] Furthermore, various implementations involve encoding. In a manner similar to the discussion above regarding "decoding," "encoding," as used in this application, may include, for example, all or part of the processing performed on an input video sequence to produce an encoded bitstream or signal. In various embodiments, such processes include one or more processes typically performed by an encoder, such as partitioning, differential encoding, transforms, quantization, and entropy encoding. In various embodiments, such processes also or alternatively include processes performed by the encoders of the various implementations described herein.
[0183] As a further example, in one embodiment, "encoding" refers only to entropy encoding, in another embodiment, "encoding" refers only to differential encoding, and in another embodiment, "encoding" refers to a combination of differential encoding and entropy encoding. Whether the phrase "encoding process" is intended to refer specifically to a subset of operations or generally to a broader encoding process will become clear based on the context of the specific description and is believed to be well understood by those skilled in the art.
[0184] Note that the syntax elements used here are descriptive terms. Therefore, they do not exclude the use of other syntax element names.
[0185] When a figure is presented as a flow chart, it should be understood that it also provides a block diagram of the corresponding apparatus. Similarly, when a figure is presented as a block diagram, it should be understood that it also provides a flow chart of the corresponding method / process.
[0186] Various embodiments relate to rate-distortion optimization. In particular, during the encoding process, a balance or trade-off between rate and distortion is often considered, often given a constraint on computational complexity. Rate-distortion optimization is typically formulated as minimizing a rate-distortion function, which is a weighted sum of rate and distortion. There are different approaches to solving the rate-distortion optimization problem. For example, these approaches can be based on extensive testing of all coding options, including all considered modes or decoding parameter values, with a complete evaluation of their decoding costs and associated distortion in the reconstructed signal after decoding and encoding. Faster approaches can also be used to save coding complexity, particularly by computing approximate distortion based on a prediction or prediction residual signal rather than the reconstructed signal. A hybrid of these two approaches can also be used, for example by using approximate distortion for only some possible coding options and the full distortion for others. Other approaches evaluate only a subset of possible coding options. More generally, many approaches employ any of a variety of techniques to perform optimization, but optimization does not necessarily require a complete evaluation of both coding cost and associated distortion.
[0187] The implementations and aspects described herein can be implemented in, for example, a method or process, a device, a software program, a data stream, or a signal. Even if discussed only in the context of a single form of implementation (e.g., discussed only as a method), the implementation of the features discussed can also be implemented in other forms (e.g., a device or program). For example, an apparatus can be implemented with appropriate hardware, software, and firmware. The method can be implemented in, for example, a processor, which generally refers to a processing device, including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. The processor also includes communication equipment, such as a computer, a cellular phone, a portable / personal digital assistant ("PDA"), and other equipment that facilitates information communication between end users.
[0188] Reference to "one embodiment," "an embodiment," or "an implementation," or "an implementation," and other variations thereof, means that a particular feature, structure, characteristic, etc., described in connection with the embodiment is included in at least one embodiment. Thus, the appearances of the phrases "in one embodiment," "in an embodiment," or "in an implementation," or "in an implementation," and any other variations thereof, in various places throughout this document are not necessarily all referring to the same embodiment.
[0189] Additionally, this document may refer to "obtaining" various information. Obtaining information may include, for example, one or more of: determining the information, estimating the information, calculating the information, predicting the information, or retrieving the information from a memory.
[0190] Additionally, this document may refer to “accessing” various information. Accessing information may include, for example, one or more of: receiving information, retrieving information (e.g., from memory), storing information, moving information, copying information, calculating information, determining information, predicting information, or estimating information.
[0191] Additionally, this document may refer to "receiving" various types of information. Like "accessing," receiving is intended to be a broad term. Receiving information can include, for example, one or more of accessing information or retrieving information (e.g., from a memory device). Furthermore, during operations such as storing information, processing information, sending information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information, "receiving" is often involved in one way or another.
[0192] It should be understood that use of any of the following “ / ,” “and / or,” “one or more of,” and “at least one of,” for example, in the context of “A / B,” “A and / or B,” “one or more of A or B,” and “at least one of A and B,” is intended to encompass selection of only the first-listed option (A), or only the second-listed option (B), or both options (A and B). As a further example, in the context of “A, B and / or C,” “one or more of A, B, or C,” and “at least one of A, B, and C,” such wording is intended to include selection of only the first-listed option (A), or only the second-listed option (B), or only the third-listed option (C), or only the first and second-listed options (A and B), or only the first and third-listed options (A and C), or only the second and third-listed options (B and C), or all three options (A, B, and C). As will be apparent to one of ordinary skill in this and related arts, this can be extended to multiple items listed.
[0193] Furthermore, as used herein, the term "signal" specifically refers to indicating something to a corresponding decoder. For example, in certain embodiments, an encoder signals a specific one of multiple parameters for refinement. Thus, in one embodiment, the same parameters are used on both the encoder and decoder sides. Thus, for example, the encoder can send (explicit signaling) specific parameters to the decoder so that the decoder can use the same specific parameters. Conversely, if the decoder already has the specific parameters along with other parameters, signaling can be used instead of sending them (implicit signaling), simply allowing the decoder to know and select the specific parameters. By avoiding the transmission of any actual functionality, bit savings are achieved in various embodiments. It should be understood that signaling can be implemented in various ways. For example, in various embodiments, one or more syntax elements, flags, etc. are used to signal information to the corresponding decoder. Although the foregoing refers to the verb form of the term "signal," the term "signal" can also be used as a noun in this document.
[0194] As will be apparent to one of ordinary skill in the art, implementations can generate various signals formatted to carry information that can be stored or transmitted, for example. The information can include, for example, instructions for performing a method, or data generated by one of the described implementations. For example, a signal can be formatted to carry a bitstream or signal of the described embodiments. Such a signal can be formatted as, for example, an electromagnetic wave (e.g., using the radio frequency portion of the spectrum) or a baseband signal. Formatting can include, for example, encoding a data stream and modulating a carrier wave with the encoded data stream. The information carried by the signal can be, for example, analog or digital information. As is known, the signal can be transmitted over a variety of different wired or wireless links. The signal can be stored on a processor-readable medium.
[0195] Various generalized and specialized embodiments are also supported and considered in the present disclosure.Examples of embodiments according to the present disclosure include, but are not limited to, the following.
[0196] Generally speaking, an example of an embodiment may relate to a method comprising: determining a decoding mode associated with a current decoding unit comprising picture information; determining, based on the decoding mode, a first precision level associated with a first portion of motion vector information and a second precision level associated with a second portion of the motion vector information, wherein the motion vector information is associated with the current decoding unit; obtaining a motion vector associated with the current decoding unit based on the motion vector information, and the first precision level and the second precision level; and decoding at least a portion of the picture information included in the current coding unit based on the decoding mode and the motion vector.
[0197] In general, another example of an embodiment may relate to a method comprising: determining a decoding mode associated with a current decoding unit comprising picture information; determining, based on the decoding mode, a first precision level associated with a first portion of motion vector information and a second precision level associated with a second portion of the motion vector information, wherein the motion vector information is associated with the current decoding unit; obtaining a motion vector associated with the current decoding unit based on the motion vector information and the first precision level and the second precision level; and encoding at least a portion of the picture information based on the decoding mode and the motion vector.
[0198] In general, another example of an embodiment may involve a device comprising one or more processors configured to determine a decoding mode associated with a current decoding unit comprising picture information; based on the decoding mode, determine a first precision level associated with a first portion of motion vector information and a second precision level associated with a second portion of the motion vector information, wherein the motion vector information is associated with the current decoding unit; based on the motion vector information, and the first precision level and the second precision level, obtain a motion vector associated with the current decoding unit; and decode at least a portion of the picture information based on the decoding mode and the motion vector.
[0199] In general, another example of an embodiment may involve a device comprising one or more processors configured to determine a decoding mode associated with a current decoding unit comprising picture information; based on the decoding mode, determine a first precision level associated with a first portion of motion vector information and a second precision level associated with a second portion of the motion vector information, wherein the motion vector information is associated with the current decoding unit; based on the motion vector information, and the first precision level and the second precision level, obtain a motion vector associated with the decoding unit of picture information; and encode at least a portion of the picture information based on the decoding mode and the motion vector.
[0200] In general, another example of an embodiment may relate to a method comprising: determining a decoding mode associated with a current decoding unit including picture information; determining, based on the decoding mode, a first precision level associated with a first component of a motion vector residual and a second precision level associated with a second component of the motion vector residual, the motion vector residual being associated with the current decoding unit; obtaining a motion vector associated with the current decoding unit based on the motion vector residual and the first precision level and the second precision level; and decoding at least a portion of the picture information included in the current decoding unit based on the decoding mode and the motion vector.
[0201] In general, another example of an embodiment may relate to a method comprising: determining a decoding mode associated with a current decoding unit comprising picture information; determining, based on the decoding mode, a first precision level associated with a first component of a motion vector residual and a second precision level associated with a second component of the motion vector residual, the motion vector residual being associated with the current decoding unit; obtaining a motion vector associated with the current decoding unit based on the motion vector residual and the first precision level and the second precision level; and encoding at least a portion of the picture information based on the decoding mode and the motion vector.
[0202] In general, another example of an embodiment may involve a device comprising one or more processors configured to determine a decoding mode associated with a current decoding unit comprising picture information; determine, based on the decoding mode, a first precision level associated with a first component of a motion vector residual and a second precision level associated with a second component of the motion vector residual, the motion vector residual being associated with the current decoding unit; obtain a motion vector associated with the current decoding unit based on the motion vector residual and the first precision level and the second precision level; and decode at least a portion of the picture information based on the decoding mode and the motion vector.
[0203] In general, another example of an embodiment may involve a device comprising one or more processors configured to determine a decoding mode associated with a current decoding unit comprising picture information; determine, based on the decoding mode, a first precision level associated with a first component of a motion vector residual and a second precision level associated with a second component of the motion vector residual, the motion vector residual being associated with the current decoding unit; obtain a motion vector associated with the decoding unit of picture information based on the motion vector residual and the first precision level and the second precision level; and encode at least a portion of the picture information based on the decoding mode and the motion vector.
[0204] In general, another example of an embodiment may relate to a method comprising: determining a decoding mode associated with a current decoding unit comprising picture information; determining, based on the decoding mode, a first precision level associated with a first motion vector residual corresponding to a first reference picture list, and a second precision level associated with a second motion vector residual corresponding to a second reference picture list; obtaining a first motion vector based on the first motion vector residual and the first precision level, and obtaining a second motion vector based on the second motion vector residual and the second precision level; and decoding at least a portion of the picture information based on the decoding mode, and the first motion vector and the second motion vector.
[0205] In general, another example of an embodiment may relate to a method comprising: determining a decoding mode associated with a current decoding unit comprising picture information; determining, based on the decoding mode, a first precision level associated with a first motion vector residual corresponding to a first reference picture list, and a second precision level associated with a second motion vector residual corresponding to a second reference picture list; obtaining a first motion vector based on the first motion vector residual and the first precision level, and obtaining a second motion vector based on the second motion vector residual and the second precision level; and encoding at least a portion of the picture information based on the decoding mode, and the first motion vector and the second motion vector.
[0206] In general, another example of an embodiment may involve a device comprising one or more processors configured to determine a decoding mode associated with a current decoding unit comprising picture information; determine, based on the decoding mode, a first precision level associated with a first motion vector residual corresponding to a first reference picture list, and a second precision level associated with a second motion vector residual corresponding to a second reference picture list; obtain a first motion vector based on the first motion vector residual and the first precision level, and obtain a second motion vector based on the second motion vector residual and the second precision level; and decode at least a portion of the picture information based on the decoding mode, and the first motion vector and the second motion vector.
[0207] In general, another example of an embodiment may involve a device comprising one or more processors configured to determine a decoding mode associated with a current decoding unit comprising picture information; determine, based on the decoding mode, a first precision level associated with a first motion vector residual corresponding to a first reference picture list, and a second precision level associated with a second motion vector residual corresponding to a second reference picture list; obtain a first motion vector based on the first motion vector residual and the first precision level, and obtain a second motion vector based on the second motion vector residual and the second precision level; and encode at least a portion of the picture information based on the decoding mode, and the first motion vector and the second motion vector.
[0208] Generally speaking, another example of an embodiment may involve a method or apparatus according to one or more examples of the embodiments described herein, wherein the first component corresponds to the x-component of the motion vector residual; and the second component corresponds to the y-component of the motion vector residual.
[0209] In general, another example of an embodiment may involve a method or apparatus according to one or more examples of the embodiments described herein, wherein the first level of precision represents a first level of accuracy of the first motion vector residual; and the second level of precision represents a second level of accuracy of the second motion vector residual.
[0210] Generally speaking, another example of an embodiment may involve a method or apparatus according to one or more examples of the embodiments described herein, wherein the first level of accuracy is different from the second level of accuracy.
[0211] In general, another example of an embodiment may relate to a method or apparatus according to one or more examples of the embodiments described herein, wherein the coding mode comprises an intra block copy mode.
[0212] In general, another example of an embodiment may relate to a method or apparatus according to one or more examples of the embodiments described herein, wherein the coding mode comprises a translation adaptive motion vector prediction mode.
[0213] In general, another instance of an embodiment may be directed to a method or apparatus according to one or more examples of the embodiments described herein, wherein the coding mode comprises an affine adaptive motion vector prediction mode.
[0214] In general, another example of an embodiment may involve a bitstream formatted to contain encoded picture information, wherein the encoded picture information is encoded by processing the picture information based on any one or more of the examples of embodiments of the method according to the present invention.
[0215] In general, one or more other examples of embodiments may also provide a computer-readable storage medium, such as a non-volatile computer-readable storage medium, having stored thereon instructions for encoding or decoding picture information such as video data according to the methods or apparatus described herein.
[0216] In general, at least one example of the embodiments may be directed to a computer program product comprising instructions that, when executed by a computer, cause the computer to perform a method according to one or more examples of the embodiments described herein.
[0217] Generally speaking, at least one example of the embodiments may involve a non-transitory computer-readable medium storing executable program instructions that cause a computer executing the instructions to perform a method according to one or more examples of the embodiments described herein.
[0218] In general, at least one example of the embodiments may involve a signal including data generated according to any one or more examples of the embodiments described herein.
[0219] In general, at least one example of an embodiment may involve a bitstream formatted to include syntax elements and coded image information generated according to any one or more of the examples of the embodiments described herein.
[0220] In general, at least one example of the embodiments may involve a computer-readable storage medium having stored thereon a bitstream generated according to a method or apparatus described herein.
[0221] In general, at least one example of an embodiment may involve transmitting or receiving a bitstream or signal generated according to a method or apparatus described herein.
[0222] In general, at least one example of an embodiment may relate to a device comprising an apparatus according to any one or more of the examples of the embodiments described herein; and (i) an antenna configured to receive a signal, the signal including data representing the image information, (ii) a frequency band limiter configured to limit the received signal to a frequency band including the data representing the image information, and (iii) a display configured to display an image based on the image information.
[0223] In general, at least one example of an embodiment may be directed to a device as described herein, wherein the device comprises one of: a television, a television signal receiver, a set-top box, a gateway device, a mobile device, a cellular phone, a tablet computer, or other electronic device.
[0224] Various embodiments have been described. Embodiments may include any of the following features or entities, alone or in any combination, across various claim categories and types:
[0225] Providing a level of accuracy for coded motion vector differences (or MV residuals) in an encoder and / or decoder, said level of accuracy being differentiated between the x-component and the y-component of said decoded motion vector differences;
[0226] In the case of bi-predicted blocks, provide a level of accuracy for coded motion vector differences (or MV residuals) in the encoder and / or decoder that differentiates between the L0 and L1 reference picture lists, where symmetric motion vector difference coding is not used;
[0227] Providing a level of accuracy of the decoded motion vector differences (or MV residuals) in the encoder and / or decoder, said level of accuracy being differentiated between the x-component and the y-component of the decoded motion vector differences, where the considered block is coded by an intra block copy mode such as VVC;
[0228] Providing a level of accuracy of decoded motion vector differences (or MV residuals) in an encoder and / or decoder, the level of accuracy being differentiated between the x-component and the y-component of the decoded motion vector differences, where the considered block is coded by an inter-translation AMVP mode such as VVC;
[0229] Providing a level of accuracy of the decoded motion vector differences (or MV residuals) in the encoder and / or decoder, the level of accuracy being differentiated between the x-component and the y-component of the decoded motion vector differences, the considered block being coded by the inter-affine AMVP mode of VVC;
[0230] providing, in the encoder and / or decoder, a level of accuracy of the coded motion vector differences (or MV residuals) that is differentiated between the L0 and L1 reference picture lists in case of bi-predicted blocks, where symmetric motion vector difference coding is not used and where the considered block is coded with an inter-translational AMVP mode such as VVC;
[0231] In case of bi-predicted blocks, provide in the encoder and / or decoder a level of accuracy of the coded motion vector differences (or MV residuals) that is differentiated between the L0 and L1 reference picture lists, where symmetric motion vector difference coding is not used, and where the considered block is coded by an inter-affine AMVP mode such as VVC.
[0232] Providing increased flexibility in AMVR tools such as VVC in encoders and / or decoders;
[0233] Providing a level of precision in the encoder and / or decoder for distinguishing between x-values and y-values of decoded motion vector differences (MVd);
[0234] • In case of bi-predictive coding units, provide a level of precision in the encoder and / or decoder for distinguishing between L0 and L1 motion vector difference coding.
[0235] Providing increased flexibility in an AMVR tool (e.g., an AMVR tool for VVC) in accordance with any embodiment, feature, or entity in an encoder and / or decoder, alone or in any combination, as described herein, further based on providing reduced complexity and / or improved compression efficiency.
[0236] • inserting syntax elements in the signaling, the syntax elements enabling the encoder and / or decoder to provide encoding and / or decoding according to any of the embodiments, features or entities, alone or in any combination, as described herein.
[0237] • Based on these syntax elements, the features or entities are selected, alone or in any combination, as described herein to be applied at the decoder.
[0238] • A bitstream or signal comprising one or more of the described syntax elements, or variations thereof.
[0239] • Inserting syntax elements in the signaling that enable the decoder to provide decoding in a manner corresponding to the encoding manner used by the encoder.
[0240] • Creating and / or sending and / or receiving and / or decoding a bitstream or signal comprising one or more of the described syntax elements or variations thereof.
[0241] • A TV, set-top box, cellular phone, tablet or other electronic device providing for applying encoding and / or decoding according to any of the embodiments, features or entities, alone or in any combination, as described herein.
[0242] A TV, set-top box, cellular phone, tablet, or other electronic device that performs encoding and / or decoding according to any embodiment, feature, or entity, alone or in any combination, as described herein, and displays (e.g., using a monitor, screen, or other type of display) the resulting image.
[0243] A TV, set-top box, cellular phone, tablet, or other electronic device that tunes (e.g., using a tuner) to a channel to receive a signal including an encoded image and performs encoding and / or decoding according to any of the embodiments, features, or entities as described herein, alone or in any combination.
[0244] A TV, set-top box, cellular phone, tablet, or other electronic device that receives over the air (e.g., using an antenna) a signal comprising an encoded image and performs encoding and / or decoding according to any of the embodiments, features, or entities, alone or in any combination, as described herein.
[0245] • A computer program product storing program codes which, when encoded and / or decoded by a computer according to any embodiment, feature or entity, are executed alone or in any combination as described herein.
[0246] A non-transitory computer-readable medium comprising executable program instructions that cause a computer to execute said instructions to implement encoding and / or decoding according to any of the embodiments, features or entities, alone or in any combination, as described herein.
[0247] Various other generalized and specialized embodiments are also supported and contemplated in this disclosure.
[0248] Syntax structure list 1: Example of coding_unit syntax structure according to embodiment 1.
[0249]
[0250]
[0251]
[0252]
[0253] Syntax structure list 2: Example of coding_unit syntax structure according to embodiment 2.
[0254]
[0255]
[0256]
[0257]
[0258] Syntax structure list 3: Example of coding_unit syntax structure according to embodiment 3.
[0259]
[0260]
[0261]
[0262]
[0263] Syntax structure list 4: Example of coding_unit syntax structure according to embodiment 4.
[0264]
[0265]
[0266]
[0267]
[0268] Syntax structure list 5: Example of coding_unit syntax structure according to embodiment 5.
[0269]
[0270]
[0271]
[0272]
Claims
1. A method for video decoding, comprising: determining a decoding mode associated with a current block, wherein the current block includes picture information; Resolving a motion vector difference associated with the current block; determining, based on the decoding mode, a first level of precision associated with a first component of the motion vector difference and a second level of precision associated with a second component of the motion vector difference; converting the first component of the motion vector difference from the first level of precision to a third level of precision used by a decoder to represent motion data, and converting the second component of the motion vector difference from the second level of precision to the third level of precision; obtaining a motion vector associated with the current block based on the motion vector difference and a motion vector prediction value; as well as The current block is reconstructed based on the decoding mode and the motion vector.
2. The method according to claim 1, comprising: A first syntax element associated with the first component and a second syntax element associated with the second component are parsed, and the first level of precision is determined using the first syntax element and the second level of precision is determined using the second syntax element.
3. The method according to claim 1 or 2, wherein The first component corresponds to the x-component of the motion vector difference; The second component corresponds to the y component of the motion vector difference; The first level of precision represents a first level of accuracy of the x-component; and The second level of precision represents a second level of accuracy of the y-component.
4. The method according to claim 3, wherein: The first level of accuracy is different from the second level of accuracy.
5. The method of claim 1 or 2, wherein the coding mode comprises an intra block copy mode.
6. The method of claim 1 or 2, wherein the coding mode comprises a translation adaptive motion vector prediction mode.
7. The method of claim 1 or 2, wherein the decoding mode comprises an affine adaptive motion vector prediction mode.
8. A method for video encoding, comprising: determining a decoding mode associated with a current block; Obtaining and encoding a motion vector difference associated with the current block to represent motion data; determining, based on the decoding mode, a first level of precision associated with a first component of the motion vector difference and a second level of precision associated with a second component of the motion vector difference; converting the first component of the motion vector difference from the first level of precision to a third level of precision used by an encoder to represent motion data, and converting the second component of the motion vector difference from the second level of precision to the third level of precision; Obtaining a motion vector associated with the current block based on the motion vector difference and a motion vector prediction value; as well as The current block is encoded based on the coding mode and the motion vector.
9. The method according to claim 8, comprising: A first syntax element associated with the first component and a second syntax element associated with the second component are encoded, the first syntax element being used to determine the first level of precision and the second syntax element being used to determine the second level of precision.
10. The method according to claim 8 or 9, wherein The first component corresponds to the x-component of the motion vector difference; The second component corresponds to the y component of the motion vector difference; The first level of precision represents a first level of accuracy of the x-component; and The second level of precision represents a second level of accuracy of the y-component.
11. The method according to claim 10, wherein: The first level of accuracy is different from the second level of accuracy.
12. The method of claim 8 or 9, wherein the coding mode comprises an intra block copy mode.
13. The method of claim 8 or 9, wherein the coding mode comprises a translation adaptive motion vector prediction mode.
14. The method of claim 8 or 9, wherein the coding mode comprises an affine adaptive motion vector prediction mode.
15. A device for video decoding, comprising: One or more processors configured to perform: determining a decoding mode associated with a current block, wherein the current block includes picture information; Resolving a motion vector difference associated with the current block; determining, based on the decoding mode, a first level of precision associated with a first component of the motion vector difference and a second level of precision associated with a second component of the motion vector difference; converting the first component of the motion vector difference from the first level of precision to a third level of precision used by a decoder to represent motion data, and converting the second component of the motion vector difference from the second level of precision to the third level of precision; Obtaining a motion vector associated with the current block based on the motion vector difference and a motion vector prediction value; as well as The current block is reconstructed based on the decoding mode and the motion vector.
16. The apparatus of claim 15, wherein the processor is configured to perform: parsing a first syntax element associated with the first component and a second syntax element associated with the second component, and determining the first level of precision using the first syntax element, and determining the second level of precision using the second syntax element.
17. The device according to claim 15 or 16, wherein The first component corresponds to the x-component of the motion vector difference; The second component corresponds to the y component of the motion vector difference; The first level of precision represents a first level of accuracy of the x-component; and The second level of precision represents a second level of accuracy of the y-component.
18. The device according to claim 17, wherein The first level of accuracy is different from the second level of accuracy.
19. The apparatus of claim 15 or 16, wherein the coding mode comprises an intra block copy mode.
20. The apparatus of claim 15 or 16, wherein the coding mode comprises a translation adaptive motion vector prediction mode.
21. The apparatus of claim 15 or 16, wherein the coding mode comprises an affine adaptive motion vector prediction mode.
22. An apparatus for video encoding, comprising: One or more processors configured to perform: determining a decoding mode associated with a current block; Obtaining and encoding a motion vector difference associated with the current block to represent motion data; determining, based on the decoding mode, a first level of precision associated with a first component of the motion vector difference and a second level of precision associated with a second component of the motion vector difference; converting the first component of the motion vector difference from the first level of precision to a third level of precision used by an encoder to represent motion data, and converting the second component of the motion vector difference from the second level of precision to the third level of precision; Obtaining a motion vector associated with the current block based on the motion vector difference and a motion vector prediction value; as well as The current block is encoded based on the coding mode and the motion vector.
23. The apparatus of claim 22, wherein the processor is configured to perform encoding a first syntax element associated with the first component and a second syntax element associated with the second component, the first syntax element being used to determine the first level of precision and the second syntax element being used to determine the second level of precision.
24. The device according to claim 22 or 23, wherein The first component corresponds to the x-component of the motion vector difference; The second component corresponds to the y component of the motion vector difference; The first level of precision represents a first level of accuracy of the x-component; and The second level of precision represents a second level of accuracy of the y-component.
25. The apparatus according to claim 24, wherein The first level of accuracy is different from the second level of accuracy.
26. The apparatus of claim 22 or 23, wherein the coding mode comprises an intra block copy mode.
27. The device of claim 22 or 23, wherein the coding mode comprises a translation adaptive motion vector prediction mode.
28. The apparatus of claim 22 or 23, wherein the coding mode comprises an affine adaptive motion vector prediction mode.
29. A computer program product comprising instructions which, when executed by a computer, cause the computer to perform the method according to any one of claims 1 to 14.
30. A non-transitory computer-readable medium storing executable program instructions, wherein the executable program instructions cause a computer to execute the method according to any one of claims 1 to 14.
Citation Information
Patent Citations
Moving picture signal coding method, decoding method, coding apparatus, and decoding apparatus
EP1469682A1