Systems and methods for deriving motion vector predictors in video coding
Patent Information
- Application Number
- CN202311654563.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-03-07
- Filing Date
- 2019-11-14
- Publication Date
- 2026-08-18
- Estimated Expiration
- 2039-11-14
Smart Images

Figure CN117692657B_ABST
Abstract
Description
[0001] This application is a divisional application of the application filed on November 14, 2019, with international application number PCT / JP2019 / 044730, Chinese application number "201980074290.5", and entitled "System and method for deriving motion vector prediction in video coding". Technical Field
[0002] This disclosure relates to video coding, and more specifically to techniques for deriving motion vector predictions. Background Technology
[0003] Digital video capabilities can be integrated into a wide variety of devices, including digital televisions, laptops or desktops, tablets, digital recording devices, digital media players, video game consoles, cellular phones (including so-called smartphones), medical imaging equipment, and more. Digital video can be encoded according to video coding standards. Video coding standards can incorporate video compression techniques. Examples of video coding standards include ISO / IEC MPEG-4 Visual and ITU-T H.264 (also known as ISO / IEC MPEG-4 AVC) and High Efficiency Video Coding (HEVC). HEVC is described in ITU-T Recommendation H.265, dated December 2016, and is incorporated herein by reference as ITU-T H.265. Extensions and improvements to ITU-T H.265 are currently under consideration for developing next-generation video coding standards. For example, the ITU-T Video Coding Experts Group (VCEG) and the ISO / IEC Moving Picture Experts Group (MPEG) (collectively referred to as the Joint Video Study Group (JVET)) are investigating the potential need for standardization of future video coding technologies with compression capabilities significantly exceeding the current HEVC standard. The Joint Exploratory Model 7 (JEM 7), the algorithmic description of Joint Exploratory Test Model 7 (JEM 7), and the ISO / IEC JTC1 / SC29 / WG11 document JVET-G1001 (July 2017, Turin, Italy), which are incorporated herein by reference, describe the coding features of JVET under the Joint Test Model study. This technique is a potential enhanced video coding technique that surpasses the capabilities of ITU-T H.265. It should be noted that the coding features of JEM 7 are implemented in the JEM reference software. As used herein, the term JEM can be used collectively to refer to the algorithms included in JEM 7 and the specific implementations in the JEM reference software. Furthermore, in response to the “Joint Call for Proposals on Video Compression with Capabilities beyond HEVC” jointly published by VCEG and MPEG, various groups presented multiple descriptions of video coding at the 10th meeting of ISO / IEC JTC1 / SC29 / WG11 held in San Diego, California, from April 16 to 20, 2018. As a result of various descriptions of video coding, the draft text of the video coding specification is described in “Versatile Video Coding (Draft 1)”, namely document JVET-J1001-v2, at the 10th meeting of ISO / IEC JTC1 / SC29 / WG11 held in San Diego, California, from April 16 to 20, 2018. This document is incorporated herein by reference and is referred to as JVET-J1001.The “Versatile Video Coding (Draft 2)” (document JVET-K1001-v7), presented at the 11th meeting of ISO / IEC JTC1 / SC29 / WG11 held in Ljubljana, Slovenia, from July 10 to 18, 2018, is an update to JVET-J1001. This document is incorporated herein by reference and is referred to as JVET-K1001. Furthermore, the “Versatile Video Coding (Draft 3)” (document JVET-L1001-v2), presented at the 12th meeting of ISO / IEC JTC1 / SC29 / WG11 held in Macau, China, from October 3 to 12, 2018, is an update to JVET-K1001. This document is incorporated herein by reference and is referred to as JVET-L1001.
[0004] Video compression technology reduces the data requirements for storing and transmitting video data. It reduces data requirements by utilizing the inherent redundancy in video sequences. Video compression can further divide a video sequence into smaller, continuous segments (e.g., frame groups within a video sequence, frames within a frame group, regions within a frame, video blocks within a region, and sub-blocks within a video block). Intra-frame predictive coding techniques (e.g., intra-picture (spatial)) and inter-frame predictive coding techniques (i.e., inter-picture (temporal)) can be used to generate the difference between the video data unit to be encoded and a reference unit of the video data. This difference can be called residual data. The residual data can be encoded as quantized transform coefficients. Syntax elements can relate to the residual data and reference coding units (e.g., intra-frame predictive mode indexes, motion vectors, and block vectors). Entropy coding can be applied to the residual data and syntax elements. The entropy-coded residual data and syntax elements can be included in the compliant bitstream. The compliant bitstream and associated metadata can be formatted according to the data structure. Summary of the Invention
[0005] In one example, a method for performing motion vector prediction for encoding video data is disclosed, the method comprising: determining a full-precision motion vector mv for generating predictions for video blocks in a first image; storing a rounded motion vector rmv having a lower precision than the full-precision motion vector mv; and generating candidate predicted motion vectors for video blocks in a second image from the stored motion vectors. Attached Figure Description
[0006] [ Figure 1 ] Figure 1 This is a conceptual diagram illustrating an example of a set of pictures encoded according to quadtree / multitree partitioning based on one or more techniques of this disclosure.
[0007] [ Figure 2A ] Figure 2AThis is a conceptual diagram illustrating an example of encoding video data blocks according to one or more techniques disclosed herein.
[0008] [ Figure 2B ] Figure 2B This is a conceptual diagram illustrating an example of encoding video data blocks according to one or more techniques disclosed herein.
[0009] [ Figure 3 ] Figure 3 This is a conceptual diagram illustrating the location of adjacent video blocks included in a set of candidates for predicting motion vectors, according to one or more techniques disclosed herein.
[0010] [ Figure 4 ] Figure 4 This is a conceptual diagram illustrating one or more techniques according to the present disclosure for containing positionally adjacent video blocks in a set of candidate predicted motion vectors.
[0011] [ Figure 5 ] Figure 5 This is a block diagram illustrating an example of a system that can be configured to encode and decode video data according to one or more techniques of this disclosure.
[0012] [ Figure 6 ] Figure 6 This is a block diagram illustrating an example of a video encoder that can be configured to encode video data according to one or more techniques of this disclosure.
[0013] [ Figure 7 ] Figure 7 This is a block diagram illustrating an example of a video decoder that can be configured to decode video data according to one or more techniques of this disclosure. Detailed Implementation
[0014] Generally, this disclosure describes various techniques for encoding video data. Specifically, this disclosure describes techniques for motion vector prediction in video coding. More specifically, this disclosure describes techniques for varying the storage precision of motion information used to generate motion vector predictions. Varying the precision of motion information according to the techniques described herein can be particularly useful for optimizing video coding performance and the memory cost of motion vector prediction. It should be noted that although the techniques of this disclosure are described relative to ITU-T H.264, ITU-T H.265, JVET-J1001, JVET-K1001, and JVET-L1001, the techniques of this disclosure are generally applicable to video coding. For example, the coding techniques described herein can be incorporated into video coding systems (including video coding systems based on future video coding standards), and these techniques include block structure, intra-frame prediction techniques, inter-frame prediction techniques, transform techniques, filtering techniques, and / or entropy coding techniques, different from those included in ITU-T H.265. Therefore, references to ITU-T H.264, ITU-T H.265, JVET-J1001, JVET-K1001, and JVET-L1001 are for descriptive purposes and should not be construed as limiting the scope of the techniques described herein. Furthermore, it should be noted that the inclusion of references herein by way of citation should not be construed as limiting or creating ambiguity with respect to the terminology used herein. For example, where a definition of a term provided in one of the incorporated references differs from that in another incorporated reference and / or as used herein, the term should be interpreted in a manner that broadly includes each corresponding definition and / or in a manner that includes each specific definition in alternatives.
[0015] In one example, an apparatus for reconstructing video data includes one or more processors configured to: determine full-precision motion vectors for generating predictions of video patches in a first image, store the motion vectors at less than the full precision, and generate candidate predicted motion vectors from the stored motion vectors for video patches in a second image.
[0016] In one example, a non-transitory computer-readable storage medium includes instructions stored thereon that, when executed, cause one or more processors of the device to: determine a full-precision motion vector for generating predictions of video patches in a first image, store the motion vector at less than the full precision, and generate candidate predicted motion vectors from the stored motion vectors for video patches in a second image.
[0017] In one example, an apparatus includes: components for determining full-precision motion vectors for generating predictions of video blocks in a first image; components for storing the motion vectors with less than the full precision; and components for generating candidate predicted motion vectors for video blocks in a second image from the stored motion vectors.
[0018] Details of one or more examples are set forth in the following figures and description. Other features, objects, and advantages will be apparent from the description, figures, and claims.
[0019] Video content typically comprises a sequence of frames (or pictures). A series of frames may also be referred to as a group of pictures (GOP). Each video frame or picture may be divided into one or more regions. A region may be defined based on a basic unit (e.g., a video block) and a set of rules defining the region (e.g., a region must be an integer number of video blocks arranged in a rectangle). As used herein, the term "video block" may generally refer to a region of a picture, or more specifically, to the largest array of sample values that can be predictably encoded, its sub-partitions, and / or corresponding structures. Furthermore, the term "current video block" may refer to the region of a picture that is being encoded or decoded. A video block may be defined as an array of sample values that can be predictably encoded. It should be noted that in some cases, pixel values may be described as sample values comprising the corresponding components of the video data, which may also be referred to as color components (e.g., luminance (Y) and chrominance (Cb and Cr) components or red, green, and blue components). It should be noted that in some cases, the terms "pixel value" and "sample value" are used interchangeably. Furthermore, in some cases, a pixel or sample may be referred to as a pel. A video sampling format (also known as a chroma format) can be defined relative to the number of chroma samples included in a video block. For example, in a 4:2:0 sampling format, the sampling rate of the luminance component is twice the sampling rate of the chroma components in both the horizontal and vertical directions. Therefore, for a video block formatted according to the 4:2:0 format, the width and height of the sample array used for the luminance component are twice the width and height of each sample array used for the chroma component. For a video block formatted according to the 4:2:2 format, the width of the sample array for the luminance component is twice the width of the sample array for each chroma component, but the height of the sample array for the luminance component is equal to the height of the sample array for each chroma component. Furthermore, for a video block formatted according to the 4:4:4 format, the sample array for the luminance component has the same width and height as the sample array for each chroma component.
[0020] Video blocks can be ordered within a picture and / or region according to a scanning mode (e.g., raster scan). A video encoder can perform predictive coding on video blocks and their sub-regions. Video blocks and their sub-regions can be referred to as nodes. ITU-T H.264 specifies macroblocks comprising 16×16 luma samples. That is, in ITU-T H.264, pictures are segmented into macroblocks. ITU-T H.265 specifies a similar Code Tree Unit (CTU) structure (also known as Maximum Code Unit (LCU)). In ITU-T H.265, pictures are segmented into CTUs. In ITU-T H.265, for pictures, the CTU size can be set to include 16×16, 32×32, or 64×64 luma samples. In ITU-T H.265, a CTU consists of a corresponding Code Tree Block (CTB) for each component of the video data (e.g., luma (Y) and chroma (Cb and Cr)). Furthermore, in ITU-T H.265, CTUs can be partitioned according to a quadtree (QT) partitioning structure, which allows the CTU's CTB to be divided into coded blocks (CBs). That is, in ITU-T H.265, a CTU can be divided into quadtree leaf nodes. According to ITU-T H.265, a luma CB, along with two corresponding chroma CBs and associated syntax elements, is called a coding unit (CU). In ITU-T H.265, the minimum permissible size of a CB that can be transmitted by signal is specified. In ITU-T H.265, the minimum permissible size of a luma CB is 8×8 luma samples. In other words, in ITU-T H.265, the decision to code a picture region using intra-frame prediction or inter-frame prediction is made at the CU level.
[0021] In ITU-T H.265, a CU (Cubic Component Unit) is associated with a Prediction Unit (PU) structure that has its root at the CU. In ITU-T H.265, the PU structure allows the segmentation of the Luminance CB (Cubic Block) and Chromaticity CB to generate corresponding reference samples. That is, in ITU-T H.265, the Luminance CB and Chromaticity CB can be segmented into corresponding Luminance Prediction Blocks (PBs) and Chromaticity Prediction Blocks (PBs), where each PB comprises a block of sample values to which the same prediction is applied. In ITU-T H.265, a CB can be divided into one, two, or four PBs. ITU-T H.265 supports PB sizes from 64×64 samples down to 4×4 samples. ITU-T H.265 supports square PBs for intra-frame prediction, where a CB can form a PB, or a CB can be segmented into four square PBs (i.e., intra-frame prediction PB types include M×M or M / 2×M / 2, where M is the height and width of the square CB). In ITU-T H.265, in addition to square prediction blocks (PBs), rectangular PBs are also supported for inter-frame prediction, where the frame cutoff block (CB) can be halved vertically or horizontally to form the PB (i.e., inter-frame prediction PB types include M×M, M / 2×M / 2, M / 2×M, or M×M / 2). Furthermore, it should be noted that ITU-T H.265 supports four asymmetric PB partitions for inter-frame prediction, where the CB is divided into two PBs at one-quarter of its height (top or bottom) or width (left or right) (i.e., asymmetric partitions include M / 4×M left, M / 4×M right, M×M / 4 top, and M×M / 4 bottom). Intra-frame prediction data (e.g., intra-frame prediction mode syntax elements) or inter-frame prediction data (e.g., motion data syntax elements) corresponding to the PB are used to generate reference and / or prediction sample values for the PB.
[0022] As described above, each video frame or image can be divided into one or more regions. For example, according to ITU-T H.265, each video frame or image can be divided into one or more slices, and further divided into one or more tiles, wherein each slice includes a CTU sequence (e.g., arranged in raster scan order), and wherein a tile is a CTU sequence corresponding to a rectangular area of the image. It should be noted that in ITU-T H.265, a slice is a sequence of one or more slice segments, starting from an independent slice fragment and preceding all subsequent subordinate slice fragments (if any) contained within the same access unit. A slice fragment (such as a slice) is a CTU sequence. Therefore, in some cases, the terms "slice" and "slice segment" are used interchangeably to indicate a CTU sequence. Furthermore, it should be noted that in ITU-T H.265, a tile may consist of CTUs contained in more than one slice, and a slice may consist of CTUs contained in more than one tile. However, ITU-T H.265 specifies that one or both of the following conditions must be met: (1) all CTUs in a segment belong to the same tile; and (2) all CTUs in a tile belong to the same slice. Compared to JVET-L1001, it has been proposed that slices must consist of an integer number of complete tiles, rather than just an integer number of complete CTUs. Therefore, some video coding techniques may or may not support slices containing a group of CTUs that do not form a picture. Furthermore, slices that require an integer number of complete tiles are called tile groups. The techniques described herein are applicable to slices, tiles, and / or tile groups. Figure 1 This is a concept diagram showing an example of a set of images including groups of tiles. Figure 1 In the example shown, Pic4 is depicted as comprising two tile groups (i.e., tile group 1 and tile group 2). It should be noted that in some cases, tile group 1 and tile group 2 may be classified as slices and / or tiles.
[0023] JEM specifies a maximum size CTU with 256×256 luminance samples. JEM specifies a Quadtree Plus Binary Tree (QTBT) block structure. In JEM, the QTBT structure allows the quadtree leaf nodes to be further divided by a binary tree (BT) structure. That is, in JEM, the binary tree structure allows for recursive vertical or horizontal partitioning of quadtree leaf nodes. In JVET-L1001, CTUs are partitioned according to a Quadtree Plus Multi-Type Tree (QTMT) structure. The QTMT in JVET-L1001 is similar to the QTBT in JEM. However, in JVET-L1001, in addition to indicating binary partitioning, the multi-type tree can also indicate so-called ternary (or ternary tree (TT)) partitioning. Ternary partitioning divides a block vertically or horizontally into three blocks. In the case of a vertical TT split, the block is divided at one-quarter of its width from the left edge and at one-quarter of its width from the right edge; and in the case of a horizontal TT split, the block is divided at one-quarter of its height from the top edge and at one-quarter of its height from the bottom edge. See again. Figure 1 , Figure 1 This illustrates an example where a CTU is partitioned into quadtree leaf nodes, and these quadtree leaf nodes are further partitioned based on either BT or TT partitioning. That is, in Figure 1 In the diagram, dashed lines indicate additional binary and ternary partitions in a quadtree.
[0024] As described above, intra-frame prediction data or inter-frame prediction data is used to generate reference sample values for the current video block. The difference between the sample values included in the prediction generated from the reference sample values and the current video block can be referred to as residual data. The residual data can include a corresponding array of differences for each component of the video data. The residual data may be in the pixel domain. Transformations such as Discrete Cosine Transform (DCT), Discrete Sine Transform (DST), integer transform, wavelet transform, or conceptually similar transforms can be applied to the array of differences to generate transform coefficients. It should be noted that in ITU-T H.265 and JVET-L1001, the CU is associated with a Transform Unit (TU) structure having its root at the CU level. That is, to generate transform coefficients, the array of differences can be partitioned (e.g., four 8×8 transforms can be applied to a 16×16 residual array). Such a subdivision of the differences for each component of the video data can be referred to as a Transform Block (TB). It should be noted that in some cases, a core transform and subsequent quadratic transforms can be applied (in the video encoder) to generate transform coefficients. For video decoders, the transformation order is reversed.
[0025] Quantization can be performed on transform coefficients. Quantization essentially scales the transform coefficients to change the amount of data needed to represent a set of transform coefficients. Quantization may include dividing the transform coefficients by a quantization scaling factor and any associated rounding function (e.g., rounding to the nearest integer). The quantized transform coefficients may be referred to as coefficient bit values. Inverse quantization (or “dequantization”) may include multiplying the coefficient bit values by the quantization scaling factor. It should be noted that, as used herein, the term quantization process may refer in some cases to dividing by a scaling factor to generate bit values, and in some cases to multiplying by a scaling factor to recover the transform coefficients. That is, quantization process may refer to quantization in some cases and inverse quantization in others.
[0026] Figures 2A to 2B This is a conceptual diagram illustrating an example of encoding video data blocks. (Example:) Figure 2A As shown, the current block of video data is encoded by generating bit-order values by subtracting a set of predicted values from the current video data block to generate a residual, performing a transformation on the residual, and quantizing the transformation coefficients. Figure 2B As shown, the current video data block is decoded by performing inverse quantization on the bit-order values, performing an inverse transform, and adding a set of predicted values to the resulting residual. It should be noted that, in Figures 2A to 2B In the example, the sample values of the reconstructed block differ from the sample values of the current video block being encoded. Thus, the encoding can be considered lossy. However, the difference in these values can be considered acceptable or imperceptible to the viewer of the reconstructed video. Additionally, as... Figures 2A to 2B As shown, scaling is performed using an array of scaling factors.
[0027] like Figure 2AAs shown, the quantized transform coefficients are encoded into a bitstream. The quantized transform coefficients and syntax elements (e.g., syntax elements indicating the coding structure of video blocks) can be entropy-coded using entropy coding techniques. Examples of entropy coding techniques include Content Adaptive Variable Length Coding (CAVLC), Context Adaptive Binary Arithmetic Coding (CABAC), Probability Interval Partition Entropy Coding (PIPE), etc. The entropy-coded quantized transform coefficients and the corresponding entropy-coded syntax elements can form a compliant bitstream that can be used to reproduce video data at the video decoder. The entropy coding process may include binarizing the syntax elements. Binarization is the process of converting the values of syntax values into a sequence of one or more bits. These bits may be referred to as "bins". Binarization is a lossless process and may include one or a combination of the following coding techniques: fixed-length coding, unary coding, truncated unary coding, truncated Rice coding, Golomb coding, k-order exponential Golomb coding, and Golomb-Rice coding. For example, binarization may include representing the integer value 5 of a syntax element as 00000101 using an 8-bit fixed-length binarization technique, or representing the integer value 5 as 11110 using a unary coding binarization technique. As used herein, each of the terms fixed-length coding, unary coding, truncated unary coding, truncated Rice coding, Golomb coding, k-order exponential Golomb coding, and Golomb-Rice coding may refer to a general implementation of these techniques and / or a more specific implementation of these coding techniques. For example, a Golomb-Rice coding implementation may be specifically defined according to a video coding standard (e.g., ITU-T H.265).
[0028] The entropy coding process also includes encoding bin values using a lossless data compression algorithm. In the CABAC examples, for a specific bin, a context model can be selected from a set of available context models associated with that bin. In some examples, the context model can be selected based on the values of previous bins and / or previous syntax elements. The context model can identify the probability that a bin has a specific value. For example, the context model might indicate a probability of 0.7 for encoding a bin with a value of 0. After selecting an available context model, the CABAC entropy encoder can perform arithmetic coding on the bins based on the identified context model. The context model can be updated based on the values of the encoded bins. The context model can also be updated based on associated variables stored with the context, such as window size adaptation and the number of encoded bins using the context. It should be noted that a CABAC entropy encoder can be implemented such that arithmetic coding can be used to entropy code some syntax elements without using an explicitly specified context model; this type of coding can be called bypass coding.
[0029] As mentioned above, intra-frame prediction data or inter-frame prediction data indicates how to generate predictions for the current video block. For intra-frame prediction coding, the intra-frame prediction mode can specify the location of reference samples within the image used to generate predictions. In ITU-T TH.265, the possible intra-frame prediction modes defined include planar (i.e., surface-fitting) prediction mode (predMode: 0), DC (i.e., flat global average) prediction mode (predMode: 1), and 33-corner prediction modes (predMode: 2-34). In JVET-L1001, the defined possible intra-frame prediction modes for luma include planar prediction mode (predMode: 0), DC prediction mode (predMode: 1), and 65-corner prediction modes (predMode: 2-66). It should be noted that the planar prediction mode and DC prediction mode can be referred to as non-directional prediction modes, and the corner prediction mode can be referred to as directional prediction modes. Furthermore, various methods may exist for deriving intra-frame prediction modes for chroma components based on the intra-frame prediction modes for luma components. It should be noted that the techniques described in this paper are generally applicable regardless of the number of possible prediction patterns that have been defined.
[0030] For inter-frame predictive coding, one or more previously decoded images are determined, i.e., reference images are identified, and motion vectors (MVs) identify samples in these reference images used to generate predictions for the current video block. For example, reference sample values located in one or more previously encoded images can be used to predict the current video block, and motion vectors are used to indicate the position of the reference block relative to the current video block. Motion vectors can describe, for example, the horizontal displacement component of the motion vector (i.e., MV). x ), the vertical displacement component of the motion vector (i.e., MV) yThe resolution of the motion vectors (e.g., quarter-pixel precision, half-pixel precision, one-pixel precision, two-pixel precision, four-pixel precision) is used. Previously decoded images (which may include images output before or after the current image) can be organized into one or more reference image lists and identified using reference image index values. Furthermore, in inter-frame predictive coding, single prediction refers to generating a prediction using sample values from a single reference image, while dual prediction refers to generating a prediction using corresponding sample values from two reference images. That is, in single prediction, a single reference image and its corresponding motion vector are used to generate a prediction for the current video block, while in dual prediction, a first reference image and its corresponding first motion vector, and a second reference image and its corresponding second motion vector are used to generate a prediction for the current video block. In dual prediction, the corresponding sample values are combined (e.g., added, rounded, and cropped, or averaged according to weights) to generate a prediction. Images and their regions can be classified based on which types of prediction patterns are available for encoding their video blocks. In other words, for regions of type B (e.g., B slices), dual prediction, single prediction, and intra-prediction modes can be used; for regions of type P (e.g., P slices), single prediction and intra-prediction modes can be used; and for regions of type I (e.g., I slices), only intra-prediction mode can be used. As described above, reference pictures are identified by reference indexes. In ITU-TH.265, for P slices, there is a single reference picture list RefPicList0; for B slices, in addition to RefPicList0, there is a second independent reference picture list RefPicList1. It should be noted that for single prediction in B slices, either RefPicList0 or RefPicList1 can be used to generate the prediction. Furthermore, it should be noted that in ITU-TH.265, during the decoding process, at the start of decoding a picture, one or more reference picture lists are generated from previously decoded pictures stored in the Decoding Picture Buffer (DPB).
[0031] Furthermore, the coding standard supports various motion vector prediction modes. Motion vector prediction enables the derivation of the value of a motion vector based on another motion vector. Examples of motion vector prediction include Advanced Motion Vector Prediction (AMVP), Temporal Motion Vector Prediction (TMVP), the so-called "merge" mode, and "skip" and "direct" motion inference. Other examples of motion vector prediction include Advanced Temporal Motion Vector Prediction (ATMVP) and Spatial-Temporal Motion Vector Prediction (STMVP). ITU-T H.265 supports two modes for motion vector prediction: the merge mode and the so-called Advanced Motion Vector Prediction (AMVP). In ITU-T H.265, for both the merge mode and AMVP of the current PB, a set of candidate blocks is derived. Both the video encoder and video decoder perform the same process to derive a set of candidates. Therefore, for the current video block, the same set of candidates is generated during encoding and decoding. The candidate blocks consist of video blocks with associated motion information from which the predicted motion information for generating the current video block can be derived. For the merge mode in ITU-T H.265, all motion information associated with the selected candidate (i.e., motion vector displacement values, reference picture indices, and reference picture lists) is inherited as the motion information of the current PB. That is, at the video encoder, candidate blocks are selected from the derived set of candidates, and the index values included in the bitstream indicate the selected candidate, and thus the motion information of the current PB. For AMVP in ITU-T H.265, the motion vector information for the selected candidate is used as the motion vector predictor (MVP) of the current PB's motion vector. That is, at the video encoder, candidate blocks are selected from the derived set of candidates, and the index values indicating the selected candidate and the incremental values indicating the difference between the predicted motion vector and the current PB's motion vector (i.e., motion vector increment (MVD)) are included in the bitstream. Furthermore, for AMVP in ITU-T H.265, the syntax elements identifying the reference pictures are included in the bitstream.
[0032] In ITU-T H.265, a set of candidate blocks can be derived from spatially adjacent blocks and temporally adjacent blocks. Furthermore, the generated (or default) motion information can be used for motion vector prediction. In ITU-T H.265, whether the motion information used for motion vector prediction of the current block PB includes motion information associated with spatially adjacent blocks, motion information associated with temporally adjacent blocks, or the generated motion information depends on the number of candidates to be included in the group, whether temporal motion vector prediction is enabled, block availability, and / or whether the motion information associated with the blocks is redundant.
[0033] For the merging mode in ITU-T H.265, the maximum number of candidates that can be included in a set of candidate blocks can be set by the video encoder and transmitted via signals, and can be up to five. Furthermore, the video encoder can disable the use of temporal motion vector candidates (e.g., to reduce the amount of memory resources required to store motion information at the video decoder), and the signal transmission indicates whether the use of temporal motion vector candidates is enabled or disabled for the picture. Figure 3 The locations of spatially adjacent blocks and temporal blocks in a set of candidate blocks that may be included in the merging mode in ITU-T H.265 are shown. The derivation of the set of candidates for the merging mode in ITU-T H.265 includes determining the availability of A1, B1, B0, A0, and B2. It should be noted that if a block is intra-frame predicted (i.e., does not have corresponding motion information) or is not included in the current segment (or tile), the block is considered unavailable. After determining the availability of A1, B1, B0, A0, and B2, a set of comparisons (such as...) are performed. Figure 3 (As shown by the dashed arrow in the image) to remove redundant entries from the group of candidates. For example, B2 is compared with B1, and if B1 has the same associated motion information as B2, it is removed from the group of candidates. Removing entries from a group of candidates can be called a pruning process. It should be noted that in Figure 3 In order to reduce complexity, a full comparison of candidates is not performed (e.g., A0 is not compared with B0), and therefore redundant entries may be included in the group of candidates.
[0034] Refer again Figure 3 The dashed block labeled "Time" refers to a temporal candidate that can be included in the candidate group. In ITU-T H.265 for merging modes, for a temporal candidate, a spatial juxtaposition unit (PU) included in the reference image is defined, and this temporal candidate includes a block located exactly outside the lower right corner of the juxtaposition unit (if available), or a block located at the center of the juxtaposition unit. As described above, the maximum number of candidates that can be included in a candidate block group can be set. If the maximum number of candidates is set to N, then if the number of available spatial (pruned) and temporal candidates is greater than or equal to N, N-1 spatial and temporal candidates are included in the group. If the number of available spatial (pruned) and temporal candidates is less than N, the generated motion information is included in the group to fill it.
[0035] For AMVP in ITU-T H265, refer to Figure 4The derivation of this candidate group involves adding one of A0 or A1 (i.e., the left candidate) and one of B0, B1, or B2 (the top candidate) to the group based on the availability of A0, A1, B0, B1, and B2. That is, the first available left candidate and the first available top candidate are added to the group. When the left and top candidates have redundant motion vector components, a redundant candidate is removed from the group. If the number of candidates included in the group is less than two, and temporal motion vector prediction is enabled, then temporal candidates (time) are included in the group. If the number of available spatial candidates (pruned) and temporal candidates included in the group is less than two, zero-valued motion vectors are included in the group to fill the group.
[0036] The following arithmetic operators can be used for the formulas used in this article:
[0037] Addition
[0038] - Subtraction
[0039] Multiplication, including matrix multiplication
[0040] x y Exponentiation. This indicates raising x to the power of y. In other contexts, this symbol is used as a superscript rather than intended to be interpreted as exponentiation.
[0041] / is an integer division operation that truncates the result towards zero. For example, it truncates 7 / 4 and -7 / -4 to 1, and -7 / 4 and 7 / -4 to 1.
[0042] In mathematical formulas, ÷ is used to represent division that is not intended to be truncated or rounded.
[0043] In mathematical formulas, x / y is used to represent division that is not intended to be truncated or rounded.
[0044] x%y modulo. The remainder when x is divided by y is defined only for integers x and y where x ≥ 0 and y > 0.
[0045] In addition, the following mathematical functions can be used:
[0046] Log2(x), the base-2 logarithm of x;
[0047]
[0048]
[0049] Ceil(x) is the smallest integer greater than or equal to x.
[0050] Floor(x), the largest integer less than or equal to x.
[0051]
[0052]
[0053]
[0054] In addition, the following logical operators can be used:
[0055] x&&y, the Boolean logical AND operation between x and y.
[0056] x||y, the Boolean logical OR operation between x and y.
[0057] ! represents the Boolean logic "NOT".
[0058] x? y : z, evaluate to y if x is TRUE or not equal to 0; otherwise, evaluate to z.
[0059] In addition, the following relational operators can be used:
[0060] >greater than
[0061] >= greater than or equal to
[0062] <less than
[0063] <= Less than or equal to
[0064] == equals
[0065] ! = not equal to
[0066] In addition, the following bitwise operators can be used:
[0067] & performs a bitwise AND operation. When operating on integer variables, it operates on the two's complement representation of the integer value. When operating on a binary variable that contains fewer bits than another variable, it expands the shorter variable by adding more significant bits equal to 0.
[0068] | Bitwise OR. When performing operations on integer variables, the operation is performed on the two's complement representation of the integer value. When performing operations on binary variables that contain fewer bits than another variable, the shorter variable is expanded by adding more significant bits equal to 0.
[0069] ^ Performs a bitwise XOR operation. When operating on integer variables, it operates on the two's complement representation of the integer value. When operating on a binary variable that contains fewer bits than another variable, it expands the shorter variable by adding more significant bits equal to 0.
[0070] x >> y is an arithmetic right shift of the binary representation of y in two's complement. This function is defined only for non-negative integer values of y. The bits shifted into the most significant bit (MSB) by the right shift have the same MSB value as x before the shift operation.
[0071] x << y is an arithmetic left shift of the binary representation of y in two's complement. This function is defined only for non-negative integer values of y. The bits shifted to the least significant bit (LSB) due to the left shift have a value equal to 0.
[0072] JVET-L1001 includes a merging mode based on the merging mode defined in ITU-T H.265 and an AMVP mode based on AMVP defined in ITU-T TH.256. It should be noted that JVET-L1001 also includes affine motion vector prediction techniques. As mentioned above, a motion vector may include its horizontal displacement component, its vertical displacement component, and its resolution. JVET-L1001 specifies the cases for deriving the luminance motion vector with 1 / 16 fractional sampling accuracy and the chrominance motion vector with 1 / 32 fractional sampling accuracy. Specifically, JVET-L1001 specifies that for luminance motion vector prediction, the luminance motion vector mvLX is derived as follows:
[0073] uLX[0]=(mvpLX[0]+mvdLX[0]+2 18 )%2 18
[0074] mvLX[0][0][0]=(uLX[0]>=2 17 )? (uLX[0]-2 18 ): uLX[0]
[0075] uLX[1]=(mvpLX[1]+mvdLX[1]+2 18 )%2 18
[0076] mvLX[0][0][1]=(uLX[1]>=2 17 )? (uLX[1]-2 18 ): uLX[1]
[0077] in,
[0078] Based on the corresponding motion prediction direction, X is replaced with 0 or 1;
[0079] mvpLX is the predicted motion vector; and
[0080] mvdLX is the motion vector increment.
[0081] It should be noted that, based on the above formula, the values of mvLX[0] (which indicates the direction (left or right) and magnitude of horizontal displacement) and mvLX[1] (which indicates the direction (up or down) and magnitude of vertical displacement) will always be in the range of -2.17 Up to 2 17 The range is -1 (including the end value).
[0082] JVET-L1001 specifies that, for affine motion vectors, the luma sub-block motion vector array is derived with a fractional sampling accuracy of 1 / 16, and the chroma sub-block motion vector array is derived with a fractional sampling accuracy of 1 / 32. Specifically, JVET-L1001 specifies that, for luma affine control point prediction motion vectors, the luma motion vector cpMvLX[cpIdx] with cpIdx in the range of 0 to NumCpMv⁻¹ is derived as follows:
[0083] uLX[cpIdx][0]=(mvpCpLX[cpIdx][0]+mvdCpLX[cpIdx][0]+2 18 )%2 18
[0084] cpMvLX[cpIdx][0]=(uLX[cpIdx][0]>=2 17 )? (uLX[cpIdx][0]-2 18 ): uLX[cpIdx][0]
[0085] uLX[cpIdx][1]=(mvpCpLX[cpIdx][1]+mvdCpLX[cpIdx][1]+2 18 )%2 18
[0086] cpMvLX[cpIdx][1]=(uLX[cpIdx][1]>=2 17 )? (uLX[cpIdx][1]-2 18 ): uLX[cpIdx][1]
[0087] in,
[0088] Based on the corresponding motion prediction direction, X is replaced with 0 or 1;
[0089] mvpCpLX is the predicted motion vector for the control points; and
[0090] mvdCpLX is the increment of the control point motion vector.
[0091] It should be noted that, based on the above formula, the values of cpmvLX[0] and cpmvLX[1] specified above will always be in the range of -2. 17 Up to 2 17 The range is -1 (including the end value).
[0092] As described above, motion vector candidates may include temporal candidates, which may include motion information associated with the juxtaposed blocks included in the reference image. Specifically, JVET-L1001 specifies the following for deriving juxtaposed motion vectors:
[0093] The input to this process is:
[0094] - The variable currCb specifies the current encoded block.
[0095] - The variable colCb specifies the juxtaposition encoding block within the juxtaposition image specified by ColPic.
[0096] -Luminance position (xColCb, yColCb), which specifies the top-left sample of the juxtaposed luminance coding block specified by colCb relative to the top-left luminance sample of the juxtaposed image specified by ColPic.
[0097] - Refer to the index refIdxLX, where X is 0 or 1.
[0098] - The flag sbFlag indicates the candidate for time merging of sub-blocks.
[0099] The output of this process is:
[0100] Predicted motion vector mvLXCol with -1 / 16 fractional sampling accuracy.
[0101] -Availability flag availableFlagLXCol.
[0102] The variable currPic specifies the current image.
[0103] Set the arrays predFlagL0Col[x][y], mvL0Col[x][y], and refIdxL0Col[x][y] to be equal to the PredFlagL0[x][y], MvL0[x][y], and RefIdxL0[x][y] of the juxtaposed images specified by ColPic, respectively. Set the arrays predFlagL1Col[x][y], mvL1Col[x][y], and refIdxL1Col[x][y] to be equal to the PredFlagL1[x][y], MvL1[x][y], and RefIdxL1[x][y] of the juxtaposed images specified by ColPic, respectively.
[0104] The variables mvLXCol and availableFlagLXCol are exported as follows:
[0105] - If colCb is encoded in intra-prediction mode, then set both components of mvLXCol to 0 and set availableFlagLXCol to 0.
[0106] - Otherwise, the motion vector mvCol, the reference index refIdxCol, and the reference list identifier listCol are derived as follows:
[0107] - If sbFlag equals 0, then set availableFlagLXCol to 1, and the following applies:
[0108] - If predFlagL0Col[xColCb][yColCb] equals 0, then set mvCol, refIdxCol, and listCol to equal mvL1Col[xColCb][yColCb], refIdxL1Col[xColCb][yColCb], and L1, respectively.
[0109] Otherwise, if predFlagL0Col[xColCb][yColCb] equals 1 and predFlagL1Col[xColCb][yColCb] equals 0, then set mvCol, refIdxCol, and listCol to equal mvL0Col[xColCb][yColCb], refIdxL0Col[xColCb][yColCb], and L0, respectively.
[0110] - Otherwise (predFlagL0Col[xColCb][yColCb] equals 1, and predFlagL1Col[xColCb][yColCb] equals 1), specify the following:
[0111] - If NoBackwardPredFlag equals 1, then set mvCol, refIdxCol, and listCol to equal mvLXCol[xColCb][yColCb], refIdxcLXCol[xColCb][yColCb], and LX, respectively.
[0112] Otherwise, set mvCol, refIdxCol, and listCol to be equal to mvLNCol[xColCb][yColCb], refIdxLNCol[xColCb][yColCb], and LN, respectively, where N is the value of collocated_from_10flag.
[0113] - Otherwise (sbFlag equals 1), the following applies:
[0114] - If PredFlagLXCol[xColCb][yColCb] equals 1, then set mvCol, refIdxCol, and listCol to equal mvLXCol[xColCb][yColCb], refIdxLXCol[xColCb][yColCb], and LX respectively, and set availableFlagLXCol to 1.
[0115] - Otherwise (PredFlagLXCol[xColCb][yColCb] equals 0), the following applies:
[0116] If for each image in the reference image list for the current slice, DiffPicOrderCnt(aPic, currPic) is less than or equal to 0, and PredFlagLYCol[xColCb][yColCb] is equal to 1, then set mvCol, refIdxCol, and listCol to mvLYCol[xColCb][yColCb], refIdxLYCol[xColCb][yColCb], and LY, respectively, where Y equals ! X, where X is the value of X that called the procedure. Set availableFlagLXCol to 1.
[0117] - Set both components of mvLXCol to 0, and set availableFlagLXCol to equal 0.
[0118] - When availableFlagLXCol equals TRUE, mvLXCol and availableFlagLXCol are derived as follows:
[0119] - If LongTermRefPic(currPic, currCb, refIdxLX, LX) is not equal to LongTermRefPic(ColPic, colCb, refIdxCol, listCol), then set both components of mvLXCol to 0 and set availableFlagLXCol to 0.
[0120] - Otherwise, set the variable availableFlagLXCol to equal 1, set refPicListCol[refIdxCol] to the image with reference index refIdxCol in the list of reference images listCol containing slices of coded blocks colCb in the juxtaposed images specified by ColPic, and the following applies:
[0121] colPocDiff=DiffPicOrderCnt(ColPic, refPicListCol[refIdxCol])
[0122] currPocDiff=DiffPicOrderCnt(currPic, RefPicListX[refIdxLX])
[0123] - If RefPicListX[refIdxLX] is a long-term reference image, or colPocDiff equals currPocDiff, then mvLXCol is exported as follows:
[0124] mvLXCol=mvCol
[0125] - Otherwise, export mvLXCol as a scaled version of the motion vector mvCol as follows:
[0126] tx=(16384+(Abs(td)>>1)) / td
[0127] distScaleFactor=Clip3(-4096, 4095, (tb*tx+32)>>6)
[0128] mvLXCol=Clip3(-32768, 32767, Sign(distScaleFactor*mvCol)*((Abs(distScaleFactor*mvCol)+127)>>8))
[0129] The td and tb are exported as follows:
[0130] td=Clip3(-128, 127, colPocDiff)
[0131] tb=Clip3(-128, 127, currPocDiff)
[0132] in,
[0133] The function LongTermRefPic(aPic, aPb, refIdx, LX), where X is 0 or 1, can be defined as follows:
[0134] - If an image with index refIdx of a list of reference images LX containing a slice of the predicted block appb from image aPic is marked as "for long-term reference" when aPic is the current image, then LongTermRefPic(aPic, aPb, refIdx, LX) equals 1.
[0135] Otherwise, LongTermRefPic(aPic, apb, refIdx, LX) equals 0.
[0136] and
[0137] The function DiffPicOrderCnt(picA, picB) is specified as follows:
[0138] DiffPicOrderCnt(picA,picB)=PicOrderCnt(picA)-PicOrderCnt(picB)
[0139] Therefore, to support temporal motion vector prediction, motion information from previously encoded images is stored. Typically, this information is stored in a temporal motion buffer. That is, the temporal motion buffer can include motion vectors determined in previously encoded images, which can be used to predict the motion vector of the current block in the current image. For example, as mentioned above, to derive mvLXCol, 18-bit MvLX[x][y] values from the juxtaposed images are required. It should be noted that motion vectors are typically stored in the temporal motion buffer using a finite set of precision (i.e., with the same precision as the derived motion vectors). Furthermore, a motion vector corresponding to each list of reference images can be stored for the sample blocks. A reference index corresponding to the image referenced by this motion vector can be stored. For each block of the previous images, a prediction mode (e.g., inter-frame prediction or non-inter-frame prediction) can be stored in the temporal buffer. Additionally, inter-frame prediction sub-modes (e.g., brightness illumination compensation mode markers) can be stored for the sample blocks.
[0140] Therefore, enabling temporal motion vector prediction using a typical temporal motion buffer can potentially incur significant storage costs. According to the techniques described herein, the resolution of the motion information can be varied to optimize the coding efficiency improvements resulting from temporal motion vector prediction, while simultaneously reducing the storage costs of implementing the temporal motion buffer. It should be noted that two methods exist to reduce the storage costs of implementing the temporal motion buffer. These methods can be used in combination or individually. One method is to utilize a dedicated implementation of the temporal motion buffer using lossless compression techniques for the motion information. For example, referring to the above-described derivation of mvLXCol, a set of 18-bit MvLX[x][y] values can be stored in the temporal motion buffer implementation using lossless compression techniques. Another method is to enable the temporal motion vector to be derived in a manner requiring less precision than the full precision of the temporal motion information. For example, according to the techniques described herein, as further described below, the above-described derivation of mvLXCol for the current image can be modified so that mvLXCol can be derived from the 16-bit values of MvLX[x][y] of the juxtaposed image.
[0141] Figure 5This is a block diagram illustrating an example of a system that can be configured to encode (e.g., encode and / or decode) video data according to one or more techniques of this disclosure. System 100 represents an example of a system that can perform video coding using motion vector prediction techniques according to one or more examples of this disclosure. Figure 5 As shown, system 100 includes source device 102, communication medium 110, and target device 120. Figure 5 In the example shown, source device 102 may include any device configured to encode video data and transmit the encoded video data to communication medium 110. Target device 120 may include any device configured to receive and decode the encoded video data via communication medium 110. Source device 102 and / or target device 120 may include computing devices equipped for wired and / or wireless communication, and may include set-top boxes, digital video recorders, televisions, desktop computers, laptops or tablets, game consoles, and mobile devices, including, for example, smartphones, cellular phones, personal gaming devices, and medical imaging equipment.
[0142] Communication medium 110 may include any combination of wireless and wired communication media and / or storage devices. Communication medium 110 may include coaxial cable, fiber optic cable, twisted-pair cable, wireless transmitters and receivers, routers, switches, repeaters, base stations, or any other device that can be used to facilitate communication between various devices and sites. Communication medium 110 may include one or more networks. For example, communication medium 110 may include a network configured to allow access to the World Wide Web, such as the Internet. The network may operate according to a combination of one or more telecommunications protocols. Telecommunication protocols may include proprietary aspects and / or may include standardized telecommunications protocols. Examples of standardized telecommunications protocols include the Digital Video Broadcasting (DVB) standard, the Advanced Television Systems Committee (ATSC) standard, the Integrated Services Digital Broadcasting (ISDB) standard, the Cable Data Services Interface Specification (DOCSIS) standard, the Global System for Mobile Communications (GSM) standard, the Code Division Multiple Access (CDMA) standard, the 3rd Generation Partnership Project (3GPP) standard, the European Telecommunications Standards Institute (ETSI) standard, the Internet Protocol (IP) standard, the Wireless Application Protocol (WAP) standard, and the Institute of Electrical and Electronics Engineers (IEEE) standard.
[0143] Storage devices can include any type of device or storage medium capable of storing data. Storage media can include tangible or non-transitory computer-readable media. Computer-readable media can include optical discs, flash memory, magnetic storage, or any other suitable digital storage medium. In some examples, a memory device or a portion thereof may be described as non-volatile memory, and in other examples, a portion of a memory device may be described as volatile memory. Examples of volatile memory can include random access memory (RAM), dynamic random access memory (DRAM), and static random access memory (SRAM). Examples of non-volatile memory can include magnetic hard disks, optical discs, floppy disks, flash memory, or electrically programmable memory (EPROM) or electrically erasable and programmable (EEPROM) memory. One or more storage devices can include memory cards (e.g., secure digital (SD) memory cards), internal / external hard disk drives, and / or internal / external solid-state drives. Data can be stored on the storage device according to a defined file format.
[0144] Refer again Figure 5 Source device 102 includes a video source 104, a video encoder 106, and an interface 108. The video source 104 may include any device configured to capture and / or store video data. For example, the video source 104 may include a camera and a storage device operatively coupled thereto. The video encoder 106 may include any device configured to receive video data and generate a compliant bitstream representing the video data. A compliant bitstream may refer to a bitstream from which a video decoder can receive and reproduce video data. Aspects of a compliant bitstream may be defined according to a video coding standard. When generating a compliant bitstream, the video encoder 106 may compress the video data. Compression may be lossy (perceptible or imperceptible) or lossless. The interface 108 may include any device configured to receive a compliant video bitstream and transmit and / or store the compliant video bitstream to a communication medium. The interface 108 may include a network interface card such as an Ethernet card and may include an optical transceiver, an RF transceiver, or any other type of device capable of transmitting and / or receiving information. Furthermore, interface 108 may include a computer system interface that enables compliant video bitstreams to be stored on a storage device. For example, interface 108 may include protocols supporting Peripheral Component Interconnect (PCI) and High-Speed Peripheral Component Interconnect (PCIe) bus protocols, dedicated bus protocols, Universal Serial Bus (USB) protocols, and I / O protocols. 2 C can be a chipset or any other logical and physical structure that can be used to interconnect peer devices.
[0145] Refer again Figure 5The target device 120 includes an interface 122, a video decoder 124, and a display 126. Interface 122 may include any device configured to receive compliant video bitstreams from a communication medium. Interface 108 may include a network interface card such as an Ethernet card, and may include an optical transceiver, an RF transceiver, or any other type of device capable of receiving and / or transmitting information. Furthermore, interface 122 may include a computer system interface enabling the retrieval of compliant video bitstreams from a storage device. For example, interface 122 may include protocols supporting PCI and PCIe bus protocols, dedicated bus protocols, USB protocols, etc. 2 The chip set of C or any other logical and physical structure that can be used to interconnect peer devices. The video decoder 124 may include any device configured to receive compliant bitstreams and / or acceptable variations thereof and reproduce video data therefrom. The display 126 may include any device configured to display video data. The display 126 may include one of a variety of display devices such as a liquid crystal display (LCD), a plasma display, an organic light-emitting diode (OLED) display, or another type of display. The display 126 may include a high-definition display or an ultra-high-definition display. It should be noted that, although in Figure 5 In the example shown, video decoder 124 is described as outputting data to display 126, but video decoder 124 can be configured to output video data to various types of devices and / or their sub-components. For example, video decoder 124 can be configured to output video data to any communication medium, as described herein.
[0146] Figure 6 This is a block diagram illustrating an example of a video encoder 200 that can implement the techniques described herein for encoding video data. It should be noted that although the exemplary video encoder 200 is shown as having different functional blocks, such illustrations are intended for descriptive purposes and do not limit the video encoder 200 and / or its sub-components to a particular hardware or software architecture. The functionality of the video encoder 200 can be implemented using any combination of hardware, firmware, and / or software implementations. In one example, the video encoder 200 can be configured to encode video data according to the techniques described herein. The video encoder 200 can perform intra-frame predictive coding and inter-frame predictive coding of picture regions, and therefore can be referred to as a hybrid video encoder. Figure 6In the example shown, video encoder 200 receives a source video block. In some examples, the source video block may include picture regions that have been partitioned according to the coding structure. For example, source video data may include macroblocks, CTUs, CBs, their sub-partitions, and / or additional equivalent coding units. In some examples, video encoder 200 may be configured to perform additional subdivision of the source video block. It should be noted that some of the techniques described herein are generally applicable to video coding, regardless of how the source video data is partitioned before and / or during encoding. Figure 6 In the example shown, the video encoder 200 includes a summer 202, a transform coefficient generator 204, a coefficient quantization unit 206, an inverse quantization / transformation processing unit 208, a summer 210, an intra-frame prediction processing unit 212, an inter-frame prediction processing unit 214, a filter unit 216, and an entropy coding unit 218.
[0147] like Figure 6 As shown, video encoder 200 receives source video blocks and outputs a bitstream. Video encoder 200 generates residual data by subtracting a predicted video block from the source video block. Summer 202 represents the component configured to perform this subtraction operation. In one example, the subtraction of the video block occurs in the pixel domain. Transform coefficient generator 204 applies transforms such as Discrete Cosine Transform (DCT), Discrete Sine Transform (DST), or conceptually similar transforms (e.g., applying four 8×8 transforms to a 16×16 residual value array) to the residual block or its sub-partitions to produce a set of residual transform coefficients. Transform coefficient generator 204 can be configured to perform any and all combinations of transforms included in the discrete trigonometric transform series. Transform coefficient generator 204 can output the transform coefficients to coefficient quantization unit 206. Coefficient quantization unit 206 can be configured to perform quantization of the transform coefficients. As mentioned above, the degree of quantization can be modified by adjusting the quantization parameters. Coefficient quantization unit 206 may be further configured to determine quantization parameters (QP) and output QP data (e.g., data for determining quantization group size and / or incremental QP values), which the video decoder can use to reconstruct the quantization parameters to perform inverse quantization during video decoding. It should be noted that in other examples, one or more additional or alternative parameters (e.g., scaling factors) may be used to determine the quantization level. The techniques described herein are generally applicable to determining the quantization level of transform coefficients corresponding to another component of video data based on the quantization level of transform coefficients corresponding to one component of the video data.
[0148] like Figure 6 As shown, the quantized transform coefficients are output to the inverse quantization / transform processing unit 208. The inverse quantization / transform processing unit 208 can be configured to apply inverse quantization and / or inverse transform to generate reconstructed residual data. For example... Figure 6As shown, at summer 210, the reconstructed residual data can be added to the predicted video block. This allows for the reconstruction of the encoded video block, which can then be used to evaluate the coding quality of a given prediction, transform, and / or quantization. The video encoder 200 can be configured to perform multiple coding rounds (e.g., coding while changing one or more of the prediction, transform, and quantization parameters). The rate distortion or other system parameters of the bitstream can be optimized based on the evaluation of the reconstructed video block. Furthermore, the reconstructed video block can be stored and used as a reference for predicting subsequent blocks.
[0149] As described above, video blocks can be encoded using intra-prediction modes. Intra-prediction processing unit 212 can be configured to select an intra-prediction mode for the current video block. Intra-prediction processing unit 212 can be configured to evaluate frames and / or regions thereof, and determine the intra-prediction mode used for encoding the current block. Figure 6 As shown, the intra-frame prediction processing unit 212 outputs intra-frame prediction data (e.g., syntax elements) to the entropy coding unit 218 and the transform coefficient generator 204. As described above, possible intra-frame prediction modes may include planar prediction mode, DC prediction mode, and angular prediction mode. The inter-frame prediction processing unit 214 can be configured to perform inter-frame prediction coding for the current video block. The inter-frame prediction processing unit 214 can be configured to receive the source video block and calculate the motion information of the PU of the video block. The motion vector can indicate the displacement of the PU (or similar coding structure) of the video block in the current video frame relative to the prediction block in the reference frame. Inter-frame prediction coding can use one or more reference pictures. For example, the inter-frame prediction processing unit 214 can locate the prediction video block in the frame buffer ( Figure 6 (Not shown in the image). It should be noted that the inter-frame prediction processing unit 214 can be further configured to apply one or more interpolation filters to the reconstructed residual block to compute sub-integer pixel values for motion estimation. Furthermore, motion prediction can be unidirectional prediction (using one motion vector) or bidirectional prediction (using two motion vectors). The inter-frame prediction processing unit 214 can be configured to select prediction blocks by computing pixel differences determined by, for example, sum of absolute differences (SAD), sum of squared differences (SSD), or other difference metrics. The inter-frame prediction processing unit 214 can output the motion prediction data of the computed motion vectors to the entropy coding unit 218.
[0150] As described above, motion information can be determined and specified using motion vector prediction techniques. The inter-frame prediction processing unit 214 can be configured to perform motion vector prediction techniques, including those described above. Furthermore, the inter-frame prediction processing unit 214 can be configured to perform motion vector prediction according to the aforementioned techniques. Specifically, the inter-frame prediction processing unit 214 can be configured to perform temporal motion vector prediction. As described above, enabling temporal motion vector prediction using a typical temporal motion buffer may require significant storage costs.
[0151] In one example, according to the techniques described herein, the inter-frame prediction processing unit 214 may be configured to store the value representing the motion vector using a reduced number of bits compared to the number of bits used to generate the motion vector (i.e., the motion vector used to generate the prediction block). It should be noted that, in addition to the methods for specifying motion vectors provided above and / or as alternatives to the methods provided above for specifying motion vectors, several other methods may exist for specifying motion vectors (e.g., polar coordinates, etc.). Regardless of how the motion vector is specified, the techniques described herein are generally applicable.
[0152] As mentioned above, enabling time motion vector prediction using a typical time motion buffer can potentially incur significant storage costs. In one example, the time motion buffer may contain the following fields for each block (e.g., a PU in ITU-T H.265) or sub-block:
[0153] - A 1-bit flag indicating whether the block is encoded using inter-frame prediction, such as the field isInter;
[0154] - A 2-bit flag indicating the direction of inter-frame prediction, such as the interDir field (e.g., list 0, list 1, or bidirectional prediction);
[0155] - An unsigned 16-bit member identity identifier for tile sets, such as the tileSetIdx field;
[0156] - Unsigned 16-bit slice index value, for example, the sliceIdx field:
[0157] - For each motion vector component, there is a signed 18-bit integer. For example, the field mv[MAX_RPL_COUNT][2], where [2] represents the horizontal and vertical displacement directions, max_RPL_COUNT is 2 and corresponds to list0 and list1.
[0158] - An unsigned 4-bit integer array indicating the reference image index, such as referenceIdx[MAX_RPL_COUNT].
[0159] It should be noted that an unsigned 4-bit value can represent a reference picture index, since the maximum number of pictures in the Reference Picture List (RPL) in JVET-L1001 is 16. Furthermore, it should be noted that in some cases, such as if an additional reference picture list is supported, max_RPL_COUNT can be increased.
[0160] As described above, in JVET-L1001, mvLX[0] and mvLX[1] are each derived as 18-bit values. Furthermore, in JVET-L1001, cpmvLX[0] and cpmvLX[1] are each derived as 18-bit values. In one example, according to the techniques described herein, mvLX[0], mvLX[1], cpmvLX[0], and cpmvLX[1] can be stored as N-bit values in the time-motion buffer, where N is less than 18, or more generally, where N is less than the number of bits in the derived value of the motion vector component. In one example, the value representing the motion vector component can be clipped to a range such that the value is represented by N bits before being stored in the time-motion buffer. For example, MV x It can be stored as Clip3(-2^ 15 ,2^ 15 -1, MV x ), MV y It can be stored as Clip3(-2^ 15 ,2^ 15 -1, MV y Compared to the juxtaposed motion vector derivation in JVET-L1001 provided above, in an example where the values representing the motion vector components are clipped, the stored value of mvCol can be as follows:
[0161] mvCol=Clip3(-(1<<(N-1)), (1<<(N-1))-1, mvCol)
[0162] In another example, only a subset of the N MSB bits of the derived motion vector components can be stored in the time motion buffer. For example, for MV x and MV y The following values can be stored in the time motion buffer:
[0163] MVDD x =MV x / 4; and
[0164] MVDD y ==MV y / 4
[0165] In this example, during reading from the time motion buffer, the MV can be derived as follows. x and MVy MV xR and MV yR Reconstructed values:
[0166] MV xR =4*MVDD x :and
[0167] MV yR =4*MVDD y
[0168] Compared to the juxtaposed motion vector export in JVET-L1001 provided above, in an example where only one subgroup of the N MSB bits of the exported motion vector components is stored in the time motion buffer, the stored value of mvCol can be as follows:
[0169] mvCol = (mvCol >> 4)
[0170] Furthermore, the exported values of mvCol can be as follows:
[0171] mvCol = (mvCol << 4)
[0172] It should be pointed out that in some cases, MV xR and / or MV yR It can be equal to MVx and MV y And in other cases, MV xR and / or MV yR Not the same as MVX and MV y Therefore, in some examples, it can be MV. xR and / or MV yR Provide incremental values conditionally, so that in MV xR and / or MV yR Not equal to MV x and MV y MV can be exported under certain circumstances x and / or MV y For example, in one example, the time motion buffer can be configured for each MV. xR and MV yR The entry includes a flag indicating whether a corresponding increment value exists in the time motion buffer. In this case, relative to the derivation of the juxtaposed motion vector in JVET-L1001 provided above, in one example, the stored value of mvCol could be as follows:
[0173] mvCol = (mvCol >> 4)
[0174] Furthermore, the exported values of mvCol can be as follows:
[0175] mvCol = (mvCol << 4) + increment value
[0176] The increment value is conditionally based on the existence of the flag value, and is inferred to be zero when it does not exist.
[0177] In another example, only a subset of N LSB bits of the derived motion vector component may be stored. For example, the motion vector component derived as X′ bit may be stored as N bits, where (NX) MSB bits are not stored. In one example, the (NX) MSB bits may be obtained using neighboring spatial motion vectors (e.g., by averaging neighboring spatial motion vectors, or, for example, by deriving (NX) MSB bits using motion vectors at predetermined locations with predetermined priorities). For example, in one example, the time motion buffer may include an incremental value for deriving (NX) MSB bits from neighboring spatial motion vectors. In this case, relative to the derivation of the juxtaposed motion vectors in JVET-L1 001 provided above, in one example, the stored value of mvCol may be as follows:
[0178] mvCol = (mvCol & LSBMask), where LSBMask == 0x000000FF
[0179] Furthermore, the exported values of mvCol can be as follows:
[0180] mvCol = mvCol | (mvAdjacent & MSBMask), where MSBMask = 0xFFFFFF00, and mvAdjacent is derived from a spatially adjacent motion vector, for example, from the left side of the current block.
[0181] Furthermore, in another example, the above techniques can be combined such that for a motion vector component of bit A, a subgroup of bits excluding B MSB bits and C LSB bits is stored (i.e., the middle bit is stored).
[0182] In another example, higher resolution can be used to store small motion vector components. For example, if the derived motion vector components are within a specific range (e.g., -2^3),... 15 up to (2^ 15If the value is -1 (including the endpoint), then the motion vector component can be determined to be a small motion vector component. In one example, this range can be empty (i.e., no signal is given to the small motion vector flag and it is inferred to be 0). In one example, a flag indicating whether a motion vector is a small motion vector can be stored in or read from the time motion buffer. In this case, in one example, when the motion vector component is a small motion vector component, the small motion vector flag is set to 1 and the small motion vector value is stored (e.g., stored at 1 / 16-peI resolution), and when the motion vector component is not a small motion vector component, the small motion vector flag is set to 0, and the motion vector component is scaled (e.g., scaled by dividing by 4 (right shift of 2, with / without a rounding offset of 1)) and stored (e.g., stored at 1 / 4-peI resolution). Furthermore, in this case, when reading a small motion vector component from the time motion buffer, when the small motion vector is marked as 1 (which can be inferred based on the range of values read from the motion buffer), the motion vector component value is read from the time motion buffer at the resolution associated with the small motion vector, and when the small motion vector is marked as 0, the motion vector value is read from the time motion buffer at another determined resolution (e.g., 1 / 4-pel resolution) and scaled (e.g., multiplied by 4 - left-shifted by 2) to the small motion vector resolution (e.g., 1 / 16-pel resolution). In this case, relative to the derivation of the juxtaposed motion vector in JVET-L100l provided above, in one example, the stored value of mvCol could be as follows:
[0183] mvCol=(smallMotionVectorFlag)? mvCol: (mvCol>>2)
[0184] Furthermore, the exported values of mvCol can be as follows:
[0185] mvCol=(smallMotionVectorFlag)? mvCol: (mvCol<<2)
[0186] In another example, if the stored value is near an N-bit wrapper edge (e.g., 8 bits for LSBs > 229 or < 25), an automatic subtraction / addition (e.g., automatic subtraction / addition of 128) can be applied to the stored value. In one example, the average of the spatial candidates for the current image can be used to predict this subtraction / addition. In one example, the motion vector can be cropped to 8 bits, and if the value is + / - 128, it is not included as a candidate. In one example, an offset factor can be signaled, and then applied before cropping. In this example, if the result is at the cropping point, it may not be included as a motion vector candidate in some cases. For example, if the derived motion vector is 0x0F, then 4 LSB bits are 0xF. If the average of the spatial candidates is 0x10, simply concatenating 0x10 to 0x0F to provide 0x1F might not be desirable. Therefore, in one example, when approaching the edge of the package, an operation can be performed to move the LSB bit away from the edge (e.g., add 0x7). This operation can then be reversed after concatenation. Thus, for the above case, in one example, 0x0F + 0x07 = 0x16 can be performed, and 0x6 can be stored. In this case, the space average becomes 0x10 + 0x07 = 0x17. Concatenation will result in 0x16, and the final result of removing the addition is: 0x16 - 0x07 = 0x0F. In one example, whether it is close to the edge of the package can be determined by comparison with a threshold (e.g., 229 and 25).
[0187] In one example, according to the techniques described herein, the inter-frame prediction processing unit 214 can be configured to store values representing motion vector components using a clipping range corresponding to the precision of the motion vector calculation. As described above, in JVET-L1001, the scaling of the temporal / spatial / derived predicted motion vector is based on the Picture Order Count (POC) distance, where the POC distance corresponds to the temporal distance between pictures. The reference motion vector can be scaled to the target picture based on the ratio of the POC distance between the predicted picture and the reference picture to the POC distance between the current predicted picture and the reference picture. The scaled motion vector can be used for prediction. The temporal / spatial / derived motion vector scaling operation can be followed by a clipping operation for each displacement direction in the scaled motion vector displacement direction (e.g., Clip3(-2^ 15 ,2^ 15-1, scaledMv), where scaledMv is the motion vector component with the scaling motion vector displacement direction). In this case, when using a 1 / 16-pel motion vector during decoding, it is expected that the clipping boundary will not become an unnecessary constraint, as this could lead to clipping of large motion vectors. In one example, according to the techniques described herein, the inter-frame prediction processing unit 214 can be configured to allow a larger bit depth (e.g., 18 bits) in the scaling motion vector displacement direction that the clipping boundary requires (e.g., Clip3(-2^ 17 ,2^ 17 -1, scaledMv). In this case, relative to the derivation of the juxtaposed motion vectors in JVET-L1001 provided above, in one example, mvCol and mvLXCol can be derived as follows:
[0188] mvCol=Clip3(-131072, 131071, Sign(distScaleFactor*mvCol)*((Abs(distScaleFactor*mvCol)+127)>>8))
[0189] mvLXCol=Clip3(-131072, 131071, Sign(distScaleFactor*mvCol)*((Abs(distScaleFactor*mvCol)+127)>>8))
[0190] Furthermore, in this case, other motion vector prediction candidates can be pruned in a similar manner.
[0191] In one example, according to the techniques described herein, the inter-frame prediction processing unit 2l4 can be configured to use a subset of bits for a pruning process comparison. As described above, the pruning process can be used to remove redundant entries from a set of motion vector candidates. In one example, according to the techniques described herein, the pruning process can use a subset of bits to determine whether a motion candidate is redundant, and thus determine whether to add the motion candidate to the candidate list. Furthermore, in one example, according to the techniques described herein, the pruning process can be used to determine whether to store motion information about the block into a temporal motion buffer. In one example, for comparison purposes, the pruning process can use only a subset of bits of the motion vector components. For example, only x LSB bits of the motion vector components are used for comparison, or only x MSB bits of the motion vector components are used for comparison.
[0192] In one example, according to the techniques described herein, the inter-frame prediction processing unit 214 may be configured to use normalization to reduce the bit depth of the motion vector stored in the temporal motion buffer. For example, each 16×16 block of an exemplary architecture of the temporal motion buffer may contain the following 82 bits (or 90 bits):
[0193] 4MV components, each 16-bit (or 18-bit).
[0194] Two incremental POC values, each 8 bits
[0195] Two LTRP (Long-Term Reference Image) markers, 1 bit each.
[0196] In one example, according to the techniques described herein, it may be desirable to reduce the storage requirement to 64 bits in order to reduce memory bandwidth. In one example, the storage requirement can be reduced by: (1) storing a normalized motion vector (e.g., normalized relative to an incremental POC, which then does not necessarily need to be stored); (2) using a floating-point representation of the normalized motion vector (presumably, 1 / 16 precision is not very useful for large motion vectors); and / or cropping to a reduced range (i.e., the normalized MV should not be so large). In one example, normalization relative to an incremental POC may include determining an initial POC difference representing the difference between the current image and a reference current image's POC value. To reduce storage, the POC difference can be scaled to a constant value (e.g., 1, 2, 4), and the corresponding motion vector can be appropriately scaled considering the ratio of the original POC difference to the new POC difference. The cropping operation can be performed after the motion vector is scaled (i.e., during storage). Besides reducing storage, choosing a constant POC difference simplifies time scaling operations, which typically involve taking the ratio of two POC differences: one corresponding to the difference between the POC of the current image and the POC of the reference image, and the other corresponding to the difference between the POC of the time image and the POC of the time image's reference image. During storage, the latter can be set to a constant, thus reducing the computational complexity of the ratio of the two POC differences. For example, if the denominator is 1, the division operation can be skipped, or if the denominator is a power of 2, the division operation becomes a shift operation. Normalization can also be based on whether the motion vector references a long-term reference image. For example, motion vectors referencing a long-term reference image may not be normalized. In this case, relative to the derivation of the juxtaposed motion vectors in JVET-L1001 provided above, in one example, mvCol can be derived as follows:
[0197] When the image corresponding to refPicListCol[refIdxCol] is not a long-term reference image, the following steps are performed in sequence:
[0198] colPocDiff=DiffPicOrderCnt[ColPic, refPicListCol[refIdxCol])
[0199] newColPocDiff = CONSTANT
[0200] tx=(16384+(Abs(td)>>1)) / td
[0201] distScaleFactor=Clip3(-4096, 4095, (tb*tx+32)>>6)
[0202] mvCol=Clip3(-131072, 131071, Sign(distScaleFactor*mvCol)*((Abs(distScaleFactor*mvCol)+127)>>8))
[0203] The td and tb are exported as follows:
[0204] td=Clip3(-128, 127, colPocDiff)
[0205] tb=Clip3(-128, 127, newColPocDiff)
[0206] - When availableFlagLXCol equals TRUE, mvLXCol and availableFlagLXCol are derived as follows:
[0207] - If LongTermRefPic(currPic, currCb, refIdxLX, LX) is not equal to LongTermRefPic(ColPic, colCb, refIdxCol, listCol), then set both components of mvLXCol to 0 and set availableFlagLXCol to 0.
[0208] - Otherwise, set the variable availableFlagLXCol to equal 1, set refPicListCol[refIdxCol] to the image with reference index refIdxCol in the list of reference images listCol containing slices of coded blocks colCb in the juxtaposed images specified by ColPic, and the following applies:
[0209] colPocDiff = newColPocDiff
[0210] currPocDiff=DiffPicOrderCnt(currPic, RefPicListX[refIdxLX])
[0211] - If RefPicListX[refIdxLX] is a long-term reference image, or colPocDiff equals currPocDiff, then mvLXCol is exported as follows:
[0212] mvLXCol=mvCol
[0213] - Otherwise, export mvLXCol as a scaled version of the motion vector mvCol as follows:
[0214] tx=(16384+(Abs(td)>>1)) / td
[0215] distScaleFactor=Clip3(-4096, 4095, (tb*tx+32)>>6)
[0216] mvLXCol=Clip3(-131072, 131071, Sign(distScaleFactor*mvCol)*((Abs(distScaleFactor*mvCol)+127)>>8))
[0217] The td and tb are exported as follows:
[0218] td=Clip3(-128, 127, colPocDiff)
[0219] tb=Clip3(-128, 127, currPocDiff)
[0220] As described above, in one example, the storage requirements of the time-motion buffer can be reduced by: storing the normalized motion vector; using a floating-point representation of the normalized motion vector; and / or cropping to a reduced range. Furthermore, in one example, the storage requirements of the time-motion buffer can be reduced by converting the motion vector displacement to a mantissa-exponential representation before storage. In one example, if the range of this representation is greater than a threshold (e.g., 18 bits), then cropping the motion vector displacement is not necessary. In one example, the mantissa can be 6 bits and range from -32 to 31 (inclusive), and the exponent can be 4 bits and range from 0 to 15 (inclusive). Additionally, in one example, the motion vector of the reference long-term image can also be scaled and / or converted to a floating-point representation before storage.
[0221] In one example, the mantissa-exponent representation of a number can be calculated as follows:
[0222]
[0223] in,
[0224] static const int MV_EXPONENT_BITCOUNT=4;
[0225] static const int MV_MANTISSA_BITCOUNT=6;
[0226] static const int MV_MANTISSA_LOWER_LIMIT=-(1<<(MV_MANTISSA_BITCOUNT-1));
[0227] static const int MV MANTISSA_UPPER_LIMIT=((1<<(MV_MANTISSA_BITCOUNT-1))-1);
[0228] static const int MV_MANTISSA_LIMIT = (1 << (MV_MANTISSA_BITCOUNT - 1)); and
[0229] static const int MV_EXPONENT_MASK=((1<<MV_EXPONENT_BITCOUNT)-1);
[0230] Furthermore, the conversion from mantissa-exponent representation to number can be performed as follows:
[0231]
[0232] Therefore, in one example, according to the techniques described in this paper, the derivation of the juxtaposed motion vector can be performed as follows:
[0233] The input to this process is:
[0234] - The variable currCb specifies the current encoded block.
[0235] - The variable colCb specifies the juxtaposition encoding block within the juxtaposition image specified by ColPic.
[0236] -Luminance position (xColCb, yColCb), which specifies the top-left sample of the juxtaposed luminance coding block specified by colCb relative to the top-left luminance sample of the juxtaposed image specified by ColPic.
[0237] - Refer to the index refIdxLX, where X is 0 or 1.
[0238] - The flag sbFlag indicates the candidate for time merging of sub-blocks.
[0239] The output of this process is:
[0240] Predicted motion vector mvLXCol with -1 / 16 fractional sampling accuracy.
[0241] -Availability flag availableFlagLXCol.
[0242] The variable currPic specifies the current image.
[0243] Set the arrays predFlagL0Col[x][y], mvL0Col[x][y], and refIdxL0Col[x][y] to be equal to the PredFlagL0[x][y], MvL0[x][y], and RefIdxL0[x][y] of the juxtaposed images specified by ColPic, respectively. Set the arrays predFlagL1Col[x][y], mvL1Col[x][y], and refIdxL1Col[x][y] to be equal to the PredFlagL1[x][y], MvL1[x][y], and RefIdxL1[x][y] of the juxtaposed images specified by ColPic, respectively.
[0244] The variables mvLXCol and availableFlagLXCol are exported as follows:
[0245] - If colCb is encoded in intra-prediction mode, then set both components of mvLXCol to 0 and set availableFlagLXCol to 0.
[0246] - Otherwise, the motion vector mvCol, the reference index refIdxCol, and the reference list identifier listCol are derived as follows:
[0247] - If sbFlag equals 0, then set availableFlagLXCol to 1, and the following applies:
[0248] - If predFlagL0Col[xColCb][yColCb] equals 0, then set mvCol, refIdxCol, and listCol to equal mvL1Col[xColCb][yColCb], refIdxL1Col[xColCb][yColCb], and L1, respectively.
[0249] Otherwise, if predFlagL0Col[xColCb][yColCb] equals 1 and predFlagL1Col[xColCb][yColCb] equals 0, then set mvCol, refIdxCol, and listCol to equal mvL0Col[xColCb][yColCb], refIdxL0Col[xColCb][yColCb], and L0, respectively.
[0250] - Otherwise (predFlagL0Col[xColCb][yColCb] equals 1, and predFlagL1Col[xColCb][yColCb] equals 1), specify the following:
[0251] - If NoBackwardPredFlag equals 1, then set mnvCol, refIdxCol, and listCol to equal mvLXCol[xColCb][yColCb], refIdxLXCol[xColCb][yColCb], and LX, respectively.
[0252] Otherwise, set mvCol, refIdxCol, and listCol to be equal to mvLNCol[xColCb][yColCb], refIdxLNCol[xColCb][yColCb], and LN, respectively, where N is the value of collocated_from_10flag.
[0253] - Otherwise (sbFlag equals 1), the following applies:
[0254] - If PredFlagLXCol[xColCb][yColCb] equals 1, then set mvCol, refIdxCol, and listCol to equal mvLXCol[xColCb][yColCb], refIdxLXCol[xColCb][yColCb], and LX respectively, and set availableFlagLXCol to 1.
[0255] - Otherwise (PredFlagLXCol[xColCb][yColCb] equals 0), the following applies:
[0256] If for each image in the reference image list for the current slice, DiffPicOrderCnt(aPic, currPic) is less than or equal to 0, and PredFlagLYCol[xColCb][yColCb] is equal to 1, then set mvCol, refIdxCol, and listCol to mvLYCol[xColCb][yColCb], refIdxLYCol[xColCb][yColCb], and LY, respectively, where Y equals ! X, where X is the value of X that called the procedure. Set availableFlagLXCol to 1.
[0257] - Set both components of mvLXCol to 0, and set availableFlagLXCol to equal 0.
[0258] - When availableFlagLXCol equals TRUE, mvLXCol and availableFlagLXCol are derived as follows:
[0259] - If LongTermRefPic(currPic, currCb, refIdxLX, LX) is not equal to LongTermRefPic(ColPic, colCb, refIdxCol, listCol), then both components of mnvLXCol are set to 0, and availableFlagLXCol is set to 0.
[0260] - Otherwise, set the variable availableFlagLXCol to equal 1, set refPicListCol[refIdxCol] to the image with reference index refIdxCol in the list of reference images listCol containing slices of coded blocks colCb in the juxtaposed images specified by ColPic, and the following applies:
[0261] colPocDiff=DiffPicOrderCnt(ColPic, refPicListCol[refIdxCol])
[0262] currPocDiff=DiffPicOrderCnt(currPic, RefPicListX[refIdxLX])
[0263] - Export mvColScaled as a scaled version of the motion vector mvCol as follows:
[0264] txScaled=(16384+(Abs(tdScaled)>>1)) / 4
[0265] distScaleFactor=Clip3(-4096, 4095, (4*txScaled+32)>>6)
[0266] mvColScaled=Clip3(-32768, 32767, Sign(distScaleFactor*mvCol)*((Abs(distScaleFactor*mvCol)+127)>>8))
[0267] The tdScaled component is exported as follows:
[0268] tdScaled=Clip3(-128, 127, colPocDiff)
[0269] - Use mvColScaled as input to call the export procedure specified in item A (below), which converts the motion vector displacement to mantissa and exponential representation, and assign the output to mvColScaledME.
[0270] - Use mvColScaledME as input to call the procedure specified in item B (below) for deriving the motion vector displacement from the mantissa and exponential representations, and assign the output to mvColScaledQ.
[0271] - If RefPicListX[refIdxLX] is a long-term reference image, or colPocDiff equals currPocDiff, then mvLXCol is exported as follows:
[0272] mvLXCol==mvColScaledQ
[0273] - Otherwise, export mvLXCol as a scaled version of the motion vector mvColScaledQ as follows:
[0274] tx = 4096
[0275] distScaleFactor=Clip3(-4096, 4095, (tb*tx+32)>>6)
[0276] mvLXCol=Clip3(-32768, 32767, Sign(distScaleFactor*mvColScaledQ)*((Abs(distScaleFactor*mvColScaledQ)+127)>>8))
[0277] The tb file is exported as follows:
[0278] tb=Clip3(-128, 127, currPocDiff)
[0279] Option A describes the process of converting motion vector displacement into a representation of the mantissa and exponent.
[0280] The input to this process is:
[0281] - Specifies the variable mvColScale for scaling the time motion vector.
[0282] The output of this process is:
[0283] - Specifies the variable mvColScaledME, which represents the mantissa exponent of the scaled time motion vector.
[0284] The variables mantissa and exponent are exported as follows:
[0285]
[0286] The variable mvColScaledME is exported as follows:
[0287] mvColScaledME=exponent|(mantissa<<4)
[0288] Option B describes the process of deriving the motion vector displacement from the mantissa and exponent representations.
[0289] The input to this process is:
[0290] - Specifies the variable mvColScaledME, which represents the scaled time motion vector using mantissa exponent.
[0291] The output of this process is:
[0292] - Specifies the variable mvColScaledQ, which is derived from its mantissa representation of the scaled time motion vector.
[0293] The variables mantissa and exponent are exported as follows:
[0294] exponent = mvColScaledME & 15
[0295] mantissa=mvColScaledME>>4
[0296] The variable mvColScaledQ is exported as follows:
[0297]
[0298] In one example, according to the techniques described herein, the mantissa and exponent, which reach predetermined values (e.g., -32 and 15, respectively) used to indicate the displacement of a temporal motion vector, are not available for motion prediction. For example, when using intra-frame mode, all four motion vector displacement direction values are assigned a mantissa of -32 and an exponent of 15. Similarly, when only one of two motion vectors is valid (e.g., inter_pred_idc[][] is PRED_L0 or PRED_L1), both displacement directions of the motion vector that does not have valid motion information are assigned a mantissa of -32 and an exponent of 15. This method of indicating availability can also be applied to temporal motion information corresponding to the current image reference.
[0299] In one example, calculating the mantissa-exponential representation of a number may involve a two-step process. The first step is to determine the quantization interval, and the second step is to add half of the quantization interval, and then quantize. The mantissa-exponential representation can be calculated as follows:
[0300]
[0301]
[0302] in,
[0303] static const int MV_EXPONENT_BITCOUNT=4;
[0304] static const int MV_MANTISSA_BITCOUNT=6;
[0305] static const int MV_MANTISSA_LOWER_LIMIT=-(1<<(MV_MANTISSA_BITCOUNT-1));
[0306] static const int MV_MANTISSA_UPPER_LIMIT=((1<<(MV_MANTISSA_BITCOUNT-1))-1);
[0307] static const int MV_MANTISSA_LIMIT = (1 << (MV_MANTISSA_BITCOUNT - 1)); and
[0308] static const int MV_EXPONENT_MASK=((1<<MV_EXPONENT_BITCOUNT)-1);
[0309] Furthermore, in one example of mantissa-exponent representation calculation, the calculation (quantization) of the mantissa and exponent for an input number with a rounding offset can be performed as follows:
[0310] int numberWithOffset=number+(((1<<(exponent-1))-((number<0)?0:1))>>1);
[0311] Furthermore, in one example of mantissa-exponent representation calculation, the calculation (quantization) of the mantissa and exponent for an input number with a rounding offset can be performed as follows:
[0312] int numberWithOffset=number+(((1<<(exponent-1))-((number<0)?1:0))>>1);
[0313] Therefore, in one example, according to the techniques described in this paper, the derivation of the juxtaposed motion vector can be performed as follows:
[0314] The call in item A is for the storage process of time-brightness motion information. :
[0315] Item A should be invoked as part of the decoding process for the current image. Item A should be invoked after all motion information (motion vectors, reference indices, prediction exploit tags, and tags from long-term reference images) for the current image has been decoded and its values determined. Item A (and the items invoked within Item A) should be the last item in the process of accessing motion information during the decoding of the current image.
[0316] The function LongTermRefPic(aPic, aCb, refIdx, LX), where X is 0 or 1, is defined as follows:
[0317] - If a picture with index refIdx of a list of reference pictures LX containing a slice of coded block aCb from picture aPic is marked as "for long-term reference" when aPic is the current picture, then LongTermRefPic(aPic, aCb, refIdx, LX) equals 1.
[0318] Otherwise, LongTermRefPic(aPic, aCb, refIdx, LX) equals 0.
[0319] Option A. The process used for storing temporal brightness motion information.
[0320] The input to this process is:
[0321] - The prediction list uses an array labeled PredFlagL0 and PredFlagL1.
[0322] - Array of brightness motion vectors MvL0 and MvL1
[0323] - Arrays referencing indices RefIdxL0 and RefIdxL1
[0324] The output of this process is:
[0325] - Arrays of stored luminance motion vectors, MvL0Mantissa, MvL0Exponent, MvL1Mantissa, and MvL1MExponent, represented by their mantissas and exponents.
[0326] - Arrays LtrpFlagL0 and LtrpFlagL1 indicating the use of markers for long-term reference images.
[0327] The variable currPic specifies the current image.
[0328] The following applies to xL values from 0 to (pic_width_in_luma_samples >> 3):
[0329] - The following applies to yL values from 0 to (pic_height_in_luma_samples >> 3):
[0330] - Variables x and y are assigned values (x << 3) and (y << 3) respectively.
[0331] - The variable currCb specifies the coded block at the luminance position (x, y).
[0332] - When PredFlagL0[x][y] equals 0, -32 is assigned to MvL0Mantissa[x][y][0] and MvL0Mantissa[x][y][1], and 15 is assigned to MvL0Exponent[x][y][0] and MvL0Exponent[x][y][1]. Otherwise, the export procedure for scaling and transforming motion vector displacements, as specified in item B, is invoked with MvL0[x][y] as input, and the output is assigned to MvL0Mantissa[x][y] and MvL0Exponent[x][y].
[0333] - Assign the output of LongTermRefPic(currPic, currCb, RefIdxL0[x][y], L0) to LtrpFlagL0[x][y].
[0334] - When PredFlagL1[x][y] equals 0, assign -32 to MvL1Mantissa[x][y][0] and MvL1Mantissa[x][y][1], and assign 15 to MvL1Exponent[x][y][0] and MvL1Exponent[x][y][1]. Otherwise, call the derived procedure for scaling the transformation motion vector displacement as specified in item B with MvLl[x][y] as input, and assign the output to MvL1Mantissa[x][y] and MvL1Exponent[x][y].
[0335] - Assign the output of LongTermRefPic(currPic, currCb, RefIdxL1[x][y], L1) to LtrpFlagL1[x][y].
[0336] The process of deriving juxtaposed motion vectors
[0337] The input to this process is:
[0338] - The variable currCb specifies the current encoded block.
[0339] - The variable colCb specifies the juxtaposition encoding block within the juxtaposition image specified by ColPic.
[0340] -Luminance position (xColCb, yColCb), which specifies the top-left sample of the juxtaposed luminance coding block specified by colCb relative to the top-left luminance sample of the juxtaposed image specified by ColPic.
[0341] - Refer to the index refIdxLX, where X is 0 or 1.
[0342] - The flag sbFlag indicates the candidate for time merging of sub-blocks.
[0343] The output of this process is:
[0344] Predicted motion vector mvLXCol with -1 / 16 fractional sampling accuracy.
[0345] -Availability flag availableFlagLXCol.
[0346] The variable currPic specifies the current image.
[0347] Set the arrays mvL0MantissaCol[x][y], mvL0ExponentCol[x][y], ltrpL0FlagCol[x][y], mvL1MantissaCol[x][y], mvL1ExponentCol[x][y], and ItrpL1FlagCol[x][y] to be equal to the juxtaposed images MvL0Mantissa[x][y], MvL0Exponent[x][y], LtrpL0Flagt[x][y], MvL1Manitssa[x][y], MvL1Exponent[x][y], and LtrpLlFlag[x][y] specified by ColPic, respectively.
[0348] When mvL0MantissaCol[x][y][0] and mvL0ExponentCol[x][y][0] are equal to -32 and 15 respectively, predFlagL0Col[x][y] is set to 0; otherwise, predFlagL0Col[x][y] is set to 1.
[0349] When mvL1MantissaCol[x][y][0] and mvL1ExponentCol[x][y][0] are equal to -32 and 15 respectively, predFlagL1Col[x][y] is set to 0; otherwise, predFlagL1Col[x][y] is set to 1.
[0350] The variables mvLXCol and availableFlagLXCol are exported as follows:
[0351] - If colCb is encoded in intra-prediction mode, then set both components of mvLXCol to 0 and set availableFlagLXCol to 0.
[0352] - Otherwise, the motion vector mvCol, the reference index refIdxCol, and the reference list identifier listCol are derived as follows:
[0353] - If sbFlag equals 0, then set availableFlagLXCol to 1, and the following applies:
[0354] - If predFlagL0Col[xColCb][yColCb] equals 0, then set mvCol and listCol to equal mvL1Col[xColCb][yColCb] and L1 respectively.
[0355] Otherwise, if predFlagL0Col[xColCb][yColCb] equals 1 and predFlagL1Col[xColCb][yColCb] equals 0, then set mvCol and listCol to equal mvL0Col[xColCb][yColCb] and L0 respectively.
[0356] - Otherwise (predFlagL0Col[xColCb][yColCb] equals 1, and predFlagL1Col[xColCb][yColCb] equals 1), specify the following:
[0357] - If NoBackwardPredFlag equals 1, then set mvCol and listCol to equal mvLXCol[xColCb][yColCb] and LX respectively.
[0358] Otherwise, set mvCol and listCol to be equal to mvLNCol[xColCb][yColCb] and LN respectively, where N is the value of collocated_from_10flag.
[0359] - Otherwise (sbFlag equals 1), the following applies:
[0360] - If PredFlagLXCol[xColCb][yColCb] equals 1, then set mvCol and listCol to equal mvLXCol[xColCb][yColCb] and LX respectively, and set availableFlagLXCol to 1.
[0361] - Otherwise (PredFlagLXCol[xColCb][yColCb] equals 0), the following applies:
[0362] - If for each image in the reference image list for the current slice, DiffPicOrderCm(aPic, currPic) is less than or equal to 0, and PredFlagLYCol[xColCb][yColCb] is equal to 1, then set mvCol and listCol to mvLYCol[xColCb][yColCb] and LY respectively, where Y equals ! X, where X is the value of X that called the procedure. Set availableFlagLXCol to 1.
[0363] - Set both components of mvLXCol to 0, and set availableFlagLXCol to equal 0.
[0364] - When availableFlagLXCol equals TRUE, mvLXCol and availableFlagLXCol are derived as follows:
[0365] - If LongTermRefPic(currPic, currCb, refIdxLX, LX) is not equal to ltrpLXFlagCol[xColCb][yColCb], then set both components of mvLXCol to 0 and set availableFlagLXCol to 0.
[0366] - Otherwise, set the variable availableFlagLXCol to equal 1, set refPicListCol[refIdxCol] to the image with reference index refIdxCol in the list of reference images listCol containing slices of coded blocks colCb in the juxtaposed images specified by ColPic, and the following applies:
[0367] colPocDiff=DiffPicOrderCnt(ColPic, refPicListCol[refIdxCol])
[0368] currPocDiff=DiffPicOrderCnt(currPic, RefPicListX[refIdxLX])
[0369] -Use mvL0MantissaCol[xColCb][yColCb] and mvL0ExponentCol[xColCb][yColCb] as input to call the procedure specified in item C for deriving the motion vector displacement from the mantissa and exponential representations, and assign the output to mvColScaledQ.
[0370] - If RefPicListX[refIdxLX] is a long-term reference image, or colPocDiff equals currPocDiff, then mvLXCol is exported as follows:
[0371] mvLXCol=mvColScaledQ
[0372] - Otherwise, export mvLXCol as a scaled version of the motion vector mvColScaledQ as follows:
[0373] tx = 4096
[0374] distScaleFactor=Clip3(-4096, 4095, (tb*tx+32)>>6)
[0375] mvLXCol=Clip3(-32768, 32767, Sign(distScaleFactor*mvColScaledQ)*((Abs(distScaleFactor*mvColScaledQ)+127)>>8))
[0376] The td is exported as follows:
[0377] tb=Clip3(-128, 127, currPocDiff)
[0378] Option B. The process of scaling the motion vector displacement and converting it to mantissa and exponent representation.
[0379] The input to this process is:
[0380] -Specify the array mvCol of input motion vectors,
[0381] The output of this process is:
[0382] - Specifies the mantissa and exponent representation of the scaled time motion vectors, such as mvColScaledMantissa and mvColScaledExponent.
[0383] The variable mvColScaled is exported as follows:
[0384] txScaled=(16384+(Abs(tdScaled)>>1)) / tdScaled
[0385] distColScaleFactor=(4*txScaled+32)>>6
[0386] mvColScaled=Clip3(-32768, 32767, Sign(distColScaleFactor*mvCol)*((Abs(distColScaleFactor*mvCol)+127)>>8))
[0387] The tdScaled component is exported as follows:
[0388] tdScaled=Clip3(-128, 127, colPocDiff)
[0389] The arrays mvColScaledMantissa and mvColScaledExponent are exported as follows:
[0390]
[0391] Alternatively, in one example, the arrays mvColScaledMantissa and mvColScaledExponent can be exported as follows:
[0392]
[0393] Alternatively, in one example, the arrays mvColScaledMantissa and mvColScaledExponent can be exported as follows:
[0394]
[0395] Option C. Used for deriving the motion vector displacement from the mantissa and exponent representations.
[0396] The input to this process is:
[0397] - Specifies the array of scaled time motion vectors, mvColScaledMantissa and mvColScaledExponent, expressed in mantissa and exponent form.
[0398] The output of this process is:
[0399] - Specifies the array of scaled time motion vectors derived from their mantissa and exponent representations, mvColScaledQ
[0400] The array mvColScaledQ is exported as follows:
[0401]
[0402] in,
[0403] dir represents the direction of displacement; for example, 0 represents horizontal and 1 represents vertical.
[0404] In one example, the clipping range used for motion vector values is modified to 18 bits, according to the techniques described herein. In one example, Clip3(-32768, 32767, ...) is modified to Clip3(-131072, 131071, ...).
[0405] In one example, the conversion from mantissa-exponent representation to the derived integer can be based on whether the exponent is a predetermined number. For example, when the exponent is equal to 0, the derived integer is equal to the mantissa.
[0406] In one example, according to the techniques described herein, the conversion from mantissa-exponent representation to integer can introduce bias. This is likely because mantissa-exponent representation is not perfectly symmetric (e.g., when the mantissa is a 6-bit signed integer). It may be desirable to have no sign-based bias. In one example, to avoid sign bias, the mantissa can include 5 unsigned integer bits and 1 sign bit. For a non-zero exponent, the reconstruction would be (MV_MANTISSA_LIMIT + mantissa) * (1 - 2 * sign) << (exponent - 1) instead of (mantissa ^ MV_MANTISSA_LIMIT) << (exponent - 1). Alternatively, if symmetry is required, it can be reconstructed from a 6-bit signed mantissa as follows: ((mantissa ^ MV_MANTISSA_LIMIT) + (mantissa < 0 ? 1 : 0) << (exponent - 1). In one example, MV_MANTISSA_LIMIT can be 32.
[0407] In one example, according to the techniques described in this paper, the mantissa should be sign-extended before applying the XOR operation. For example, the 6-bit signed mantissa should be extended to an 18-bit signed value before applying the XOR operation.
[0408] In one example, according to the technique described herein, the XOR operation between the mantissa and the number is replaced by adding the number to the mantissa to obtain a positive (i.e., greater than or equal to 0) value for the mantissa and subtracting the number from the mantissa to obtain a negative value for the mantissa.
[0409] In one example, according to the techniques described in this paper, the derivation of the juxtaposed motion vector can be performed as follows:
[0410] The call in item A is for the storage process of time-brightness motion information. :
[0411] Item A should be invoked as part of the decoding process for the current image. Item A should be invoked after all motion information (motion vectors, reference indices, prediction exploit tags, and tags from long-term reference images) for the current image has been decoded and its values determined. Item A (and the items invoked within Item A) should be the last item in the process of accessing motion information during the decoding of the current image.
[0412] Option A. The process used for storing temporal brightness motion information.
[0413] The input to this process is:
[0414] - The prediction list uses an array labeled PredFlagL0 and PredFlagL1.
[0415] - Array of brightness motion vectors MvL0 and MvL1
[0416] - Arrays referencing indices RefIdxL0 and RefIdxL1
[0417] The output of this process is:
[0418] - Arrays of stored luminance motion vectors, MvL0Mantissa, MvL0Exponent, MvL1Mantissa, and MvL1MExponent, represented by their mantissas and exponents.
[0419] The variable currPic specifies the current image.
[0420] The following applies to xL values from 0 to (pic_width_in_luma_samples >> 3):
[0421] - The following applies to yL values from 0 to (pic_height_in_luma_samples >> 3):
[0422] - Variables x and y are assigned values (x << 3) and (y << 3) respectively.
[0423] - The variable currCb specifies the coded block at the luminance position (x, y).
[0424] - Using MvL0[x][y] as input, call the derivation procedure specified in section B to convert the motion vector displacement to mantissa and exponential representation, and assign the output to MvL0Mantissa[x][y] and MvL0Exponent[x][y].
[0425] - Using MvL1[x][y] as input, call the derivation procedure specified in section B to convert the motion vector displacement to mantissa and exponential representation, and assign the output to MvL1Mantissa[x][y] and MvL1Exponent[x][y].
[0426] The process of deriving juxtaposed motion vectors
[0427] The input to this process is:
[0428] - The variable currCb specifies the current encoded block.
[0429] - The variable colCb specifies the juxtaposition encoding block within the juxtaposition image specified by ColPic.
[0430] -Luminance position (xColCb, yColCb), which specifies the top-left sample of the juxtaposed luminance coding block specified by colCb relative to the top-left luminance sample of the juxtaposed image specified by ColPic.
[0431] - Refer to the index refIdxLX, where X is 0 or 1.
[0432] - The flag sbFlag indicates the candidate for time merging of sub-blocks.
[0433] The output of this process is:
[0434] Predicted motion vector mvLXCol with -1 / 16 fractional sampling accuracy.
[0435] -Availability flag availableFlagLXCol.
[0436] The variable currPic specifies the current image.
[0437] Set the arrays predFlagL0Col[x][y] and refIdxL0Col[x][y] to be equal to the PredFlagL0[x][y] and RefIdxL0[x][y] of the juxtaposed image specified by ColPic, respectively, and set the arrays predFlagLlCol[x][y] and refIdxLlCol[x][y] to be equal to the PredFlagL1[x][y] and RefIdxL1[x][y] of the juxtaposed image specified by ColPic, respectively.
[0438] Set the arrays mvL0MantissaCol[x][y], mvL0ExponentCol[x][y], mvL1MantissaCol[x][y], and mvL1ExponentCol[x][y] to be equal to the juxtaposed images MvL0Mantissa[x][y], MvL0Exponent[x][y], MvL1Manitssa[x][y], and MvL1Exponent[x][y] specified by ColPic, respectively.
[0439] The variables mvLXCol and availableFlagLXCol are exported as follows:
[0440] - If colCb is encoded in intra-prediction mode, then set both components of mvLXCol to 0 and set availableFlagLXCol to 0.
[0441] - Otherwise, the motion vector mvCol, the reference index refIdxCol, and the reference list identifier listCol are derived as follows:
[0442] - If sbFlag equals 0, then set availableFlagLXCol to 1, and the following applies:
[0443] - If predFlagL0Col[xColCb][yColCb] equals 0, then set mvCol, refIdxCol, and listCol to equal mvL1Col[xColCb][yColCb], refIdxL1Col[xColCb][yColCb], and L1, respectively.
[0444] Otherwise, if predFlagL0Col[xColCb][yColCb] equals 1 and predFlagL1Col[xColCb][yColCb] equals 0, then set mvCol, refIdxCol, and listCol to equal mvL0Col[xColCb][yColCb], refIdxL0Col[xColCb][yColCb], and L0, respectively.
[0445] - Otherwise (predFlagL0Col[xColCb][yColCb] equals 1, and predFlagL1Col[xColCb][yColCb] equals 1), specify the following:
[0446] - If NoBackwardPredFlag equals 1, then set mvCol, refIdxCol, and listCol to equal mvLXCol[xColCb][yColCb], refIdxLXCol[xColCb][yColCb], and LX, respectively.
[0447] Otherwise, set mvCol, refIdxCol, and listCol to be equal to mvLNCol[xColCb][yColCb], refldxLNCol[xColCb][yColCb], and LN, respectively, where N is the value of collocated_from_10flag.
[0448] - Otherwise (sbFlag equals 1), the following applies:
[0449] - If PredFlagLXCol[xColCb][yColCb] equals 1, then set mvCol, refldxCol, and listCol to equal mvLXCol[xColCb][yColCb], refIdxLXCol[xColCb][yColCb], and LX respectively, and set availableFlagLXCol to 1.
[0450] - Otherwise (PredFlagLXCol[xColCb][yColCb] equals 0), the following applies:
[0451] If for each image in the reference image list for the current slice, DiffPicOrderCnt(aPic, currPic) is less than or equal to 0, and PredFlagLYCol[xColCb][yColCb] is equal to 1, then set mvCol, refIdxCol, and listCol to mvLYCol[xColCb][yColCb], refIdxLYCol[xColCb][yColCb], and LY, respectively, where Y equals ! X, where X is the value of X that called the procedure. Set availableFlagLXCol to 1.
[0452] - Set both components of mvLXCol to 0, and set availableFlagLXCol to equal 0.
[0453] - When availableFlagLXCol equals TRUE, mvLXCol and availableFlagLXCol are derived as follows:
[0454] - If LongTermRefPic(currPic, currCb, refIdxLX, LX) is not equal to LongTermRefPic(ColPic, colCb, refIdxCol, listCol), then set both components of mvLXCol to 0 and set availableFlagLXCol to 0.
[0455] - Otherwise, set the variable availableFlagLXCol to equal 1, set refPicListCol[refIdxCol] to the image with reference index refIdxCol in the list of reference images listCol containing slices of coded blocks colCb in the juxtaposed images specified by ColPic, and the following applies:
[0456] colPocDiff=DiffPicOrderCnt(ColPic, refPicListCol[refIdxCol])
[0457] currPocDiff=DiffPicOrderCnt(currPic, RefPicListX[refIdxLX])
[0458] -Use mvL0MantissaCol[xColCb][yColCb] and mvL0ExponentCol[xColCb][yColCb] as input to call the procedure specified in item C for deriving the motion vector displacement from the mantissa and exponent representations, and assign the output to mvColQ.
[0459] - If RefPicListX[refIdxLX] is a long-term reference image, or colPocDiff equals currPocDiff, then mvLXCol is exported as follows:
[0460] mvLXCol=mvColQ
[0461] - Otherwise, export mvLXCol as a scaled version of the motion vector mvColQ as follows:
[0462] tx=(16384+(Abs(td)>>1)) / td
[0463] distScaleFactor=Clip3(-4096, 4095, (tb*tx+32)>>6)
[0464] mvLXCol=Clip3(-32768, 32767, Sign(distScaleFactor*mvColQ)*((Abs(distScaleFactor*mvColQ)+127)>>8))
[0465] The td and tb are exported as follows:
[0466] td=Clip3(-128, 127, colPocDiff)
[0467] tb=Clip3(-128, 127, currPocDiff)
[0468] Option B. The process of deriving the conversion of motion vector displacement into mantissa and exponent representation.
[0469] The input to this process is:
[0470] - Specify the array mvCol of input motion vectors. The output of this process is:
[0471] - Arrays representing the mantissa and exponent of the specified time motion vector: mvColMantissa, mvColExponent
[0472] The arrays mvColMantissa and mvColExponent are exported as follows:
[0473]
[0474] Option C. Used for deriving the motion vector displacement from the mantissa and exponent representations.
[0475] The input to this process is:
[0476] -Specify arrays of time motion vectors using mantissa and exponent representation, such as mvColMantissa and mvColExponent.
[0477] The output of this process is:
[0478] - Specifies the array mvColQ of time motion vectors derived from their mantissa and exponent representations.
[0479] The array mvColQ is exported as follows:
[0480]
[0481] In one example, scaling of the motion vector is skipped, in which case mvColScaled equals mvCol in item B.
[0482] In one example, according to the techniques described in this paper, the derivation of the juxtaposed motion vector can be performed as follows:
[0483] The input to this process is:
[0484] - The variable currCb specifies the current encoded block.
[0485] - The variable colCb specifies the juxtaposition encoding block within the juxtaposition image specified by ColPic.
[0486] -Luminance position (xColCb, yColCb), which specifies the top-left sample of the juxtaposed luminance coding block specified by colCb relative to the top-left luminance sample of the juxtaposed image specified by ColPic.
[0487] - Refer to the index refIdxLX, where X is 0 or 1.
[0488] - The flag sbFlag indicates the candidate for time merging of sub-blocks.
[0489] The output of this process is:
[0490] Predicted motion vector mvLXCol with -1 / 16 fractional sampling accuracy.
[0491] -Availability flag LXCoL
[0492] The variable currPic specifies the current image.
[0493] Set the arrays predFlagL0Col[x][y], mvL0Col[x][y], and refIdxL0Col[x][y] to be equal to the PredFlagL0[x][y], MvL0[x][y], and RefIdxL0[x][y] of the juxtaposed images specified by ColPic, respectively. Set the arrays predFlagL1Col[x][y], mvL1Col[x][y], and refIdxL1Col[x][y] to be equal to the PredFlagL1[x][y], MvL1[x][y], and RefIdxL1[x][y] of the juxtaposed images specified by ColPic, respectively.
[0494] The variables mvLXCol and availableFlagLXCol are exported as follows:
[0495] - If colCb is encoded in intra-frame prediction mode or its reference image is ColPic, then set both components of mvLXCol to 0 and set availableFlagLXCol to 0.
[0496] - Otherwise, the motion vector mvCol, the reference index refIdxCol, and the reference list identifier listCol are derived as follows:
[0497] - If sbFlag equals 0, then set availableFlagLXCol to 1, and the following applies:
[0498] - If predFlagL0Col[xColCb][yColCb] equals 0, then set mvCol, refIdxCol, and listCol to equal mvL1Col[xColCb][yColCb], refIdxL1Col[xColCb][yColCb], and L1, respectively.
[0499] Otherwise, if predFlagL0Col[xColCb][yColCb] equals 1 and predFlagL1Col[xColCb][yColCb] equals 0, then set mvCol, refIdxCol, and listCol to equal mvL0Col[xColCb][yColCb], refIdxL0Col[xColCb][yColCb], and L0, respectively.
[0500] - Otherwise (predFlagL0Col[xColCb][yColCb] equals 1, and predFlagL1Col[xColCb][yColCb] equals 1), specify the following:
[0501] - If NoBackwardPredFlag equals 1, then set mvCol, refIdxCol, and listCol to equal mvLXCol[xColCb][yColCb], refIdxLXCol[xColCb][yColCb], and LX, respectively.
[0502] Otherwise, set mvCol, refIdxCol, and listCol to be equal to mvLNCol[xColCb][yColCb], refIdxLNCol[xColCb][yColCb], and LN, respectively, where N is the value of collocated_from_10flag.
[0503] - Otherwise (sbFlag equals 1), the following applies:
[0504] - If PredFlagLXCol[xColCb][yColCb] equals 1, then set mvCol, refIdxCol, and listCol to equal mvLXCol[xColCb][yColCb], refIdxLXCol[xColCb][yColCb], and LX respectively, and set availableFlagLXCol to 1.
[0505] - Otherwise (PredFlagLXCol[xColCb][yColCb] equals 0), the following applies:
[0506] - If for each image in the reference image list of the current tile group, DiffPicOrderCnt(aPic, currPic) is less than or equal to 0, and PredFlagLYCol[xColCb][yColCb] is equal to 1, then set mvCol, refIdxCol, and listCol to mvLYCol[xColCb][yColCb], refIdxLYCol[xColCb][yColCb], and LY, respectively, where Y equals ! X, where X is the value of X that called the procedure. Set availableFlagLXCol to 1.
[0507] - Set both components of mvLXCol to 0, and set availableFlagLXCol to equal 0.
[0508] - When availableFlagLXCol equals TRUE, mvLXCol and availableFlagLXCol are derived as follows:
[0509] - If LongTermRefPic(currPic, currCb, refIdxLX, LX) is not equal to LongTermRefPic(ColPic, colCb, refIdxCol, listCol), then set both components of mvLXCol to 0 and set availableFlagLXCol to 0.
[0510] - Otherwise, set the variable availableFlagLXCol to equal 1, set refPicListCol[refIdxCol] to the picture with reference index refIdxCol in the listCol of reference pictures that contains the coded block colCb in the juxtaposed picture specified by ColPic, and the following applies:
[0511] colPocDiff=DiffPicOrderCnt(ColPic, refPicListCol[refIdxCol])
[0512] currPocDiff=DiffPicOrderCnt(currPic, RefPicListX[refIdxLX])
[0513] - Use mvCol as input to call the time motion buffer compression process for juxtaposing motion vectors, and assign the output to mvCol.
[0514] - If RefPicListX[refIdxLX] is a long-term reference image, or colPocDiff equals currPocDiff, then mvLXCol is exported as follows:
[0515] mvLXCol=mvCol
[0516] - Otherwise, export mvLXCol as a scaled version of the motion vector mvCol as follows:
[0517] tx=(16384+(Abs(td)>>1)) / td
[0518] distScaleFactor=Clip3(-4096, 4095, (tb*tx+32)>>6)
[0519] mvLXCol=Clip3(-32768, 32767, Sign(distScaleFactor*mvCol)*((Abs(distScaleFactor*mvCol)+127)>>8))
[0520] The td and tb are exported as follows:
[0521] td=Clip3(-128, 127, colPocDiff)
[0522] tb=Clip3(-128, 127, currPocDiff)
[0523] Compression process of time motion buffer for juxtaposing motion vectors
[0524] The input to this process is:
[0525] - Motion vector mv,
[0526] The output of this process is:
[0527] - Rounding motion vector rmv
[0528] For each motion vector component compIdx, rmv[compIdx] is derived from mv[compIdx] as follows:
[0529] s = mv[compIdx] >> 17
[0530] f=Floor(Log2((mv[compIdx]^s)|31))-4
[0531] mask = (-1 << f) >> 1
[0532] round = (1 << f) >> 2
[0533] rmv[compIdx]=(mv[compIdx]+round)&mask
[0534] It should be noted that this process enables the use of bit-simplified representation to store juxtaposed motion vectors. Each signed 18-bit motion vector component can be represented using a mantissa and exponent format with a 6-bit signed mantissa and a 4-bit exponent.
[0535] In another example, the time-motion buffer compression process used to concatenate motion vectors exhibits symmetry about 0. That is, if the input is -mv[compIdx] instead of mv[compIdx], the output is -rmv[compIdx] instead of rmv[compIdx]. Therefore, the time-motion buffer compression process used to concatenate motion vectors can be performed as follows:
[0536] Compression process of time motion buffer for juxtaposing motion vectors
[0537] The input to this process is:
[0538] - Motion vector mv,
[0539] The output of this process is:
[0540] - Rounding motion vector rmv
[0541] For each motion vector component compIdx, rmv[compIdx] is derived from mv[compIdx] as follows:
[0542] S = mv[compIdx] >> 17
[0543] F=Floor(Log2((mv[compIdx]^s)|31))-4
[0544] Mask = (-1 << f) >> 1
[0545] round=((1<<f)+s)>>2
[0546] rmv[compIdx]=(mv[compIdx]+round)&mask
[0547] In another example, the compression process of the time motion buffer used to juxtapose motion vectors can be performed as follows:
[0548] The input to this process is:
[0549] - Motion vector mv,
[0550] The output of this process is:
[0551] - Rounding motion vector rmv
[0552] For each motion vector component compIdx, rmv[compIdx] is derived from mv[compIdx] as follows:
[0553] s = mv[compIdx] >> 17
[0554] f=Floor(Log2((mv[compIdx]^s)|31))-4
[0555] mask = (-1 << f) >> 1
[0556] round=((1<<f)-1-s)>>2
[0557] rmv[compIdx]=(mv[compIdx]+round)&mask
[0558] In another example, the compression process of the time motion buffer used to juxtapose motion vectors can be performed as follows:
[0559] The input to this process is:
[0560] - Motion vector mv,
[0561] The output of this process is:
[0562] - Rounding motion vector rmv
[0563] For each motion vector component compIdx, rmv[compIdx] is derived from mv[compIdx] as follows:
[0564] s = mv[compIdx] >> 17
[0565] f=Floor(Log2((mv[compIdx]^s)|31))-4
[0566] mask = (-1 << f) >> 1
[0567] round=((1<<f)+s)>>2
[0568] rmv[compIdx]=Clip3(-(1<<17), (1<<17)-1, (mv[compIdx]+round)&mask)
[0569] In another example, the compression process of the time motion buffer used to juxtapose motion vectors can be performed as follows:
[0570] The input to this process is:
[0571] - Motion vector mv,
[0572] The output of this process is:
[0573] - Rounding motion vector rmv
[0574] For each motion vector component compIdx, rmv[compIdx] is derived from mv[compIdx] as follows:
[0575] s = mv[compIdx] >> 17
[0576] f=Floor(Log2((mv[compIdx]^s)|31))-4
[0577] mask = (-1 << f) >> 1
[0578] round=((1<<f)+s)>>2
[0579] rmv[compIdx]=Clip3(-(64<<11), (63<<11), (mv[compIdx]+round)&mask)
[0580] It should be noted that in the above case, where rmv[compIdx] = Clip3(-(64<<11), (63<<11), (mv[compIdx]+round)&mask), the following set of operations is performed:
[0581] s = mv[compIdx] >> 17
[0582] f=Floor(Log2((mv[compIdx]^s)|31))-4
[0583] mask = (-1 << f) >> 1
[0584] round=((1<<f)+s)>>2
[0585] rmv[compIdx]=Clip3(-(64<<11), (63<<11), (mv[compIdx]+round)&mask)
[0586] Equivalent to:
[0587] cmv[compIdx]=Clip3(-(64<<11), (63<<11), mv[compIdx])
[0588] s = mv[compIdx] >> 17
[0589] f=Floor(Log2((cmv[compIdx]^s)|31))-4
[0590] mask = (-1 << f) >> 1
[0591] round=((1<<f)+s)>>2
[0592] rmv[compIdx]=(cmv[compIdx]+round)&mask
[0593] Therefore, in the case of rmv[compIdx] = Clip3(-(64<<11),(63<<11),(mv[compIdx]+round)&mask), the clipping operation can be performed earlier in the export process if needed (e.g., to improve computational throughput).
[0594] Furthermore, it should be noted that, in one example, mvLXCol can be clipped in the export of the juxtaposed motion vector process as follows:
[0595] - If RefPicListX[refIdxLX] is a long-term reference image, or colPocDiff equals currPocDiff, then mvLXCol is exported as follows:
[0596] mvLXCol=Clip3(-(1<<17), (1<<17)-1, mvCo1)
[0597] - Otherwise, export mvLXCol as a scaled version of the motion vector mvCol as follows:
[0598] tx=(16384+(Abs(td)>>1)) / td
[0599] distScaleFactor=Clip3(-4096, 4095, (tb*tx+32)>>6)
[0600] mvLXCol=Clip3(-(1<<17), (1<<17)-1, Sign(distScaleFactor*mvCol)*((Abs(distScaleFactor*mvCol)+127)>>8))
[0601] The td and tb are exported as follows:
[0602] td=Clip3(-128, 127, colPocDiff)
[0603] tb=Clip3(-128, 127, currPocDiff)
[0604] Refer again Figure 6 ,like Figure 6As shown, the inter-frame prediction processing unit 214 can receive reconstructed video blocks via a filter unit 216, which can be part of an in-loop filtering process. The filter unit 216 can be configured to perform deblocking and / or Sample Adaptive Offset (SAO) filtering. Deblocking refers to the process of smoothing the boundaries of the reconstructed video blocks (e.g., making the boundaries less perceptible to the observer). SAO filtering is a nonlinear amplitude mapping that can be used to improve the reconstruction by adding an offset to the reconstructed video data. The entropy coding unit 218 receives quantized transform coefficients and prediction syntax data (i.e., intra-frame prediction data, motion prediction data, and QP data, etc.). The entropy coding unit 218 can be configured to perform entropy coding according to one or more of the techniques described herein. The entropy coding unit 218 can be configured to output a compliant bitstream (i.e., a bitstream from which the video decoder can receive and reproduce the video data). Thus, the video encoder 200 represents an example of a device configured to: determine full-precision motion vectors for generating predictions of video blocks in a first image, store the motion vectors with less than the full precision, and generate candidate predicted motion vectors from the stored motion vectors for video blocks in a second image.
[0605] Figure 7 This is a block diagram illustrating an example of a video decoder configured to decode video data according to one or more techniques described herein. In one example, the video decoder 300 may be configured to reconstruct video data based on one or more of the techniques described above. That is, the video decoder 300 may operate in a manner reversible from the video encoder 200 described above. The video decoder 300 may be configured to perform intra-frame predictive decoding and inter-frame predictive decoding, and thus may be referred to as a hybrid decoder. Figure 7 In the example shown, the video decoder 300 includes an entropy decoding unit 302, an inverse quantization unit 304, an inverse transform processing unit 306, an intra-frame prediction processing unit 308, an inter-frame prediction processing unit 310, a summer 312, a filter unit 314, and a reference buffer 316. The video decoder 300 can be configured to decode video data in a manner consistent with a video coding system that implements one or more aspects of a video coding standard. It should be noted that although the exemplary video decoder 300 is shown with different functional blocks, such illustrations are for descriptive purposes and do not limit the video decoder 300 and / or its sub-components to a particular hardware or software architecture. The functionality of the video decoder 300 can be implemented using any combination of hardware, firmware, and / or software implementations.
[0606] like Figure 7As shown, the entropy decoding unit 302 receives an entropy-encoded bitstream. The entropy decoding unit 302 can be configured to decode the quantization syntax elements and quantization coefficients from the bitstream according to a process that is the inverse of the entropy encoding process. The entropy decoding unit 302 can be configured to perform entropy decoding according to any of the entropy encoding techniques described above. The entropy decoding unit 302 can parse the encoded bitstream in a manner consistent with video coding standards. The video decoder 300 can be configured to parse the encoded bitstream, wherein the encoded bitstream is generated based on the techniques described above. The inverse quantization unit 304 receives quantization transform coefficients (i.e., bit values) and quantization parameter data from the entropy decoding unit 302. The quantization parameter data can include any and all combinations of the aforementioned increment QP values and / or quantization group size values. The video decoder 300 and / or the inverse quantization unit 304 can be configured to determine the QP value for inverse quantization based on the value sent by the video encoder with a signal and / or through video attributes and / or encoding parameters. That is, the inverse quantization unit 304 can operate in a manner that is the inverse of the aforementioned coefficient quantization unit 206. Inverse quantization unit 304 can be configured to apply inverse quantization. Inverse transform processing unit 306 can be configured to perform an inverse transform to generate reconstructed residual data. The techniques performed by inverse quantization unit 304 and inverse transform processing unit 306 respectively can be similar to those performed by inverse quantization / transform processing unit 208 described above. Inverse transform processing unit 306 can be configured to apply inverse DCT, inverse DST, inverse integer transform, indivisible quadratic transform (NSST), or conceptually similar inverse transform procedures to transform the coefficients in order to generate residual blocks in the pixel domain. Furthermore, as mentioned above, whether a specific transform (or the type of specific transform) is performed can depend on the intra-frame prediction mode. Figure 7 As shown, the reconstructed residual data can be provided to the summer 312. The summer 312 can add the reconstructed residual data to the predicted video block and generate reconstructed video data.
[0607] As described above, predicted video blocks can be determined based on predictive video techniques (i.e., intra-frame prediction and inter-frame prediction). Intra-frame prediction processing unit 308 can be configured to receive intra-frame prediction syntax elements and retrieve predicted video blocks from reference buffer 316. Reference buffer 316 may include a memory device configured to store one or more video data frames. Intra-frame prediction syntax elements can identify intra-frame prediction modes, such as those described above. In one example, intra-frame prediction processing unit 308 may use one or more of the intra-frame prediction coding techniques described herein to reconstruct video blocks. Inter-frame prediction processing unit 310 can receive inter-frame prediction syntax elements and generate motion vectors to identify predicted blocks in one or more reference frames stored in reference buffer 316. Inter-frame prediction processing unit 310 can generate motion-compensated blocks, possibly performing interpolation based on interpolation filters. Identifiers for interpolation filters used for motion estimation with sub-pixel precision can be included in the syntax elements. Inter-frame prediction processing unit 310 may use interpolation filters to compute interpolated values for sub-integer pixels of the reference blocks.
[0608] As described above, the video decoder 300 can parse an encoded bitstream generated based on the aforementioned technique, and the video encoder 200 can generate the bitstream according to the aforementioned motion vector prediction technique. Therefore, the video decoder 300 can be configured to perform motion vector prediction according to the aforementioned technique. Thus, the video decoder 300 represents an example of a device configured to: determine full-precision motion vectors for generating predictions of video blocks in a first image, store the motion vectors with less than the full precision, and generate candidate predicted motion vectors from the stored motion vectors for video blocks in a second image.
[0609] Refer again Figure 7 Filter unit 314 can be configured to perform filtering on the reconstructed video data. For example, filter unit 314 can be configured to perform deblocking and / or SAO filtering, as described above with respect to filter unit 216. Furthermore, it should be noted that in some examples, filter unit 314 can be configured to perform dedicated arbitrary filtering (e.g., visual enhancement). Figure 7 As shown, the video block can be reconstructed by the output of the video decoder 300.
[0610] In one or more examples, the functionality may be implemented by hardware, software, firmware, or any combination thereof. If implemented in software, the functionality may be stored as one or more instructions or code on or transmitted over a computer-readable medium and executed by a hardware-based processing unit. A computer-readable medium may include a computer-readable storage medium corresponding to a tangible medium such as a data storage medium, or a propagation medium that includes, for example, any medium facilitating the transfer of a computer program from one place to another according to a communication protocol. Thus, a computer-readable medium may generally correspond to: (1) a non-transitory tangible computer-readable storage medium, or (2) a communication medium such as a signal or carrier wave. A data storage medium may be any available medium that can be accessed by one or more computers or one or more processors to retrieve instructions, code, and / or data structures for implementing the techniques described in this disclosure. Computer program products may include computer-readable media.
[0611] By way of example, and not limitation, such computer-readable storage media may include RAM, ROM, EEPROM, CD-ROM or other optical disc storage devices, disk storage devices or other magnetic storage devices, flash memory, or any other medium that can be used to store required program code in the form of instructions or data structures and that is accessible by a computer. Furthermore, any connection is appropriately referred to as a computer-readable medium. For example, coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of medium if instructions are transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave. However, it should be understood that computer-readable storage media and data storage media do not include connections, carrier waves, signals, or other transient media, but rather refer to non-transient tangible storage media. As used herein, disks and optical discs include compact optical discs (CDs), laser discs, optical discs, digital versatile optical discs (DVDs), floppy disks, and Blu-ray discs, where disks typically copy data magnetically, while optical discs use lasers to copy data optically. Combinations of the above should also be included within the scope of computer-readable media.
[0612] Instructions can be executed by one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Therefore, the term "processor" as used herein can refer to any of the foregoing structures or any other structure suitable for implementing the techniques described herein. Furthermore, in some aspects, the functionality described herein can be provided within dedicated hardware and / or software modules configured for encoding and decoding, or incorporated into combined codecs. Moreover, these techniques can be implemented entirely within one or more circuit or logic elements.
[0613] The techniques disclosed herein can be implemented in various devices or apparatuses, including wireless mobile phones, integrated circuits (ICs), or a set of ICs (e.g., chipsets). Various components, modules, or units are described in this disclosure to emphasize functional aspects of a device configured to perform the disclosed techniques, but they do not necessarily need to be implemented through different hardware units. Rather, as described above, various units can be combined in a codec hardware unit, or provided through an interoperable hardware unit comprising a collection of one or more processors as described above, combined with suitable software and / or firmware.
[0614] Furthermore, each functional block or feature of the base station equipment and terminal equipment used in each of the above embodiments can be implemented or executed by circuitry (typically one or more integrated circuits). Circuitry designed to perform the functions described in this specification can include general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components, or combinations thereof. The general-purpose processor can be a microprocessor, or alternatively, it can be a conventional processor, controller, microcontroller, or state machine. The general-purpose processor or each of the above circuitry can be configured by digital circuitry or by analog circuitry. Furthermore, when advancements in semiconductor technology lead to the development of technologies for manufacturing integrated circuits that replace current integrated circuits, integrated circuits produced using such technologies can also be used.
[0615] Various examples have been described. These and other examples are within the scope of the following claims.
[0616] <Cross-reference>
[0617] This non-provisional application claims priority under 35 U.S.C., section 119, for the following provisional applications: 62 / 768,772, November 16, 2018; 62 / 787,695, January 2, 2019; 62 / 792,872, January 15, 2019; 62 / 793,080, January 16, 2019; 62 / 793,311, January 16, 2019; and 62 / 815,109, March 7, 2019; the entire contents of which are incorporated herein by reference.
Claims
1. A method for performing motion vector prediction for encoding / decoding video data, the method comprising: Determine the motion vector mv; Determine the rounding motion vector, which is equal to the value of rmv in the following equation: s = mv >> 17; f = Floor(Log2((mv ^ s) | 31)) – 4; mask = (-1 << f) >> 1; round = (1 << f) >> 2; rmv = (mv + round) & mask; and Candidate predictive motion vectors are generated based on the rounded motion vectors.
2. The method according to claim 1, wherein, The motion vector mv is exported as a motion vector corresponding to the juxtaposed luminance coding block within the specified juxtaposed image.
3. The method according to claim 1, wherein, The predicted motion vector candidates are generated based on the rounded motion vectors by cropping them to values in the range of -131072 to 131071.
4. An apparatus comprising one or more processors, said one or more processors being configured to: Determine the motion vector mv; Determine the rounding motion vector, which is equal to the value of rmv in the following equation: s = mv >> 17; f = Floor(Log2((mv ^ s) | 31)) – 4; mask = (-1 << f) >> 1; round = (1 << f) >> 2; rmv = (mv + round) & mask; and Candidate predictive motion vectors are generated based on the rounded motion vectors.
Citation Information
Patent Citations
Inter-view residual prediction in multi-view or 3-dimensional video coding
EP2965511A1
Method and apparatus for encoding / decoding images using a motion vector
US20130294522A1