Selective use of alternative interpolation filters in video processing
By selectively using alternative interpolation filters and the merge mode of motion vector difference (MMVD), the video encoding and decoding process is optimized, solving the problem of low efficiency in existing technologies and achieving more efficient video processing and resource utilization.
Patent Information
- Application Number
- CN202080059204.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-08-20
- Filing Date
- 2020-08-20
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2040-08-20
AI Technical Summary
Existing video codec standards suffer from inefficiency and resource waste when processing video data, especially in motion vector derivation and interpolation filter selection, leading to increased bandwidth requirements and higher codec complexity.
By employing a merge mode (MMVD) that selectively uses alternative interpolation filters and motion vector differences, the appropriate interpolation filters and motion vector differences are selected to optimize the video processing flow by determining the suitability between the current video block and the codec representation.
It improves video encoding and decoding efficiency, reduces bandwidth requirements and encoding/decoding complexity, and enhances video processing performance and resource utilization.
Smart Images

Figure CN114270856B_ABST
Abstract
Description
[0001] Cross-references to related applications
[0002] This is the national phase of International Patent Application No. PCT / CN2020 / 110147, filed on August 20, 2020, which claims priority and benefit to International Patent Application No. PCT / CN2019 / 101541, filed on August 20, 2019. For all purposes under that law, the entire disclosure of the foregoing application is incorporated herein by reference as part of the disclosure of this application. Technical Field
[0003] This patent document relates to video processing technologies, equipment, and systems. Background Technology
[0004] Despite advancements in video compression, digital video still accounts for the largest share of bandwidth usage on the internet and other digital communication networks. As the number of connected user devices capable of receiving and displaying video increases, the bandwidth demand for digital video is expected to continue to grow. Summary of the Invention
[0005] Devices, systems, and methods relating to digital video processing, and particularly to motion vector derivation, are described. The methods can be applied to existing video codec standards (e.g., High Efficiency Video Codec (HEVC) or Multi-Functional Video Codec) as well as future video codec standards or codecs.
[0006] In one representative aspect, the disclosed techniques can be used to provide a method for video processing. This method includes: a conversion between a current video block of a current frame of a video and a codec representation of the video; determining the suitability of candidate interpolation filters, wherein the suitability of the candidate interpolation filters indicates whether to apply the candidate interpolation filters in the conversion; and performing the conversion based on the determination; wherein the suitability of the candidate interpolation filters is determined based on whether to use reference image resampling, in which a reference image of the current frame is resampled to perform the conversion.
[0007] In another representative aspect, the disclosed techniques can be used to provide a method for video processing. This method includes: determining a codec mode for representing a current video block as employing a merge mode with motion vector difference (MMVD), which includes a motion vector representation providing information about the distance between motion candidates and starting points; and performing a conversion between the current video block and the codec representation based on this determination, wherein the conversion is performed using a prediction block of the current video block, the prediction block being computed using a half-pixel interpolation filter selected according to a first rule and a second rule, the first rule defining a first condition that the half-pixel interpolation filter is an alternative half-pixel interpolation filter different from a default half-pixel interpolation filter, and the second rule defining a second condition regarding whether to inherit the alternative half-pixel interpolation filter.
[0008] In another representative aspect, the disclosed techniques can be used to provide a method for video processing. This method includes: determining a codec mode used in a current video region of the video; determining, based on the codec mode, the precision of the motion vectors or motion vector differences used to represent the current video region; and performing a conversion between the current video block and the codec representation of the video.
[0009] In another representative aspect, the disclosed techniques can be used to provide a method for video processing. This method includes: determining the codec mode of a current video block of a video to be a merge mode employing motion vector difference (MMVD); for the current video block, determining a distance table relating a specified distance index to a predefined offset based on motion precision information of basic merge candidates associated with the current video block; and using the distance table to perform a conversion between the current video block and the codec representation of the video.
[0010] In another representative aspect, the disclosed technique can be used to provide a method for video processing. This method includes: making a first determination: the motion vector difference used in a decoder-side motion vector refinement (DMVR) calculation for a first video region has a finer resolution than X pixel resolution, and the motion vector difference is determined using an alternative half-pixel interpolation filter, where X is an integer fraction; making a second determination based on the first determination: either not storing information about the alternative half-pixel interpolation filter associated with the first video region or making such information unusable for a second video region subsequently processed; and performing a conversion between the video including the first and second video regions and a codec representation of the video based on the first and second determinations.
[0011] In another representative aspect, the disclosed techniques can be used to provide a method for video processing. This method includes: determining, according to a rule, whether an alternative half-pixel interpolation filter is available for a current video block; and performing a conversion between the current video block and a codec representation of the video based on the determination; wherein the rule specifies that, in the case of encoding / decoding the current video block as a merge block or skipping a codec block into the codec representation, the determination is independent of whether the alternative half-pixel interpolation filter was used to process a previous block encoded or decoded before the current video block.
[0012] In another representative aspect, the disclosed techniques can be used to provide a method for video processing. This method includes: determining, according to rules, whether alternative half-pixel interpolation filters are available for a current video block; and performing a conversion between the current video block and the codec representation of the video based on the determination, wherein the rules specify the applicability of alternative half-pixel interpolation filters, different from the default interpolation filter, based on the codec information of the current video block.
[0013] In another representative aspect, the disclosed techniques can be used to provide a method for video processing. This method includes: determining coefficients of candidate half-pixel interpolation filters for a current video block according to rules; and performing a conversion between the current video block and the codec representation of the video based on the determination, wherein the rules specify the relationship between the candidate half-pixel interpolation filters and the half-pixel interpolation filters used in a certain codec mode.
[0014] In another representative aspect, the disclosed techniques can be used to provide a method for video processing. This method includes: applying an alternative half-pixel interpolation filter according to a rule for a conversion between a current video block and the codec representation of the video; and performing the conversion between the current video block and the codec representation of the video, wherein the rule specifies that the alternative interpolation filter is applied to a position at pixel X, where X is different from 1 / 2.
[0015] In another representative aspect, the disclosed technique can be used to provide a method for video processing. This method includes: making a first determination: selecting an alternative interpolation filter for a first motion vector component having an accuracy of X pixels from a first video block of the video; making a second determination due to the first determination: applying another alternative interpolation filter for a second motion vector component having an accuracy different from X pixels, where X is an integer fraction; and performing a conversion between the video including the first and second video blocks and a codec representation of the video.
[0016] Furthermore, in one representative aspect, any of the disclosed methods is a codec-side implementation.
[0017] Furthermore, in one representative aspect, any of the disclosed methods is a decoder-side implementation.
[0018] One of the above methods is embodied in the form of processor executable code and stored in a computer-readable program medium.
[0019] In another representative aspect, an apparatus for a video system is disclosed, comprising a processor and a non-transitory memory having instructions thereon. When executed by the processor, these instructions cause the processor to perform the disclosed methods.
[0020] The above and other aspects and features of the disclosed technology are described in more detail in the accompanying drawings, description and claims. Attached Figure Description
[0021] Figure 1A and Figure 1B An example of a quadtree plus binary tree (QTBT) block structure is shown.
[0022] Figure 2 An example of constructing a merge candidate list is shown.
[0023] Figure 3 An example of a candidate location for the airspace is shown.
[0024] Figure 4 An example of a candidate pair that has undergone a redundancy check for spatial merge candidates is shown.
[0025] Figure 5A and Figure 5B An example of the position of a second prediction unit (PU) based on the size and shape of the current block is shown.
[0026] Figure 6 An example of motion vector scaling for temporal merge candidates is shown.
[0027] Figure 7 An example of candidate positions for temporal merge candidates is shown.
[0028] Figure 8 An example of creating a combined bidirectional prediction merge candidate is shown.
[0029] Figure 9 An example of constructing motion vector prediction candidates is shown.
[0030] Figure 10 An example of motion vector scaling for spatial motion vector candidates is shown.
[0031] Figure 11A and Figure 11BThis is a block diagram of an example hardware platform used to implement the visual media decoding or visual media encoding described in this document.
[0032] Figure 12 A flowchart of an example method for video processing is shown.
[0033] Figure 13 An example of MMVD search is shown.
[0034] Figures 14A to 14I A flowchart illustrating an exemplary method for video processing is shown. Detailed Implementation
[0035] 1. Video encoding and decoding in HEVC / H.265
[0036] Video codec standards have primarily evolved through the development of well-known ITU-T and ISO / IEC standards. ITU-T developed H.261 and H.263, while ISO / IEC developed MPEG-1 and MPEG-4 Vision. The two organizations jointly developed the H.262 / MPEG-2 Video, H.264 / MPEG-4 Advanced Video Codec (AVC), and H.265 / HEVC standards. Since H.262, video codec standards have been based on a hybrid video codec architecture, employing temporal prediction plus transform coding. To explore future video codec technologies beyond HEVC, VCEG and MPEG jointly established the Joint Video Exploration Team (JVET) in 2015. Since then, JVET has adopted many new methods and applied them to reference software called the Joint Exploration Model (JEM). In April 2018, a Joint Video Experts Team (JVET) was created between VCEG (Q6 / 16) and ISO / IEC JTC1 SC29 / WG11 (MPEG) to research a VVC standard with a target bitrate reduction of 50% compared to HEVC.
[0037] 2.1. Quadtree plus binary tree (QTBT) block structure with a large CTU
[0038] In HEVC, the CTU is divided into CUs using a quadtree structure (represented as an encoder-decoder tree) to accommodate various local characteristics. At the CU level, it is determined whether to use inter-image (temporal) prediction or intra-image (spatial) prediction to encode and decode image regions. Depending on the PU partitioning type, each CU can be further divided into one, two, or four PUs. Within a PU, the same prediction process is applied, and relevant information is transmitted to the decoder based on the PU. After obtaining residual blocks by applying the prediction process based on the PU partitioning type, the CU can be partitioned into transform units (TUs) according to another quadtree structure similar to the CU's encoder-decoder tree. A key feature of the HEVC structure is that it has multiple partitioning concepts, including CUs, PUs, and TUs.
[0039] Figure 1 illustrates an example of a Quadtree Plus Binary Tree (QTBT) block structure. The QTBT structure eliminates the concept of multiple partition types; that is, it eliminates the separation of the concepts of CU, PU, and TU, and supports greater flexibility in the shape of CU partitions. In the QTBT block structure, CUs can be square or rectangular. As shown in Figure 1, the codec tree unit (CTU) is first partitioned using a quadtree structure. The leaf nodes of the quadtree are then further partitioned using a binary tree structure. There are two types of partitioning in the binary tree: symmetrical horizontal partitioning and symmetrical vertical partitioning. The leaf nodes of the binary tree are called codec units (CUs), and this partition is used for prediction and transform processing without further partitioning. This means that in the QTBT codec block structure, CUs, PUs, and TUs have the same block size. In JEM, a CU is sometimes composed of codec blocks (CBs) of different color components. For example, in the case of P-strips and B-strips of 4:2:0 chroma format, a CU contains one luma CB and two chroma CBs. And a CU is sometimes composed of a CB of a single component. For example, in the case of I-strips, a CU contains only one luma CB or only two chroma CBs.
[0040] The following parameters are defined for the QTBT segmentation scheme.
[0041] –CTU size: The size of the root node of the quadtree, the same concept as in HEVC.
[0042] –MinQTSize: Minimum allowed size of quadtree leaf nodes
[0043] –MaxBTSize: The maximum allowed size of the root node of the binary tree.
[0044] –MaxBTDepth: Maximum allowed binary tree depth
[0045] –MinBTSize: Minimum allowed size of binary leaf nodes
[0046] In one example of a QTBT segmentation structure, the CTU size is set to 128×128 luminance samples with two corresponding 64×64 chroma sample blocks, the MinQTSize is set to 16×16, the MaxBTSize to 64×64, the MinBTSize (for both width and height) is set to 4×4, and the MaxBTDepth is set to 4. First, a quadtree segmentation is applied to this CTU to generate quadtree leaf nodes. Quadtree leaf nodes can have sizes ranging from 16×16 (i.e., MinQTSize) to 128×128 (i.e., the CTU size). If a leaf quadtree node is 128×128, it will not be further segmented by a binary tree because its size exceeds MaxBTSize (i.e., 64×64). Otherwise, the leaf quadtree node can be further segmented by a binary tree. Therefore, the quadtree leaf node is also the root node of the binary tree, and its binary tree depth is 0. When the binary tree depth reaches MaxBTDepth (i.e., 4), further partitioning is not considered. When a binary tree node has a width equal to MinBTSize (i.e., 4), further horizontal partitioning is not considered. Similarly, when a binary tree node has a height equal to MinBTSize, further vertical partitioning is not considered. Leaf nodes of the binary tree are further processed through prediction and transformation, without further segmentation. In JEM, the maximum CTU size is 256×256 luminance samples.
[0047] Figure 1A An example of block partitioning using QTBT is shown. Figure 1B The corresponding tree representation is shown. Solid lines indicate quadtree partitions, and dashed lines indicate binary tree partitions. In each partition (i.e., non-leaf) node of a binary tree, a signaling flag indicates which partition type (i.e., horizontal or vertical) was used, where 0 indicates a horizontal partition and 1 indicates a vertical partition. For quadtree partitions, it is not necessary to indicate the partition type because quadtree partitions always divide blocks both horizontally and vertically to produce four sub-blocks of equal size.
[0048] Furthermore, the QTBT scheme supports the ability to have separate QTBT structures for luma and chroma. Currently, for P-strip and B-strip, the luma CTB and chroma CTB within a CTU share the same QTBT structure. However, for I-strip, the luma CTB is divided into CUs using one QTBT structure, and the chroma CTB is divided into chroma CUs using another QTBT structure. This means that the CUs in I-strip consist of codec blocks for the luma component or codec blocks for the two chroma components, while the CUs in P-strip or B-strip consist of codec blocks for all three color components.
[0049] In HEVC, inter-frame prediction for small blocks is restricted to reduce memory accesses for motion compensation, thus bidirectional prediction is not supported for 4×8 and 8×4 blocks, and inter-frame prediction is not supported for 4×4 blocks. These restrictions are removed in JEM's QTBT.
[0050] 2.2. Inter-frame prediction in HEVC / H.265
[0051] Each inter-frame prediction PU has motion parameters for one or two lists of reference images. The motion parameters include motion vectors and reference image indices. The use of one of the two reference image lists can also be notified using inter_pred_idc signaling. The motion vectors can be explicitly encoded as increments relative to the predictor.
[0052] When encoding and decoding a CU using the skip mode, a PU is associated with that CU, and there are no significant residual coefficients, no encoded motion vector increments, or reference picture indices. A merge mode is specified, where the motion parameters of the current PU are obtained from neighboring PUs, including spatial and temporal candidates. The merge mode can be applied to any inter-frame prediction PU, not just the skip mode. An alternative to the merge mode is explicit transmission of motion parameters, where explicit signaling is performed for each PU for motion vectors (more precisely, the motion vector difference relative to the motion vector predictors), the corresponding reference picture index for each reference picture list, and the use of the reference picture list. In this disclosure, such a mode is referred to as Advanced Motion Vector Prediction (AMVP).
[0053] When the signaling indicates that one of these two lists of reference images will be used, a PU is generated from a sample block. This practice is called "one-way prediction". One-way prediction is available for both P-strips and B-strips.
[0054] When signaling indicates that both reference image lists should be used, a PU is generated from the two sample blocks. This practice is called "bidirectional prediction." Bidirectional prediction is only available for B-strips.
[0055] The following section provides details about these inter-frame prediction modes specified in HEVC. The description will begin with the merge mode.
[0056] 2.2.1. Merge Mode
[0057] 2.2.1.1. Derivation of candidate merge patterns
[0058] When predicting PU using merge mode, the indices pointing to entries in the merge candidate list are parsed from the bitstream, and these indices are used to retrieve motion information. The construction of this list is specified in the HEVC standard and can be summarized according to the following sequence of steps:
[0059] Step 1: Initial Candidate Derivation
[0060] Step 1.1: Spatial Candidate Derivation
[0061] Step 1.2: Redundancy check of airspace candidates
[0062] Step 1.3: Time-domain candidate derivation
[0063] Step 2: Adding candidate insertions
[0064] Step 2.1: Creation of bidirectional prediction candidates
[0065] Step 2.2: Insertion of zero-motion candidates
[0066] Still Figure 2 These steps are illustrated schematically. For spatial merge candidate derivation, up to four merge candidates are selected from candidates located at five different positions. For temporal merge candidate derivation, up to one merge candidate is selected from two candidates. Since the number of candidates is assumed to be constant for each PU at the decoder, additional candidates are generated when the number of candidates obtained from step 1 does not reach the maximum number of merge candidates (MaxNumMergeCand) signaled in the stripe header. Because the number of candidates is constant, the index encoding and decoding of the best merge candidate is performed using truncated unary code binarization (TU). If the size of the CU is equal to 8, all PUs of the current CU share a single merge candidate list, which is equivalent to the merge candidate list of a 2N×2N prediction unit.
[0067] The operations associated with the foregoing steps will be described in detail below.
[0068] 2.2.1.2. Derivation of Airspace Candidates
[0069] In the derivation of spatial merge candidates, in the location Figure 3At most four merge candidates are selected from the candidates at the indicated positions. The derivation order is A1, B1, B0, A0, and B2. Position B2 is considered only if any PU at positions A1, B1, B0, or A0 is unavailable (e.g., because it belongs to another stripe or slice) or if it is an intra-frame encoding / decoding operation. After the candidate at position A1 is added, the addition of the remaining candidates is subject to a redundancy check, which ensures that candidates with the same motion information are excluded from the list, thereby improving encoding / decoding efficiency. To reduce computational complexity, not all possible candidate pairs are considered in the mentioned redundancy check. Instead, only those with... Figure 4 The arrows link the pairs, and a candidate is added to the list only if the corresponding candidate used for redundancy checking does not have the same motion information. Another source of duplicate motion information is a “second PU” associated with a segmentation different from 2N×2N. As an example, Figure 5 shows the second PUs for N×2N and 2N×N cases, respectively. When the current PU is segmented into N×2N, the candidate at position A1 is not considered for list construction. In fact, adding this candidate would cause two prediction units to have the same motion information, which is redundant for ensuring that only one PU within the encoder-decoder unit is redundant. Similarly, when the current PU is segmented into 2N×N, position B1 is not considered.
[0070] 2.2.1.3. Derivation of Time-Domain Candidates
[0071] In this step, only one candidate is added to the list. Specifically, in the derivation of this temporal merge candidate, a scaled motion vector is derived based on the co-located PU, which belongs to the image with the smallest POC difference between the current image and a given list of reference images. The list of reference images to be used for the derivation of this co-located PU is explicitly signaled within the strip header. The scaled motion vector used for the temporal merge candidate is as follows: Figure 6 The dashed lines in the diagram show the results obtained by scaling the motion vectors of the co-located PU using the POC distances (i.e., tb and td). Here, tb is defined as the POC distance between the current image and its reference image, and td is defined as the POC distance between the reference image and its co-located image. The reference image index for the temporal merge candidate is set to zero. The actual implementation of the scaling process is described in the HEVC specification. For B-strips, two motion vectors are obtained (one for reference image list 0 and the other for reference image list 1), and they are combined to generate bidirectional predictive merge candidates.
[0072] In the co-located PU(Y) belonging to the reference frame, the position of the temporal candidate is selected between candidate C0 and C1, such as... Figure 7As shown. If the PU at position C0 is unavailable, is intra-frame encoded, or is outside the current CTU line, then position C1 is used. Otherwise, position C0 is used in the derivation of the temporal merge candidate.
[0073] 2.2.1.4. Additional Candidate Insertion
[0074] In addition to spatial and temporal merge candidates, there are two additional types of merge candidates: combined bidirectional prediction merge candidates and zero merge candidates. Combined bidirectional prediction merge candidates are generated using both spatial and temporal merge candidates. Combined bidirectional prediction merge candidates are only used for B-strips. Combined bidirectional prediction candidates are generated by combining the motion parameters of the first reference image list of the initial candidate with the motion parameters of the second reference image list of the other. If these two tuples provide different motion hypotheses, they will form a new bidirectional prediction candidate. As an example, Figure 8 This illustrates the case where two candidates from the original list (left side) (which have mvL0 and refIdxL0 or mvL1 and refIdxL1) are used to create combined bidirectional predictive merge candidates that are added to the final list (right side). There are many rules regarding the combinations considered for generating these additional merge candidates.
[0075] Zero-motion candidates are inserted to fill the remaining entries in the merge candidate list, thus reaching the MaxNumMergeCand capacity. These candidates have zero spatial displacement and a reference image index that starts at zero and increases each time a new zero-motion candidate is added to the list. For unidirectional and bidirectional prediction, these candidates use one and two reference frames, respectively. Finally, no redundancy checks are performed on these candidates.
[0076] 2.2.1.5. Motion estimation region for parallel processing
[0077] To accelerate the encoding process, motion estimation can be performed in parallel, thereby simultaneously deriving the motion vectors of all prediction units within a given region. Deriving merge candidates from spatial neighbors can interfere with parallel processing because a prediction unit cannot derive motion parameters from neighboring PUs until its associated motion estimation is complete. To mitigate the trade-off between encoding / decoding efficiency and processing latency, HEVC defines a Motion Estimation Region (MER), whose size is signaled in the image parameter set using the "log2_parallel_merge_level_minus2" syntax element. When defining the MER, merge candidates falling within the same region are marked as unavailable and therefore not considered when constructing the list.
[0078] 2.2.2.AMVP
[0079] AMVP utilizes the spatiotemporal correlation between motion vectors and adjacent PUs, which is used for the explicit transmission of motion parameters. For each reference image list, the motion vector candidate list is constructed as follows: first, the availability of temporally adjacent PU positions on the left and top sides is checked, redundant candidates are removed, and zero vectors are added to ensure a constant length for the candidate list. Then, the encoder can select the best predictor from the candidate list and transmit the corresponding index indicating the selected candidate. Similar to merge index signaling, a truncated unary code is used to encode the index of the best motion vector candidate. In this case, the maximum value to be encoded is 2 (reference...). Figure 9 The following sections will provide details on the derivation process of the motion vector prediction candidates.
[0080] 2.2.2.1. Derivation of AMVP Candidates
[0081] Figure 9 The derivation process of motion vector prediction candidates is summarized.
[0082] In motion vector prediction, two types of motion vector candidates are considered: spatial motion vector candidates and temporal motion vector candidates. For the derivation of spatial motion vector candidates, based on the location as... Figure 3 The motion vectors of each PU at the five different positions shown are used to derive two candidate motion vectors.
[0083] For temporal motion vector candidate derivation, one motion vector candidate is selected from two candidates derived based on two different co-locations. After creating a first list of space-time candidates, duplicate motion vector candidates in this list are removed. If the number of possible candidates is greater than two, motion vector candidates whose reference image index in the associated reference image list is greater than 1 are removed from the list. If the number of space-time motion vector candidates is less than two, additional zero motion vector candidates are added to the list.
[0084] 2.2.2.2. Candidate Spatial Motion Vectors
[0085] In the derivation of the spatial motion vector candidates, at most two candidates are considered from five possible candidates. These five possible candidates are selected from those located in the spatial domain, such as... Figure 3The positions of the PUs shown are derived from these positions, which are the same as those positions merged in the motion. The derivation order for the left side of the current PU is defined as A0, A1, and scaled A0, scaled A1. The derivation order for the upper side of the current PU is defined as B0, B1, B2, scaled B0, scaled B1, scaled B2. For each side, there are four cases that can be used as motion vector candidates, two of which are not associated with spatial scaling, and two of which use spatial scaling. These four different cases are summarized as follows:
[0086] • No spatial scaling
[0087] –(1) List of identical reference images and index of identical reference images (identical POCs)
[0088] –(2) Different lists of reference images, but with the same reference image (same POC)
[0089] • Spatial scaling
[0090] –(3) Same list of reference images, but different reference image indices (different POCs)
[0091] –(4) List of different reference images and different reference images (different POCs)
[0092] First, check for cases without spatial scaling, then proceed with spatial scaling. Spatial scaling is considered when the POC differs between the reference image of the adjacent PU and the reference image of the current PU, regardless of the reference image list. Scaling for the upper motion vector is allowed if all candidate PUs on the left are unavailable or intra-frame encoded / decoded, thus facilitating parallel derivation of the left and upper MV candidates. Otherwise, spatial scaling for the upper motion vector is not allowed.
[0093] During spatial scaling, the motion vectors of adjacent PUs are scaled in a manner similar to temporal scaling, such as... Figure 10 As shown. The main difference is that the current PU's reference image list and index are given as input; the actual scaling process is the same as temporal scaling.
[0094] 2.2.2.3. Candidate Motion Vectors in the Temporal Domain
[0095] Aside from the derivation of the reference image index, the entire derivation process of the temporal merge candidate is identical to the derivation of the spatial motion vector candidate (see [link to derivation]). Figure 7 (Same as above.) The reference image index signaling is sent to the decoder.
[0096] 2.3. Adaptive Motion Vector Differential Resolution (AMVR)
[0097] In VVC, for regular inter-frame modes, MVD encoding and decoding can be performed in units of quarter-lumen, whole-lumen, and four-lumen samples. MVD resolution is controlled at the codec unit (CU) level, and an MVD resolution flag is conditionally signaled for each CU with at least one non-zero MVD component.
[0098] For a CU with at least one non-zero MVD component, signaling notifies a first flag to indicate whether quarter-luminance sample MV accuracy is used in the CU. When the first flag (equal to 1) indicates that quarter-luminance sample MV accuracy is not used, signaling notifies another flag to indicate whether whole-luminance sample MV accuracy or four-luminance sample MV accuracy is used.
[0099] When the first MVD resolution flag of the CU is zero, or when the CU is not encoded or decoded (meaning all MVDs within the CU are zero), a quarter-luminance sample MV resolution is used for the CU. When the CU uses integer luminance sample MV precision or four-luminance sample MV precision, the MVPs in the CU's AMVP candidate list are rounded relative to the corresponding precision.
[0100] 2.4. Interpolation Filters in VVC
[0101] For luminance interpolation filtering, an 8-tap divisible interpolation filter is used for 1 / 16 pixel precision samples, as shown in Table 1.
[0102] Table 1: 8-tap coefficients f used for 1 / 16 pixel brightness interpolation L
[0103]
[0104] Similarly, a 4-tap divisible interpolation filter is used for chroma samples with 1 / 32 pixel precision, as shown in Table 2.
[0105] Table 2: 4-tap interpolation coefficients f used for 1 / 32 pixel chroma interpolation C
[0106]
[0107]
[0108] For the 4:2:2 vertical interpolation and the 4:4:4 chroma channel horizontal and vertical interpolation, the odd positions in Table 2 are not used, resulting in 1 / 16 pixel chroma interpolation.
[0109] 2.5. Alternative Luminance Half-Pixel Interpolation Filter
[0110] In JVET-N0309, an alternative half-pixel interpolation filter was proposed.
[0111] The switching of the half-pixel interpolation filter depends on the motion vector accuracy. In addition to the existing quarter-pixel, full-pixel, and 4-pixel AMVR modes, a new half-pixel accuracy AMVR mode has been introduced. Only with half-pixel motion vector accuracy can the alternative half-pixel luminance interpolation filter be selected.
[0112] 2.5.1. Half-pixel AMVR mode
[0113] An additional AMVR mode for non-affine, non-merge inter-frame codec CUs is proposed, which allows motion vector differences to be notified with half-pixel precision signaling. The existing AMVR scheme in the current VVC draft is explicitly extended as follows: immediately following the syntax element `amvr_flag`, if `amvr_flag == 1`, there exists a new context model binary syntax element `hpel_amvr_flag`, which, if `hpel_amvr_flag == 1`, indicates the use of the new half-pixel AMVR mode. Otherwise, i.e., `hpel_amvr_flag == 0`, the choice between integer-pixel and 4-pixel AMVR modes is indicated by the syntax element `amvr_precision_flag`, as in the current VVC draft.
[0114] 2.5.2. Alternative Luminance Half-Pixel Interpolation Filter
[0115] For non-affine non-merge inter-frame codecs (CUs) using half-pixel motion vector accuracy (i.e., half-pixel AMVR mode), switching between the HEVC / VVC half-pixel luma interpolation filter and one or more alternative half-pixel interpolations is made based on the value of the new syntax element if_idx. Signaling notification of the syntax element if_idx is only required in half-pixel AMVR mode. For skip / merge modes using spatial merging candidates, the value of the syntax element if_idx is inherited from adjacent blocks.
[0116] 2.5.2.1. Test 1: An alternative half-pixel interpolation filter
[0117] In this test case, a 6-tap interpolation filter exists as an alternative to the standard HEVC / VVC half-pixel interpolation filter. The following diagram illustrates the mapping between the value of the syntax element `if_idx` and the selected half-pixel luma interpolation filter:
[0118]
[0119] 2.5.2.2. Test 2: Two Alternative Half-Pixel Interpolation Filters
[0120] In this test case, two 8-tap interpolation filters exist as alternatives to the standard HEVC / VVC half-pixel interpolation filters. The following diagram illustrates the mapping between the value of the syntax element `if_idx` and the selected half-pixel luma interpolation filter:
[0121]
[0122]
[0123] The signaling notifies amvr_precision_idx whether the current CU uses 1 / 2 pixel MV precision, 1 pixel MV precision, or 4 pixel MV precision. There are two binary bits (bins) to be encoded and decoded.
[0124] The signaling notifies hpel_if_idx whether to use the default half-pixel interpolation filter or an alternative half-pixel interpolation filter. When using two alternative half-pixel interpolation filters, there are two bits to encode and decode.
[0125] 2.6. Generalized Two-Way Forecasting
[0126] In conventional bidirectional prediction, the predictors from L0 and L1 are averaged to generate the final predictor with an equal weight of 0.5. The predictor generation formula is shown in equation (3).
[0127] P TraditionalBiPred =(P L0 +P L1 +RoundingOffset)>>shiftNum, (1)
[0128] In equation (3), P TraditionalBiPred It is the final predictor used for conventional bidirectional prediction, P L0 and P L1 These are the predictors from L0 and L1 respectively, and RoundingOffset and shiftNum are used to normalize the final predictor.
[0129] A generalized bidirectional prediction (GBI) is proposed to allow different weights to be applied to predictors from L0 and L1. Predictor generation is shown in equation (4).
[0130] P GBi =((1-w1)*P L0 +w1*P L1 +RoundingOffset GBi >>shiftNum GBi (2)
[0131] In equation (4), PGBi This is the final predictor for GBi. (1-w1) and w1 are the selected GBi weights applied to the predictors for L0 and L1, respectively. RoundingOffset GBi and shiftNum GBi Used to normalize the final predictors in GBi.
[0132] The support weights for w1 are {-1 / 4, 3 / 8, 1 / 2, 5 / 8, 5 / 4}. It supports one equal-weight set and four unequal-weight sets. For the equal-weight case, the process of generating the final predictor is exactly the same as in the regular bidirectional prediction mode. For the true bidirectional prediction case under random access (RA) conditions, the number of candidate weight sets is reduced to three.
[0133] For Advanced Motion Vector Prediction (AMVP) mode, if the CU is a bidirectional predictive codec, the weight selection in the GBI is explicitly signaled at the CU level. For merge mode, the weight selection is inherited from the merge candidate.
[0134] 2.7. Employing the MVD merge pattern (MMVD)
[0135] In addition to the merge mode, which directly uses implicitly derived motion information to generate prediction samples for the current CU, a merge mode using motion vector difference (MMVD) is introduced in VVC. The MMVD flag is signaled immediately after the skip flag and merge flag are sent to indicate whether the MMVD mode is used for the CU.
[0136] In MMVD, after selecting a merge candidate, the MVD information notified via signaling is used to further refine it. This additional information includes merge candidate flags, an index specifying the motion amplitude, and an index indicating the motion direction. In MMVD mode, the first two candidates in the merge list are selected as the MV basis. The merge candidate flags are signaled to specify which one to use.
[0137] Figure 13 An example of an MMVD search is shown. The distance index specifies motion amplitude information and indicates a predefined offset from the starting point. For example... Figure 13 As shown, the offset is added to the horizontal or vertical component of the starting MV. Table 3 specifies the relationship between the distance index and the predefined offset.
[0138] Table 3. Relationship between distance index and predefined offset
[0139]
[0140] The direction index indicates the direction of the MVD relative to the starting point. The direction index can represent the four directions shown in Table 4. It should be noted that the meaning of the MVD symbol can vary depending on the information of the starting MV. When the starting MV is a unidirectional prediction MV or a bidirectional prediction MV where both lists point to the same side of the current image (i.e., both reference POCs are greater than or less than the current image's POC), the symbols in Table 4 specify the sign of the MV offset to be added to the starting MV. When the starting MV is a bidirectional prediction MV where the two MVs point to different sides of the current image (i.e., one reference POC is greater than the current image's POC, and the other reference POC is less than the current image's POC), the symbols in Table 4 specify the sign of the MV offset added to the list 0 MV component of the starting MV, and the signs for list 1 MVs have the opposite value.
[0141] Table 4. Sign of MV Offset Specifyed by Direction Index
[0142] Directional IDX 00 01 10 11 x-axis + – N / A N / A y-axis N / A N / A + –
[0143] The slice_fpel_mmvd_enabled_flag can be signaled in the slice header to indicate whether fractional distances can be used in MMVD. If fractional distances are not allowed for slices, the above distances can be multiplied by 4 to generate the distance table below.
[0144] Table 5. Relationship between distance index and predefined offset when slice_fpel_mmvd_enabled_flag equals 1.
[0145]
[0146] A slice_fpel_mmvd_enabled_flag value of 1 specifies that the Merge mode using motion vector difference uses integer sample precision in the current slice. A slice_fpel_mmvd_enabled_flag value of 0 specifies that the Merge mode using motion vector difference can use fractional sample precision in the current slice. When the value of slice_fpel_mmvd_enabled_flag does not exist, it is inferred to be 0.
[0147] 3. Problems in conventional implementation methods
[0148] It is unreasonable that alternative half-pixel interpolation filters might be inherited in the merge (MMVD) mode which uses motion vector difference, even though the MV derived in the MMVD mode does not have 1 / 2 pixel accuracy.
[0149] When encoding and decoding amvr_precision_idx and hpel_if_idx, all binary bit contexts are encoded and decoded.
[0150] If an associated merge candidate derived from a spatially neighboring block is selected, the candidate interpolation filter flag may be inherited from the spatially neighboring block. However, when further refining the inherited MV (e.g., using DMVR), the final MV may not point to half-pixel precision. Therefore, even if the flag is true, the candidate interpolation filter is still disabled because the use of a new filter depends on the flag being true and the MV being half-pixel.
[0151] Similarly, during AMVR, if the AMVR precision is half-pixel, then both MVP and MVD conform to half-pixel accuracy. However, the final MV may not be half-pixel. In this case, even if the AMVR precision indicated by the signaling specifies half-pixel accuracy, the use of a new filter is still not permitted.
[0152] 4. Exemplary embodiments and technologies
[0153] The embodiments described below should be considered as examples to illustrate general principles. These embodiments should not be interpreted narrowly. Furthermore, these embodiments can be combined in any way.
[0154] The decoder-side motion vector derivation (DMVD) is used to represent BDOF (bidirectional optical flow) and / or DMVR (decoder-side motion vector refinement) and / or other tools at the decoder associated with refined motion vectors or prediction samples.
[0155] In the following text, the default interpolation filter may refer to an interpolation filter defined in HEVC / VVC. Newly introduced interpolation filters (e.g., those proposed in JVET-N0309) may also be referred to as alternative interpolation filters in the following description. The term MMVD can refer to any encoding / decoding method that employs some additional MVD to update the decoded motion information.
[0156] 1. Alternative 1 / N pixel interpolation filters can be used for different N values, where N is not equal to 2.
[0157] a. In one example, N can be equal to 4, 16, etc.
[0158] b. In one example, the index of the 1 / N pixel interpolation filter can be signaled in AMVR mode.
[0159] i. Alternatively, in addition, the index of the 1 / N pixel interpolation filter is signaled only when the block selects 1 / N pixel MV / MVD precision.
[0160] c. In one example, the alternative 1 / N pixel interpolation filter may not be inherited in merge mode or / and MMVD mode.
[0161] i. Alternatively, in addition, the default 1 / N pixel interpolation filter can be used only in merge mode and / or MMVD mode.
[0162] d. In one example, alternative 1 / N pixel interpolation filters can be inherited in merge mode.
[0163] e. In one example, an alternative 1 / N pixel interpolation filter can be inherited in MMVD mode.
[0164] i. In one example, when the final derived MV has 1 / N pixel precision, i.e. no MV component has finer MV precision, the alternative 1 / N pixel interpolation filter can be inherited in the MMVD mode.
[0165] ii. In one example, when the final derived MV has K (K>=1) MV components with 1 / N pixel precision, the alternative 1 / N pixel interpolation filter can be inherited in MMVD mode.
[0166] f. In one example, an alternative 1 / N pixel interpolation filter can be inherited in MMVD mode; however, the alternative 1 / N pixel interpolation filter is only used for motion compensation. The index of the alternative 1 / N pixel interpolation filter may not be stored for this block, and this index cannot be used by subsequent codec blocks.
[0167] 2. Indices for interpolation filters (e.g., default half-pixel interpolation filter, alternative half-pixel interpolation filters) can be stored together with other motion information (such as motion vectors, reference indices).
[0168] a. In one example, for a block to be encoded / decoded, when it accesses a second block located in a different region (such as in a different CTU line, in a different VPDU), the interpolation filter associated with that second block is not allowed to be used to encode / decode the current block.
[0169] 3. In one example, alternative half-pixel interpolation filters may not be inherited in merge mode and / or MMVD mode.
[0170] a. In one example, alternative half-pixel interpolation filters may not be inherited in MMVD mode.
[0171] i. Alternatively, the default half-pixel interpolation filter in VVC can always be used in MMVD mode.
[0172] ii. Alternatively, alternative half-pixel interpolation filters can be inherited in MMVD mode. That is, for MMVD mode, alternative half-pixel interpolation filters associated with the base merge candidate can be inherited.
[0173] b. In one example, alternative half-pixel interpolation filters can be inherited under certain conditions in MMVD mode.
[0174] i. In one example, when the final derived MV has 1 / 2 pixel precision, that is, when no MV component has a finer MV precision (such as 1 / 4 pixel precision, 1 / 16 pixel precision), the alternative half-pixel interpolation filter can be inherited.
[0175] ii. In one example, when the final derived MV has K (K>=1) MV components with 1 / 2 pixel precision, the alternative half-pixel interpolation filter can be inherited in MMVD mode.
[0176] iii. In one example, when the selected distance in the MMVD (e.g., the distance defined in Table 3) has an accuracy of X pixels or a coarser accuracy than X pixels (e.g., X is 1 / 2, and 1 pixel, 2 pixels, 4 pixels, etc. are coarser than X pixels), the alternative half-pixel interpolation filter can be inherited.
[0177] iv. In one example, when the selected distance (e.g., the distance defined in Table 3) has an accuracy of X pixels or a finer accuracy than X pixels (e.g., X is 1 / 4 and 1 / 16 pixels is a finer accuracy than X pixels), the alternative half-pixel interpolation filter may not be inherited.
[0178] v. When an alternative half-pixel interpolation filter is not used in MMVD, the information about the alternative half-pixel interpolation filter may not be stored in the block, and this information may not be used by subsequent codec blocks.
[0179] c. In one example, an alternative half-pixel interpolation filter can be inherited in MMVD mode and / or merge mode; however, the alternative half-pixel interpolation filter is only used for motion compensation. The block can store the index of the default half-pixel interpolation filter instead of the index of the alternative half-pixel interpolation filter, and this index of the default half-pixel interpolation filter can be used by subsequent codec blocks.
[0180] d. The above method can be applied to other situations where multiple interpolation filters with 1 / N pixel precision can be used.
[0181] 4. MV / MVD precision information can be stored for CU / PU / blocks encoded and decoded in AMVP mode and / or affine inter-frame mode, and this MV / MVD precision information can be inherited by CU / PU / blocks encoded and decoded in merge mode and / or affine merge mode.
[0182] a. In one example, for candidates derived from spatially adjacent (adjacent or non-adjacent) blocks, the MV / MVD accuracy of the associated neighboring blocks can be inherited.
[0183] b. In one example, for a pair of merge candidates, if two associated spatial merge candidates have the same MV / MVD precision, such MV / MVD precision can be assigned to the pair of merge candidates; otherwise, a fixed MV / MVD precision (e.g., 1 / 4 pixel or 1 / 16 pixel) can be assigned.
[0184] c. In one example, a fixed MV / MVD precision (e.g., 1 / 4 pixel or 1 / 16 pixel) can be assigned to temporal merge candidates.
[0185] d. In one example, MV / MVD accuracy can be stored in an HMVP (History-Based Motion Vector Prediction) table, and it can be inherited by HMVP merge candidates.
[0186] e. In one example, the MV / MVD precision of the basic merge candidate can be inherited by the CU / PU / block encoded and decoded according to the MMVD mode.
[0187] f. In one example, inherited MV / MVD accuracy can be used to predict the MV / MVD accuracy of subsequent blocks.
[0188] g. In one example, the MV of the merged codec block can be rounded to the inherited MV / MVD precision.
[0189] 5. Which distance table will be used in MMVD can depend on the accuracy of the MV / MVD of the base merge candidate.
[0190] a. In one example, if the MV / MVD of the basic merge candidate has 1 / N pixel precision, then the distance used in MMVD should have 1 / N pixel precision or a coarser precision than 1 / N pixels.
[0191] i. For example, if N=2, the distance in MMVD can have only 1 / 2 pixel precision, 1 pixel precision, 2 pixel precision, etc.
[0192] b. In one example, the distance table including fractional distances (defined in Table 3) can be modified based on the MV / MVD accuracy of the base merge candidate, for example, the distance table defined when slice_fpel_mmvd_enabled_flag equals 0.
[0193] i. For example, if the MV / MVD precision of the basic merge candidate is 1 / N pixels, and the finest MVD precision of the distance table is 1 / M pixels (e.g., M=4), then all distances in the distance table can be multiplied by M / N.
[0194] 6. For a CU / PU / block that selects an alternative half-pixel interpolation filter (e.g., encoded and decoded in merge mode), when the MVD derived in the DMVR has a finer precision than X (e.g., X = 1 / 2) pixels, the information of the alternative half-pixel interpolation filter may not be stored in the CU / PU / block, and this information may not be used by subsequent blocks.
[0195] a. In one example, if the DMVR is performed at the sub-block level, then the decision on whether to store information about alternative half-pixel interpolation filters is made independently for each sub-block based on the MVD accuracy of the sub-block derived in the DMVR.
[0196] b. In one example, if the derived MVD has a finer precision than X pixels for at least N (e.g., N=1) sub-blocks / blocks, then the information of the alternative half-pixel interpolation filters may not be stored, and that information may not be used for subsequent blocks.
[0197] 7. Interpolation filter information can be stored in a history-based motion vector prediction (HMVP) table, and this information can be inherited by HMVP merge candidates.
[0198] a. In one example, when inserting a new candidate into the HMVP lookup table, interpolation filter information can be considered. For instance, two candidates with the same motion information but different interpolation filter information can be considered as two different candidates.
[0199] b. In one example, when inserting a new candidate into the HMVP lookup table, two candidates with the same motion information but different interpolation filter information can be considered as the same candidate.
[0200] 8. When inserting merge candidates into the merge candidate list, interpolation filter information can be considered during the pruning process.
[0201] a. In one example, two candidates with different interpolation filters can be considered as two different merge candidates.
[0202] b. In one example, when inserting HMVP merge candidates into the merge list, interpolation filter information can be considered during the pruning process.
[0203] c. In one example, when inserting HMVP merge candidates into the merge list, interpolation filter information can be disregarded during the pruning process.
[0204] 9. When generating paired merge candidates and / or combined merge candidates and / or zero motion vector candidates and / or other default candidates, interpolation filter information can be considered instead of always using the default interpolation filter.
[0205] a. In one example, if both candidates (involved in generating pairwise merge candidates and / or combined merge candidates) use the same alternative interpolation filter, then such interpolation filter can be inherited in the pairwise merge candidate and / or combined merge candidate.
[0206] b. In one example, if one of the two candidates (involved in generating pairwise merge candidates and / or combined merge candidates) does not use the default interpolation filter, then its interpolation filter can be inherited in that pairwise merge candidate and / or combined merge candidate.
[0207] c. In one example, if one of the two candidates (involved in generating the combined merge candidate) does not use the default interpolation filter, its interpolation filter can be inherited in the combined merge candidate. However, such an interpolation filter can be used only for the corresponding prediction direction.
[0208] d. In one example, if two candidates (involved in generating the combined merge candidate) use different interpolation filters, their interpolation filters can be inherited in the combined merge candidate. In this case, different interpolation filters can be used for different prediction directions.
[0209] e. In one example, no more than K (K>=0) pairs of merge candidates or / and combined merge candidates can use alternative interpolation filters.
[0210] f. In one example, the default interpolation filter is always used for pairwise merge candidates and / or combined merge candidates.
[0211] 10. It is proposed to prohibit the use of half-pixel motion vector / motion vector difference precision when encoding and decoding the current block using IBC mode.
[0212] a. Alternatively, there is no need for signaling notification of instructions regarding the use of half-pixel MV / MVD precision.
[0213] b. In one example, if the current block is encoded and decoded in IBC mode, the alternative half-pixel interpolation filter is always disabled.
[0214] c. Alternatively, there is no need to signal the indication of the half-pixel interpolation filter.
[0215] d. In one example, the condition "the current block is encoded in IBC mode" can be replaced by "the current block is encoded in a certain mode". Such a mode can be defined as triangle mode, merge mode, etc.
[0216] 11. When encoding amvr_precision_idx and / or hpel_if_idx, only the first binary bit is context-coded.
[0217] a. Alternatively, other binary bits can be bypassed for encoding and decoding.
[0218] b. In one example, the first binary bit of amvr_precision_idx can be bypassed for encoding and decoding.
[0219] c. In one example, the first binary bit of hpel_if_idx can be bypassed for encoding and decoding.
[0220] d. In one example, the first binary bit of amvr_precision_idx can be encoded and decoded using only one context.
[0221] e. In one example, the first binary bit of hpel_if_idx can be encoded and decoded using only one context.
[0222] f. In one example, all bits of amvr_precision_idx can share the same context.
[0223] g. In one example, all bits of hpel_if_idx can share the same context.
[0224] 12. When using alternative interpolation filters, some encoding / decoding tools may not be allowed.
[0225] a. In one example, bidirectional optical flow (BDOF) may be disallowed when using an alternative interpolation filter.
[0226] b. In one example, when using alternative interpolation filters, DMVR and / or DMVD may not be allowed.
[0227] c. In one example, CIIP (Combined Inter-Frame Intra-Frame Prediction) may not be allowed when using alternative interpolation filters.
[0228] i. In one example, when merging candidate inheritance alternative interpolation filters, the CIIP flag can be skipped and inferred as false.
[0229] ii. Alternatively, when the CIIP flag is true, the default interpolation filter can always be used.
[0230] d. In one example, SMVD (Symmetric Motion Vector Difference) may not be allowed when using an alternative interpolation filter.
[0231] i. In one example, when using SMVD, the default interpolation filter can always be used, and the syntax elements related to alternative interpolation filters can be notified without signaling.
[0232] ii. Alternatively, when the syntax element related to the alternative interpolation filter indicates the use of the alternative interpolation filter, the SMVD-related syntax element may not be signaled and the SMVD mode may not be used.
[0233] e. In one example, SBT (subblock transform) may not be allowed when using alternative interpolation filters.
[0234] i. In one example, when using SBT, the default interpolation filter can always be used, and no signaling notification is required for syntax elements related to alternative interpolation filters.
[0235] ii. Alternatively, when a syntax element related to an alternative interpolation filter indicates the use of an alternative interpolation filter, the relevant syntax element of the SBT may not be signaled and the SBT may not be used.
[0236] f. In one example, when using an alternative interpolation filter, triangular prediction may not be allowed.
[0237] i. In one example, interpolation filter information may not be inherited in triangle prediction, and only the default interpolation filter can be used.
[0238] g. Alternatively, triangular prediction can be allowed when an alternative interpolation filter is used.
[0239] i. In one example, interpolation filter information can be inherited in triangle prediction.
[0240] h. Alternatively, for the encoding / decoding tools mentioned above, if they are enabled, the alternative half-pixel interpolation filter can be disabled.
[0241] 13. Filters can be applied to MV with N-pixel precision.
[0242] a. In one example, N can be equal to 1, 2, or 4, etc.
[0243] b. In one example, the filter could be a low-pass filter.
[0244] c. In one example, the filter could be a 1-d (one-dimensional) filter.
[0245] i. For example, the filter could be a 1-d level filter.
[0246] ii. For example, the filter could be a 1-d vertical filter.
[0247] d. In one example, a signaling flag can be used to indicate whether such a filter is employed.
[0248] i. Alternatively, in addition, a signaling notification flag may be used only when N-pixel MVD precision (signaling notification in AMVR mode) is used for a block.
[0249] 14. Different sets of weighting factors can be used for regular inter-frame mode and affine mode in GBI mode.
[0250] a. In one example, signaling notifications for the weighted factor set used in regular inter-frame mode and affine mode can be provided in SPS / group header / strip header / VPS / PPS, etc.
[0251] b. In one example, the set of weighting factors used for regular inter-frame mode and affine mode can be predefined at the codec and decoder.
[0252] 15. How to define / select alternative interpolation filters can depend on the encoding / decoding mode information.
[0253] a. In one example, the set of allowed alternative interpolation filters can be different for affine and non-affine modes.
[0254] b. In one example, the set of allowed alternative interpolation filters can be different for IBC mode and non-IBC mode.
[0255] 16. In one example, whether an alternative half-pixel interpolation filter is applied in a merge or skipped codec block is independent of whether an alternative half-pixel interpolation filter is applied in an adjacent block.
[0256] a. In one example, hpel_if_idx is not stored in any block.
[0257] b. In one example, whether to inherit or discard the alternative half-pixel interpolation filter may depend on the codec information and / or on the enabling / disabling of other codec tools.
[0258] 17. In one example, whether to apply an alternative half-pixel interpolation filter (e.g., in a merge or skipped codec block) may depend on the codec information of the current block.
[0259] a. In one example, if the merge index is equal to K, then an alternative half-pixel interpolation filter is applied in the merge or skip codec block, where K is an integer such as 2, 3, 4, 5.
[0260] b. In one example, if the merge index Idx satisfies Idx%S equals K, then an alternative half-pixel interpolation filter is applied in the merge or skip codec block, where S and K are integers. For example, S equals 2 and K equals 1.
[0261] c. In one example, a candidate half-pixel interpolation filter can be applied to a specific merge candidate. For example, a candidate half-pixel interpolation filter can be applied to a pair of merge candidates.
[0262] 18. In one example, the alternative half-pixel interpolation filter can be made consistent with the half-pixel interpolation filter used in affine inter-frame prediction.
[0263] a. For example, uniform interpolation filter coefficients could be [3, -11, 40, 40, -11, 3].
[0264] b. For example, the uniform interpolation filter coefficients could be [3,9,20,20,9,3]. 19. In one example, the alternative half-pixel interpolation filter could be the same as the half-pixel interpolation filter used in chroma inter-frame prediction.
[0265] a. For example, uniform interpolation filter coefficients can be [-4, 36, 36, -4].
[0266] 20. In one example, alternative interpolation filters can be applied not only at half-pixel locations.
[0267] a. For example, if it is explicitly or implicitly indicated that an alternative interpolation filter will be applied (e.g., hpel_if_idx equals 1), and MV relates to a position at pixel X, where X is not equal to 1 / 2 (e.g., 1 / 4 or 3 / 4), then an alternative interpolation filter for the position X should be applied. The alternative interpolation filter should have at least one coefficient that is different from the original interpolation filter for the position X.
[0268] b. In one example, apply the above bullet points after decoder-side motion vector refinement (DMVR) or MMVD or SMVD or other decoder-side motion derivation processes.
[0269] c. In one example, if reference image resampling (RPR) is used, apply the bullet points above if the reference image has a different resolution than the current image.
[0270] 21. In the case of bidirectional prediction, an alternative interpolation filter can be applied to only one prediction direction.
[0271] 22. When a block selects an alternative interpolation filter for an MV component with a precision of X pixels (e.g., X = 1 / 2), if there are MV components with a precision other than X pixels (e.g., 1 / 4, 1 / 16), then other alternative interpolation filters can be used instead of the default interpolation filter for such MV components.
[0272] a. In one example, if a Gaussian interpolation filter is selected for the MV component with an accuracy of X pixels, the Gaussian interpolation filter can also be used for the MV component with an accuracy of more than X pixels.
[0273] b. In one example, if a Flattop interpolation filter is selected for the MV component with an accuracy of X pixels, the Flattop interpolation filter can also be used for the MV component with an accuracy of more than X pixels.
[0274] c. In one example, such a constraint can be applied only to certain MV precisions other than the X pixel precision.
[0275] i. For example, such constraints can be applied only to MV components that have a finer precision than X pixels.
[0276] ii. For example, such constraints can be applied only to MV components that have a coarser precision than X pixels.
[0277] 23. Whether and / or how to apply alternative interpolation filters may depend on whether RPR is used.
[0278] a. In one example, if RPR is used, the alternative interpolation filter should not be used.
[0279] b. In one example, if the reference image of the current block has a different resolution compared to the current image, then alternative interpolation filters (e.g., 1 / 2 pixel interpolation filters) should not be used.
[0280] 5. Example of an implementation plan
[0281] An example for bullet point 3 is shown based on JVET-O2001-vE. Newly added sections are highlighted by bold, underlined text.
[0282] 8.5.2.2. Derivation of the Luminance Motion Vector for Merge Mode
[0283] This procedure is called only when general_merge_flag[xCb][yCb] equals 1, where (xCb, yCb) specifies the top-left sample of the current luminance codec block relative to the top-left luminance sample of the current image.
[0284] The input to this process is:
[0285] – The brightness position (xCb, yCb) of the current brightness codec block relative to the current top-left brightness sample of the current image.
[0286] – The variable cbWidth specifies the width of the current codec block in units of luminance samples.
[0287] – The variable cbHeight specifies the height of the current codec block in luminance samples. The output of this process is:
[0288] – Brightness motion vectors mvL0[0][0] and mvL1[0][0] with 1 / 16 fractional sample point accuracy,
[0289] –Refer to indices refIdxL0 and refIdxL1,
[0290] – The prediction list uses the flags predFlagL0[0][0] and predFlagL1[0][0].
[0291] – Half-sample interpolation filter index hpelIfIdx,
[0292] – Bidirectional prediction weight index bcwIdx.
[0293] –merging candidate list mergeCandList
[0294] The bidirectional prediction weight index bcwIdx is set to equal to 0.
[0295] The motion vectors mvL0[0][0] and mvL1[0][0], the reference indices refIdxL0 and refIdxL1, and the prediction utilization flags predFlagL0[0][0] and predFlagL1[0][0] can be derived by following the ordered steps below.
[0296] 1. With the luma codec block position (xCb, yCb), luma codec block width cbWidth, and luma codec block height cbHeight as input, invoke the derivation process for spatial merging candidates from adjacent codec units specified in Item 8.5.2.3, outputting availability flags availableFlagA0, availableFlagA1, availableFlagB0, availableFlagB1, and availableFlagB2, and referencing indices refIdxLXA0, refIdxLXA1, refIdxLXB0, refIdxLXB1, and refIdxLXB2. The prediction list utilizes the flags predFlagLXA0, predFlagLXA1, predFlagLXB0, predFlagLXB1, and predFlagLXB2, as well as the motion vectors mvLXA0, mvLXA1, mvLXB0, mvLXB1, and mvLXB2, where X is 0 or 1; the half-sample interpolation filter indices hpelIfIdxA0, hpelIfIdxA1, hpelIfIdxB0, hpelIfIdxB1, and hpelIfIdxB2; and the bidirectional prediction weight indices bcwIdxA0, bcwIdxA1, bcwIdxB0, bcwIdxB1, and bcwIdxB2.
[0297] 2. Set the reference index refIdxLXCol (where X is 0 or 1) and the bidirectional prediction weight index bcwIdxCol of the temporal merging candidate Col to equal 0, and set hpelIfIdxCol to equal 0.
[0298] 3. With the luminance position (xCb, yCb), luminance block width cbWidth, luminance block height cbHeight, and variable refIdxL0Col as input, invoke the derivation process for predicting the temporal luminance motion vector specified in clause 8.5.2.11. The outputs are the availability flag availableFlagL0Col and the temporal motion vector mvL0Col. The variables availableFlagCol, predFlagL0Col, and predFlagL1Col are derived as follows:
[0299] availableFlagCol=availableFlagL0Col (8-303)
[0300] predFlagL0Col=availableFlagL0Col (8-304)
[0301] predFlagL1Col=0 (8-305)
[0302] 4. When slice_type equals B, with the luminance position (xCb, yCb), luminance codec block width cbWidth, luminance codec block height cbHeight, and variable refIdxL1Col as input, invoke the derivation process for predicting the temporal luminance motion vector specified in Item 8.5.2.11. Its output is the availability flag availableFlagL1Col and the temporal motion vector mvL1Col. The variables availableFlagCol and predFlagL1Col are derived as follows:
[0303] availableFlagCol=availableFlagL0Col||availableFlagL1Col(8-306)predFlagL1Col=availableFlagL1Col (8-307)
[0304] 5. Construct the merging candidate list mergeCandList as follows:
[0305]
[0306] 6. Set the variables numCurrMergeCand and numOrigMergeCand to be equal to the number of merging candidates in mergeCandList.
[0307] 7. When numCurrMergeCand is less than (MaxNumMergeCand-1) and NumHmvpCand is greater than 0, the following applies:
[0308] – Invoke the derivation process for history-based merging candidates specified in 8.5.2.6 with mergeCandList and numCurrMergeCand as inputs and modified mergeCandList and numCurrMergeCand as outputs.
[0309] – Set numOrigMergeCand to be equal to numCurrMergeCand.
[0310] 8. When numCurrMergeCand is less than MaxNumMergeCand and greater than 1, the following applies:
[0311] – The derivation process for pairwise averaging merging candidates specified in Clause 8.5.2.4 is invoked with the following inputs: the reference indices refIdxL0N and refIdxL1N for each candidate N in mergeCandList, the prediction list using flags predFlagL0N and predFlagL1N, the motion vectors mvL0N and mvL1N, the half-sample interpolation filter index hpelIfIdxN, and numCurrMergeCand. The output is then assigned to mergeCandList, numCurrMergeCand, the reference indices refIdxL0avgCand and refIdxL1avgCand of the candidate avgCands added to mergeCandList, the prediction list using flags predFlagL0avgCand and predFlagL1avgCand, and the motion vectors mvL0avgCand and mvL1avgCand. The bidirectional prediction weight index bcwIdx of the candidate avgCand to be added to mergeCandList is set to equal to 0.
[0312] – Set numOrigMergeCand to be equal to numCurrMergeCand.
[0313] 9. With mergeCandList; the reference indices refIdxL0N and refIdxL1N for each candidate N in mergeCandList, the prediction list using flags predFlagL0N and predFlagL1N, the motion vectors mvL0N and mvL1N; and numCurrMergeCand as input, invoke the derivation procedure for merging candidates for zero motion vectors specified in Item 8.5.2.5, and assign the output to mergeCandList, numCurrMergeCand, and each new candidate zeroCand added to mergeCandList. m Reference index refIdxL0zeroCand m and refIdxL1zeroCand m The prediction list uses the flags predFlagL0zeroCand m and predFlagL1zeroCand m and motion vector mvL0zeroCand m and mvL1zeroCand m Each new zeroCand candidate will be added to mergeCandList. mThe half-sample interpolation filter index hpelIfIdx is set to equal to 0. Each new candidate zeroCand will be added to mergeCandList. m The bidirectional prediction weight index bcwIdx is set to equal to 0. The number of candidates added, numZeroMergeCand, is set to equal to (numCurrMergeCand - numOrigMergeCand). When numZeroMergeCand is greater than 0, m is in the range of 0 to numZeroMergeCand-1 (inclusive of endpoints).
[0314] 10. When N is a candidate at position merge_idx[xCb][yCb] in the merging candidate list mergeCandList (N = mergeCandList[merge_idx[xCb][yCb]]) and X is replaced by 0 or 1, make the following assignment:
[0315] refIdxLX=refIdxLXN (8-309)
[0316] predFlagLX[0][0]=predFlagLXN (8-310)
[0317] mvLX[0][0][0]=mvLXN[0] (8-311)
[0318] mvLX[0][0][1]=mvLXN[1] (8-312)
[0319] hpelIfIdx=hpelIfIdxN (8-313)
[0320] bcwIdx=bcwIdxN (8-314)
[0321] 11. When mmvd_merge_flag[xCb][yCb] equals 1, the following applies:
[0322] – The derivation process for merging motion vector differences specified in 8.5.2.7 is invoked with the brightness position (xCb, yCb), reference index refIdxL0, refIdxL1, and prediction list as inputs using flags predFlagL0[0][0] and predFlagL1[0][0] and outputs motion vector differences mMvdL0 and mMvdL1.
[0323] – The motion vector difference mMvdLX is added to the merge motion vector mvLX (where X is 0 and 1) as follows:
[0324] mvLX[0][0][0]+=mMvdLX[0] (8-315)
[0325] mvLX[0][0][1]+=mMvdLX[1] (8-316)
[0326] mvLX[0][0][0]=Clip3(-2 17 ,2 17 -1,mvLX[0][0][0]) (8-317)
[0327] mvLX[0][0][1]=Clip3(-2 17 ,2 17 -1,mvLX[0][0][1]) (8-318)
[0328] – If any of the following conditions are true, then hpelIfIdx will be set to zero:
[0329] 1. mMvdL0[0]%8 is not equal to zero.
[0330] 2. mMvdL1[0]%8 is not equal to zero.
[0331] 3. mMvdL0[1]%8 is not equal to zero.
[0332] 4. mMvdL1[1]%8 is not equal to zero.
[0333] 6. Exemplary Implementations of the Disclosed Technology
[0334] Figure 11A This is a block diagram of a video processing apparatus 1100. Apparatus 1100 can be used to implement one or more of the methods described herein. Apparatus 1100 can be embodied in smartphones, tablets, computers, Internet of Things (IoT) receivers, etc. Apparatus 1100 may include one or more processors 1102, one or more memories 1104, and video processing hardware 1106. The processors 1102(s) may be configured to implement one or more methods described herein. The memories 1104(s) may be used to store data and code for implementing the methods and techniques described herein. The video processing hardware 1106 may be used to implement some of the techniques described herein in hardware circuitry and may be part or entirely part of the processor 1102 (e.g., a graphics processing unit (GPU) core or other signal processing circuitry).
[0335] Figure 11B This is another example of a block diagram of a video processing system in which the disclosed technology can be implemented. Figure 11BThis is a block diagram illustrating an exemplary video processing system 1150 in which various techniques disclosed herein may be implemented. Various implementations may include some or all of the components of system 1150. System 1150 may include an input 1152 for receiving video content. The video content may be received in a raw or uncompressed format, such as 8-bit or 10-bit multi-component pixel values, or may have a compressed or encoded format. Input 1152 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet, Passive Optical Network (PON), and wireless interfaces such as Wi-Fi or cellular interfaces.
[0336] System 1150 may include codec component 1154, which may implement the various codec or encoding methods described herein. Codec component 1154 may reduce the average bit rate of the video from input 1152 to the output of codec component 1154 to produce a codec representation of the video. Therefore, codec techniques are sometimes referred to as video compression or video transcoding techniques. The output of codec component 1154 may be stored or transmitted via connected communication, as represented by component 1156. The stored or transmitted bitstream (or codec) representation of the video received at input 1152 may be used by component 1158 to generate pixel values or sent to displayable video at display interface 1160. The process of generating user-viewable video from the bitstream is sometimes referred to as video decompression. Furthermore, although some video processing operations are referred to as “codec” operations or tools, it should be understood that codec tools or operations are used at the encoder, and the corresponding decoding tools or operations that reverse the codec results will be performed by the decoder.
[0337] Examples of peripheral bus interfaces or display interfaces may include Universal Serial Bus (USB), High Definition Multimedia Interface (HDMI), or DisplayPort. Examples of storage interfaces include SATA (Serial Advanced Technology Accessory), PCI, IDE, etc. The technologies described in this document can be found in a variety of electronic devices, such as mobile phones, laptops, smartphones, or other devices capable of performing digital data processing and / or video display.
[0338] In some embodiments, it can be used in relation to Figure 11A The video processing method discussed in this patent document may be implemented using a device implemented on the hardware platform described in 11B.
[0339] Some embodiments of the disclosed technology include making a decision or determination to enable a video processing tool or mode. In one example, when a video processing tool or mode is enabled, the encoder will use or implement the tool or mode in the processing of video blocks, but not necessarily modify the resulting bitstream based on the use of the tool or mode. That is, when a video processing tool or mode is enabled based on a decision or determination, the conversion from video blocks to a video bitstream will use that video processing tool or mode. In another example, when a video processing tool or mode is enabled, the decoder will process the bitstream knowing that it has been modified based on the video processing tool or mode. That is, the conversion from the video bitstream to video blocks will be performed using the video processing tool or mode enabled based on a decision or determination.
[0340] Some embodiments of the disclosed technology include making a decision or determination to disable a video processing tool or mode. In one example, when a video processing tool or mode is disabled, the encoder will not use the tool or mode in the bitstream of converting video blocks into video. In another example, when a video processing tool or mode is disabled, the decoder will process the bitstream knowing that no modifications have been made to the bitstream using a video processing tool or mode that was disabled based on the decision or determination.
[0341] In this document, the term "video processing" can refer to video encoding, video decoding, video compression, or video decompression. For example, a video compression algorithm may be applied during the conversion from the pixel representation of a video to the corresponding bitstream, or vice versa. The bitstream of the current video block may, for example, correspond to bits located in one place or scattered in different places within the bitstream, as defined by the syntax. For example, a macroblock may be encoded based on the error residual values after transformation and encoding / decoding, and also using bits in the header and other fields of the bitstream.
[0342] It should be understood that the disclosed methods and techniques will be beneficial to video encoder and / or decoder embodiments incorporated within a video processing device by allowing the use of the techniques disclosed in this document, such as a smartphone, laptop, desktop computer, and similar device.
[0343] Figure 12 This is a flowchart of an exemplary video processing method 2000. Method 1200 includes, at 1210, determining a single set of motion information for a current video block based on applying at least one interpolation filter to a set of adjacent blocks, wherein the at least one interpolation filter can be configured to integer pixel accuracy or subpixel accuracy of the single set of motion information. Method 1200 further includes, at 1220, performing a conversion between the current video block and a bitstream of the current video block, wherein the conversion includes a decoder motion vector refinement (DMVR) step for refining the single set of motion information signaled in the bitstream.
[0344] The first set of clauses describes some embodiments of the disclosed technologies in the previous chapters.
[0345] 1. A video processing method, comprising: determining a single set of motion information for a current video block based on applying at least one interpolation filter to a set of adjacent blocks, wherein the at least one interpolation filter can be configured to integer pixel accuracy or subpixel accuracy of the single set of motion information; and performing a conversion between the current video block and a bitstream of the current video block, wherein the conversion includes a decoder motion vector refinement (DMVR) step for refining the single set of motion information signaled in the bitstream.
[0346] 2. The method according to Clause 1, wherein, when the single set of motion information is associated with an adaptive motion vector resolution (AMVR) mode, the index of the at least one interpolation filter is signaled in the bitstream.
[0347] 3. The method according to Clause 1, wherein, when the single set of motion information is associated with a merge mode, the index of the at least one interpolation filter is inherited from the previous video block.
[0348] 4. The method according to Clause 1, wherein, when the single set of motion information is associated with a merge (MMVD) mode employing motion vector difference, the index of the at least one interpolation filter is inherited from the previous video block.
[0349] 5. The method according to Clause 1, wherein the at least one interpolation filter corresponds to a default filter with sub-pixel accuracy, and wherein the single set of motion information is associated with a merge (MMVD) mode employing motion vector difference.
[0350] 6. The method according to any of the provisions of 2-4, wherein the coefficients of the at least one interpolation filter used for the current video block are inherited from the previous video block.
[0351] 7. The method according to any of the provisions of clauses 1-6, wherein the subpixel precision of the single set of motion information is equal to 1 / 4 pixel or 1 / 16 pixel.
[0352] 8. The method according to any of the provisions of clauses 1-7, wherein one or more components of the single set of motion information have subpixel accuracy.
[0353] 9. The method according to Clause 1, wherein the at least one interpolation filter is expressed using 6 or 8 coefficients.
[0354] 10. The method according to Clause 5, wherein the index of the at least one interpolation filter is associated only with the current video block and is not associated with subsequent video blocks.
[0355] 11. The method according to Clause 1, wherein information relating to the at least one interpolation filter is stored together with the single set of motion information.
[0356] 12. The method according to Clause 11, wherein information relating to the at least one interpolation filter identifies the at least one interpolation filter as a default filter.
[0357] 13. The method according to Clause 1, wherein the coefficients of the at least one interpolation filter used for the current video block are prevented from being used by the interpolation filter of another video block.
[0358] 14. The method according to any one of the provisions of 1-13, wherein the at least one interpolation filter corresponds to a plurality of filters, each of the plurality of filters being associated with the sub-pixel accuracy of the single set of motion information.
[0359] 15. The method according to any of the provisions of 1-14, wherein the precision of each component of the single set of motion information is equal to or less than the subpixel precision of the single set of motion information.
[0360] 16. The method according to any of the provisions of 1-15, wherein the at least one interpolation filter is stored in a history-based motion vector prediction (HMVP) lookup table.
[0361] 17. The method described under Clause 16 further includes:
[0362] When the motion information of another video block is detected to be the same as the single set of motion information of the current video block, the single set of motion information is inserted into the HMVP table, but the motion information of the other video block is not inserted.
[0363] 18. The method described under Clause 16 further includes:
[0364] When the motion information of another video block is detected to be the same as the single set of motion information of the current video block, the single set of motion information of the current video block and the motion information of the other video block are inserted into the HMVP table.
[0365] 19. The method according to any of the provisions of 17-18, wherein the other video block is associated with an interpolation filter, wherein the interpolation is based at least in part on the at least one interpolation filter of the current video block and / or the interpolation filter of the other video block.
[0366] 20. The method according to any of the provisions of 17-19, wherein the current video block and the other video block correspond to a pair of candidates or a combination of merge candidates.
[0367] 21. The method according to any of the provisions of 17-20, wherein the at least one interpolation filter of the current video block is the same as the interpolation filter of the other video block.
[0368] 22. The method according to any of the provisions of 17-20, wherein the at least one interpolation filter of the current video block is different from the interpolation filter of the other video block.
[0369] 23. The method according to Clause 1, wherein the current video block is encoded and decoded in an intra-block copy (IBC) mode, and wherein the use of sub-pixel precision in the representation of the single set of motion information is prohibited.
[0370] 24. The method described under any of the provisions of Clauses 1-23, wherein the use of the at least one interpolation filter is prohibited.
[0371] 25. The method according to Clause 1, wherein the at least one interpolation filter is associated with amvr_precision_idxflag and / or a hpel_if_idx flag.
[0372] 26. The method according to Clause 25, wherein the amvr_precision_idx flag and / or the hpel_if_idx flag are associated with bypass codec bits or context codec bits.
[0373] 27. The method according to Clause 26, wherein the first binary bit is a bypass codec binary bit or a context codec binary bit.
[0374] 28. The method described in any of the provisions of 25-27, wherein all binary bits share the same context.
[0375] 29. The method according to Clause 1, wherein one or more video processing steps are disabled based on the use of the at least one interpolation filter.
[0376] 30. The method according to Clause 29, wherein the one or more video processing steps include a decoder motion vector refinement (DMVR) step, a bidirectional optical flow (BDOF) step, a combined inter-frame intra-frame prediction (CIIP) step, a symmetric motion vector difference (SMVD) step, a sub-block transform (SBT) step, or a triangle prediction step.
[0377] 31. The method according to clause 30, wherein the at least one interpolation filter corresponds to a default filter.
[0378] 32. The method according to any of the provisions of 30-31, wherein prohibiting the one or more video processing steps includes an instruction to prohibit the one or more video processing steps in the bitstream.
[0379] 33. The method according to any of the provisions of 30-31, wherein prohibiting the one or more video processing steps includes prohibiting the inheritance of the at least one interpolation filter of the current video block to another video block.
[0380] 34. The method according to any one of the provisions of clauses 1-6, wherein the integer pixel precision of the single set of motion information corresponds to 1 pixel, 2 pixels or 4 pixels.
[0381] 35. The method according to clause 34, wherein the at least one interpolation filter is a low-pass filter.
[0382] 36. The method according to Clause 34, wherein the at least one interpolation filter is a one-dimensional filter.
[0383] 37. The method according to Clause 36, wherein the one-dimensional filter is a horizontal filter or a vertical filter.
[0384] 38. The method according to clause 34, wherein a flag in the bitstream indicates whether the at least one interpolation filter is used.
[0385] 39. The method described in any of the provisions of 1-38, wherein the current video block is associated with an Adaptive Motion Vector Resolution (AMVR) mode.
[0386] 40. The method according to any of the provisions of 1-39, wherein the at least one interpolation filter for the inter-frame mode of the current video block is different from the generalized bidirectional prediction (GBI) mode of the current video block.
[0387] 41. The method according to any one of the provisions of clauses 1-40, wherein the at least one interpolation filter is predetermined.
[0388] 42. The method described under any of the provisions of Clauses 1-41, wherein the bitstream includes a video parameter set (VPS), a picture parameter set (PPS), a picture header, a slice header, or a strip header associated with the current video block.
[0389] 43. The method according to any one of the provisions 1 to 42, wherein the video processing is a codec-side implementation.
[0390] 44. The method according to any one of the provisions 1 to 70, wherein the video processing is a decoder-side implementation.
[0391] 45. An apparatus in a video system, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of the provisions 1 to 44.
[0392] 46. A computer program product stored on a non-transitory computer-readable medium, the computer program product comprising program code for implementing the methods described in any one or more of clauses 1 to 45.
[0393] The second set of clauses describes some embodiments of the disclosed technologies in the previous sections (e.g., exemplary embodiments 3-6 and 16-23).
[0394] 1. A method for video processing (e.g., Figure 14A The method 1410 shown includes: a conversion between a current video block of a current image of a video and a codec representation of the video; determining (1412) the suitability of an alternative interpolation filter, wherein the suitability of the alternative interpolation filter indicates whether the alternative interpolation filter is applied in the conversion; and performing (1414) the conversion based on the determination; wherein the suitability of the alternative interpolation filter is determined based on whether a reference image resampling is used to perform the conversion by resampling a reference image of the current image.
[0395] 2. The method according to Clause 1, wherein an alternative interpolation filter is not applied due to the use of reference image resampling.
[0396] 3. The method according to Clause 2, wherein the reference image has a resolution different from that of the current image.
[0397] 4. The method according to Clause 1, wherein a default interpolation filter is applied if the determination determines that no alternative interpolation filter should be applied.
[0398] 5. The method according to Clause 4, wherein the default interpolation filter corresponds to an 8-tap filter with filter coefficients [-1 4 -11 40 40 -11 4 -1].
[0399] 6. The method according to any of the provisions 1 to 5, wherein the alternative interpolation filter corresponds to a 6-tap filter having filter coefficients [3 9 20 20 9 3].
[0400] 7. The method according to any one of the provisions 1 to 6, wherein the alternative interpolation filter is a half-pixel interpolation filter used in the prediction of the current video block.
[0401] 8. The method according to any of the provisions 1 to 7, wherein, if an alternative interpolation filter is used, at least one motion vector pointing to a half-pixel location is used to encode and decode the current video block.
[0402] 9. The method according to any one of the provisions 1 to 8, wherein, if an alternative interpolation filter is used, the current video block is encoded or decoded using at least one motion vector pointing to a horizontal half-pixel position or a vertical half-pixel position or both horizontal and vertical half-pixel positions.
[0403] 10. A method for video processing (e.g., Figure 14B The method 1420 shown includes: determining (1422) the codec mode for representing a current video block as a merge mode employing motion vector difference (MMVD), the MMVD including a motion vector representation providing information about the distance between a motion candidate and a starting point; and performing (1424) a conversion between the current video block and the codec representation based on the determination, wherein the conversion is performed using a prediction block of the current video block, the prediction block being computed using a half-pixel interpolation filter selected according to a first rule and a second rule, the first rule defining a first condition under which the half-pixel interpolation filter is an alternative half-pixel interpolation filter different from the default half-pixel interpolation filter, and the second rule defining a second condition regarding whether the alternative half-pixel interpolation filter is inherited.
[0404] 11. The method according to Clause 10, wherein the second rule specifies that the alternative half-pixel interpolation filter is inherited if the distance has either X-pixel precision or a coarser precision than X-pixel precision.
[0405] 12. The method described in Clause 11, wherein X is 1 / 2, and a coarser precision is 1, 2, or 4.
[0406] 13. The method according to Clause 10, wherein the second rule specifies that the alternative half-pixel interpolation filter is not inherited if the distance has either X-pixel precision or a finer precision than X-pixel precision.
[0407] 14. The method according to Clause 13, wherein X is 1 / 4, and for a finer precision, 1 / 16.
[0408] 15. A method for video processing (e.g., Figure 14CThe method 1430 shown includes: determining (1432) the encoding / decoding mode used in the current video region of the video; making a determination (1434) based on the encoding / decoding mode regarding the accuracy of the motion vector or motion vector difference used to represent the current video region; and performing (1436) a conversion between the current video block and the encoding / decoding representation of the video.
[0409] 16. The method according to Clause 15, wherein making a determination includes determining whether to inherit the precision.
[0410] 17. The method according to Clause 15, wherein, without inheriting the precision, the method further includes storing the precision.
[0411] 18. The method according to Clause 15, wherein, due to the determination of the precision inherited, the conversion includes encoding and decoding the current video region into the codec representation without directly signaling the precision.
[0412] 19. The method according to Clause 15, wherein, since the determination of the precision is not inherited, the conversion includes encoding the current video region into the codec representation by signaling the precision.
[0413] 20. The method according to Clause 15, wherein the video region corresponds to a video block, a codec unit, or a prediction unit.
[0414] 21. The method according to Clause 15, wherein the video region corresponds to the encoding / decoding unit.
[0415] 22. The method according to Clause 15, wherein the video region corresponds to a prediction unit.
[0416] 23. The method according to Clause 15, wherein the encoding / decoding mode corresponds to an Advanced Motion Vector Prediction (AMVP) mode, an affine inter-frame mode, a merge mode, and / or an affine merge mode, wherein motion candidates are derived from spatially adjacent blocks and temporally adjacent blocks.
[0417] 24. The method according to Clause 15, wherein, for candidates derived from neighboring blocks, the precision used to represent the motion vectors or motion vector differences used by neighboring blocks is inherited.
[0418] 25. The method according to Clause 15, wherein, for a pair of merge candidates associated with two spatial merge candidates having the same precision, the same precision is assigned to the pair of merge candidates, otherwise fixed motion precision information is assigned to the pair of merge candidates.
[0419] 26. The method according to Clause 15, wherein a fixed precision is assigned to temporal merge candidates.
[0420] 27. The method according to Clause 15 further comprises: maintaining at least one history-based motion vector prediction (HMVP) table before encoding / decoding the current video region or generating the decoded representation, wherein the HMVP table includes one or more entries corresponding to motion information of one or more previously processed blocks, and wherein the precision is stored in the HMVP table and inherited by the HMVP merge candidate.
[0421] 28. The method according to Clause 15, wherein, for the current video region encoded and decoded according to the merge mode (MMVD) with motion vector difference in which motion information is updated with motion vector difference, the accuracy of the basic merge candidate is inherited.
[0422] 29. The method according to Clause 15 or 16, wherein the accuracy of subsequent video regions is predicted using the accuracy that has been determined to be inherited.
[0423] 30. The method according to Clause 15, wherein the motion vector of the video block encoded and decoded in merge mode is rounded to the precision of the current video region.
[0424] 31. A method for video processing (e.g., Figure 14D The method 1440 shown includes: determining (1442) the encoding / decoding mode of the current video block of the video as a Merge mode using motion vector difference (MMVD); determining (1444) a distance table relating a specified distance index to a predefined offset for the current video block based on motion precision information of the basic merge candidates associated with the current video block; and using the distance table to perform (1446) a conversion between the current video block and the encoding / decoding representation of the video.
[0425] 32. The method according to clause 31, wherein, when the basic merge candidate has 1 / N pixel precision, a distance with 1 / N pixel precision or a coarser precision than 1 / N pixels is used, where N is a positive integer.
[0426] 33. The method according to Clause 31, wherein N is 2, and the distance is 1 / 2 pixel, 1 pixel, or 2 pixels.
[0427] 34. The method according to Clause 31, wherein the distance including the fractional distance is modified based on the motion accuracy information of the basic merge candidate.
[0428] 35. The method according to Clause 34, wherein a signaling notification flag is used to indicate whether fractional distance is permitted for MMVD mode.
[0429] 36. The method according to Clause 34, wherein, if the basic merge candidate has 1 / N pixel accuracy and the finest motion vector difference accuracy of the distance table has 1 / M pixel accuracy, the distance in the distance table is multiplied by M / N, where N and M are positive integers.
[0430] 37. A method for video processing (e.g., Figure 14E The method 1450 shown includes: making a first determination (1452): the motion vector difference used in the decoder-side motion vector refinement (DMVR) calculation for the first video region has a finer resolution than X pixel resolution, and the motion vector difference is determined using an alternative half-pixel interpolation filter, where X is an integer fraction; making a second determination (1454) based on the first determination: not storing or making the information of the alternative half-pixel interpolation filter associated with the first video region available for use in a second video region to be processed subsequently; and performing (1456) a conversion between the video including the first video region and the second video region and the codec representation of the video based on the first determination and the second determination.
[0431] 38. The method according to Clause 37, wherein the video region corresponds to a video block, a codec unit, or a prediction unit.
[0432] 39. The method according to Clause 37, wherein the second determination is made for each sub-block of the video when DMVR is performed at the sub-block level.
[0433] 40. The method according to Clause 37, wherein, for at least N sub-blocks or blocks of the video, the motion vector difference has a finer resolution than X pixel resolution.
[0434] 41. A method for video processing (e.g., Figure 14F The method 1460 shown includes:
[0435] According to rule (1462), it is determined whether the candidate half-pixel interpolation filter is available for the current video block of the video; based on this determination, a conversion between the current video block and the codec representation of the video is performed (1464); wherein the rule specifies that the determination is independent of whether the candidate half-pixel interpolation filter was used to process a previous block that was encoded or decoded before the current video block, in the case of encoding or decoding the current video block as a merge block or skipping a codec block in the codec representation.
[0436] 42. The method according to Clause 41, wherein parameters indicating whether to use an alternative half-pixel interpolation filter for the transformation are not stored for any block of the video.
[0437] 43. The method described in Clause 41, wherein whether to inherit or discard the alternative half-pixel interpolation filter depends on the codec information and / or the codec tools enabled or disabled in the current video block.
[0438] 44. A method for video processing (e.g., Figure 14F The method 1460 shown includes: determining (1462) whether an alternative half-pixel interpolation filter is available for the current video block of the video according to a rule; and performing (1464) a conversion between the current video block and the codec representation of the video based on the determination, wherein the rule specifies the applicability of an alternative half-pixel interpolation filter different from the default interpolation filter based on the codec information of the current video block.
[0439] 45. The method according to Clause 44, wherein the current video block is encoded or decoded in the codec representation as a merge block or a skipped codec block.
[0440] 46. The method according to Clause 44, wherein, in the case that the codec representation includes a merge index equal to K, an alternative half-pixel interpolation filter is applied in the current video block encoded in merge or skip mode, wherein K is a positive integer.
[0441] 47. The method according to Clause 44, wherein, in the case that the codec representation includes a merge index Idx satisfying Idx%S equal to K, an alternative half-pixel interpolation filter is applied in the current video block encoded in merge or skip mode, wherein S and K are positive integers.
[0442] 48. The method according to clause 44, wherein an alternative half-pixel interpolation filter is applied to a specific merge candidate.
[0443] 49. The method according to clause 48, wherein the particular merge candidate corresponds to a pair of merge candidates.
[0444] 50. A method for video processing (e.g., Figure 14G The method 1470 shown includes: determining (1472) coefficients of candidate half-pixel interpolation filters for a current video block of the video according to a rule; and performing (1474) a conversion between the current video block and the codec representation of the video based on the determination, wherein the rule specifies the relationship between the candidate half-pixel interpolation filters and the half-pixel interpolation filters used in a certain codec mode.
[0445] 51. The method according to Clause 50, wherein the rule specifies that the alternative half-pixel interpolation filter is consistent with the half-pixel interpolation filter used in affine inter-frame prediction.
[0446] 52. The method according to Clause 51, wherein the alternative interpolation filter has a set of coefficients of [3, -11, 40, 40, -11, 3].
[0447] 53. The method according to Clause 51, wherein the alternative interpolation filter has a set of coefficients of [3,9,20,20,9,3].
[0448] 54. The method according to Clause 50, wherein the rule specifies that the alternative half-pixel interpolation filter is the same as the half-pixel interpolation filter used in chroma inter-frame prediction.
[0449] 55. The method according to Clause 54, wherein the alternative interpolation filter has a set of coefficients of [-4, 36, 36, -4].
[0450] 56. A method for video processing (e.g., Figure 14H The method 1480 shown includes: applying (1482) an alternative half-pixel interpolation filter according to a rule for the conversion between the current video block and the codec representation of the video; and performing (1484) the conversion between the current video block and the codec representation of the video, wherein the rule specifies that the alternative interpolation filter is applied to the position at pixel X, wherein X is different from 1 / 2.
[0451] 57. The method according to Clause 56, wherein the use of alternative interpolation filters is explicitly or implicitly indicated.
[0452] 58. The method according to Clause 57, wherein the alternative interpolation filter has at least one coefficient that is different from the coefficient of the default interpolation filter already assigned to the position at pixel X.
[0453] 59. The method according to Clause 56, wherein the application of an alternative half-pixel interpolation filter is performed after decoder-side motion vector refinement (DMVR) or merge mode of motion vector difference (MMVD) or symmetric motion vector difference (SMVD) or other decoder-side motion derivation process.
[0454] 60. The method according to Clause 56, wherein an alternative half-pixel interpolation filter is applied when using reference picture resampling (RPR) and the reference picture has a resolution different from that of the current picture including the current video block.
[0455] 61. The method according to any of the provisions of clauses 56 to 60, wherein, when bidirectional prediction is used for the current video block, an alternative half-pixel interpolation filter is applied to only one prediction direction.
[0456] 62. A method for video processing (e.g., Figure 14I The method 1490 shown includes: making a first determination (1492): selecting an alternative interpolation filter for a first motion vector component having an accuracy of X pixels from a first video block of the video; making a second determination (1494) due to the first determination: applying another alternative interpolation filter for a second motion vector component having an accuracy different from X pixels, where X is an integer fraction; and performing (1496) a conversion between the video including the first video block and the second video block and the codec representation of the video.
[0457] 63. The method according to Clause 62, wherein the alternative interpolation filter and the other alternative interpolation filter correspond to Gaussian interpolation filters.
[0458] 64. The method according to Clause 62, wherein the alternative interpolation filter and the other alternative interpolation filter correspond to the Flattop interpolation filter.
[0459] 65. The method according to Clause 62, wherein the second motion vector component has a finer precision than the X pixel precision.
[0460] 66. The method according to Clause 62, wherein the second motion vector component has a coarser precision than the X pixel precision.
[0461] 67. The method described under any of the provisions 1 to 66, wherein the performance of the transformation includes generating a codec representation from the current video block.
[0462] 68. The method described under any of the provisions 1 to 66, wherein the performance of the conversion includes generating the current video block from the codec representation.
[0463] 69. An apparatus in a video system, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of claims 1 to 68.
[0464] 70. A computer program product stored on a non-transitory computer-readable medium, the computer program product comprising program code for implementing the methods described in any one of clauses 1 to 68.
[0465] The disclosed and other solutions, examples, embodiments, modules, and functional operations described in this document may be implemented in digital electronic circuits or in computer software, firmware, or hardware (including the structures disclosed in this document and their structural equivalents), or in a combination thereof. The disclosed embodiments and other embodiments may be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a computer-readable medium for execution by a data processing apparatus or for controlling the operation of the data processing apparatus. The computer-readable medium may be a machine-readable storage device, a machine-readable storage substrate, a storage device, a material composition that influences machine-readable propagated signals, or a combination thereof. The term "data processing apparatus" encompasses all means, devices, and machines for processing data, including, for example, a programmable processor, a computer, or multiple processors or computers. In addition to hardware, the apparatus may also include code that creates an execution environment for the computer program under consideration, such as code constituting processor firmware, a protocol stack, a database management system, an operating system, or a combination thereof. The propagated signals are artificially generated signals, such as machine-generated electrical, optical, or electromagnetic signals, which are generated to encode information for transmission to a suitable receiver device.
[0466] Computer programs (also known as programs, software, software applications, scripts, or code) can be written in any programming language (including compiled or interpreted languages) and can be deployed in any form, including as standalone programs or as modules, components, subroutines, or other units suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinating files (e.g., files storing one or more modules, subroutines, or code portions). Computer programs can be deployed to execute on one or more computers located at a single site or distributed across multiple sites interconnected via a communication network.
[0467] The processes and logic flows described in this document can be executed by one or more programmable processors executing one or more computer programs, thereby performing functions by manipulating input data and generating outputs. These processes and logic flows can also be executed by dedicated logic circuits, and the devices can be implemented as dedicated logic circuits, such as FPGAs (Field-Programmable Gate Arrays) or ASICs (Application-Specific Integrated Circuits).
[0468] For example, processors suitable for executing computer programs include general-purpose and special-purpose microprocessors, as well as any one or more processors in any kind of digital computer. Generally, a processor receives instructions and data from read-only memory or random access memory, or both. The basic components of a computer are a processor that executes instructions and one or more storage devices that store the instructions and data. Typically, a computer will also include one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks, or operatively coupled to receive data from or transfer data to one or more mass storage devices, or both. However, a computer does not necessarily have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices, such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and CD-ROMs and DVD-ROMs. The processor and memory may be supplemented by or incorporated into special-purpose logic circuitry.
[0469] While this patent document contains numerous details, it should not be construed as limiting any subject matter or scope of the claims, but rather as a description of specific features of particular embodiments of a particular technology. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments. Furthermore, while certain features may be described above as functioning in certain combinations and even initially claimed in this manner, one or more features from the claimed combination may be removed from that combination in certain circumstances, and the claimed combination may involve sub-combinations or variations thereof.
[0470] Similarly, although the operations are described in a specific order in the accompanying drawings, this should not be construed as requiring the specific order or sequential execution of such operations to obtain the desired result, or requiring the execution of all illustrated operations. Furthermore, the division of various system components in the embodiments described in this patent document should not be construed as requiring such division in all embodiments.
[0471] Only a few implementation methods and examples have been described. Other implementation methods, enhancements and variations can be made based on the content described and illustrated in this patent document.
Claims
1. A video processing method, comprising: For the conversion between the current video block of the current image of the video and the bitstream of the video, the first half-pixel interpolation filter is used based on whether the reference image resampling tool is enabled; as well as The conversion is performed based on the determination; If the reference image resampling tool is enabled, the resolution of the current image will be different from the resolution of the reference image of the current video block.
2. The method according to claim 1, wherein, Based on the activation of the reference image resampling tool, the first half-pixel interpolation filter is not applied to the current video block.
3. The method according to claim 1, wherein, Based on the activation of the reference image resampling tool, a second half-pixel interpolation filter, different from the first half-pixel interpolation filter, is applied to the current video block.
4. The method according to claim 1, wherein, The index of the first half-pixel interpolation filter is equal to 1.
5. The method according to claim 3, wherein, The index of the second half-pixel interpolation filter is equal to 0.
6. The method according to claim 3, wherein, The second half-pixel interpolation filter corresponds to an 8-tap filter with filter coefficients [-1 4-11 40 40-11 4-1].
7. The method according to claim 1, wherein, The first half-pixel interpolation filter corresponds to a 6-tap filter with filter coefficients [3 9 20 20 9 3].
8. The method according to claim 1, wherein, If the first half-pixel interpolation filter is used, the current video block is encoded and decoded using at least one motion vector pointing to the half-pixel position.
9. The method according to claim 1, wherein, If the first half-pixel interpolation filter is used, the current video block is encoded and decoded using at least one motion vector pointing to a horizontal half-pixel position, a vertical half-pixel position, or both horizontal and vertical half-pixel positions.
10. The method according to claim 1, wherein, The conversion includes encoding the current video block into the bitstream.
11. The method according to claim 1, wherein, The conversion includes decoding the current video block from the bitstream.
12. The method according to claim 1, further comprising, before performing the conversion based on the determination: The suitability of alternative interpolation filters is determined, wherein the suitability of the alternative interpolation filters indicates whether the alternative interpolation filters are applied in the transformation, and the suitability of the alternative interpolation filters is determined based on whether reference image resampling is used to resample the reference image of the current image to perform the transformation.
13. The method according to claim 12, wherein, The alternative interpolation filter is not applied due to the use of the reference image resampling.
14. The method according to claim 13, wherein, The resolution of the reference image is different from the resolution of the current image.
15. The method according to claim 12, wherein, If it is determined that the alternative interpolation filters should not be applied, the default interpolation filter shall be applied.
16. The method according to claim 15, wherein, The default interpolation filter corresponds to an 8-tap filter with filter coefficients [-1 4-11 40 40-11 4-1].
17. The method according to any one of claims 12 to 16, wherein, The alternative interpolation filter corresponds to a 6-tap filter with filter coefficients [3 9 20 20 9 3].
18. The method according to any one of claims 12 to 16, wherein, The alternative interpolation filter is a half-pixel interpolation filter used in the prediction of the current video block.
19. The method according to any one of claims 12 to 16, wherein, If the alternative interpolation filter is used, the current video block is encoded and decoded using at least one motion vector pointing to a half-pixel position.
20. The method according to any one of claims 12 to 16, wherein, If the alternative interpolation filter is used, the current video block is encoded and decoded using at least one motion vector pointing to a horizontal half-pixel position, a vertical half-pixel position, or both horizontal and vertical half-pixel positions.
21. The method of claim 1, further comprising, before performing the conversion based on the determination: The encoding / decoding mode of the current video block used to represent the video is determined to be a merge mode using motion vector difference (MMVD), the MMVD including a motion vector representation that provides information about the distance between the motion candidate and the starting point; The transformation is performed using a prediction block of the current video block, which is calculated using a half-pixel interpolation filter selected according to a first rule and a second rule. The first rule defines a first condition under which the half-pixel interpolation filter is an alternative half-pixel interpolation filter that is different from the default half-pixel interpolation filter. The second rule defines a second condition regarding whether to inherit the alternative half-pixel interpolation filter.
22. The method according to claim 21, wherein, The second rule specifies that the alternative half-pixel interpolation filter is inherited if the distance has either X-pixel precision or a coarser precision than X-pixel precision.
23. The method according to claim 22, wherein, X is 1 / 2, and the coarser precision is 1, 2, or 4.
24. The method according to claim 21, wherein, The second rule specifies that the alternative half-pixel interpolation filter is not inherited if the distance has either X-pixel accuracy or a finer accuracy than X-pixel accuracy.
25. The method according to claim 24, wherein, X is 1 / 4, and the finer precision is 1 / 16.
26. The method of claim 1, further comprising, before performing the conversion based on the determination: Determine the encoding / decoding mode used in the current video region of the video; Based on the encoding / decoding mode, a determination is made regarding the accuracy of the motion vector or motion vector difference used to represent the current video region.
27. The method according to claim 26, wherein, The determination mentioned above includes determining whether to inherit the precision.
28. The method according to claim 26, wherein, Without inheriting the precision, the method further includes storing the precision.
29. The method according to claim 26, wherein, Since the precision is inherited, the conversion includes encoding and decoding the current video region into the bitstream without directly signaling the precision.
30. The method according to claim 26, wherein, Since the determination of the precision is not inherited, the conversion includes encoding and decoding the current video region into the bitstream by signaling the precision.
31. The method according to claim 26, wherein, The video region corresponds to a video block, encoding / decoding unit, or prediction unit.
32. The method according to claim 26, wherein, The video region corresponds to the encoding / decoding unit.
33. The method according to claim 26, wherein, The video region corresponds to the prediction unit.
34. The method according to claim 26, wherein, The encoding / decoding modes correspond to Advanced Motion Vector Prediction (AMVP) mode, Affine Inter-Frame mode, merge mode, and / or Affine Merge mode, wherein motion candidates are derived from spatially adjacent blocks and temporally adjacent blocks.
35. The method according to claim 26, wherein, For candidates derived from neighboring blocks, the precision used to represent the motion vector or motion vector difference used by the neighboring blocks is inherited.
36. The method according to claim 26, wherein, For a pair of merge candidates associated with two spatial merge candidates having the same precision, the same precision is assigned to the pair of merge candidates; otherwise, fixed motion precision information is assigned to the pair of merge candidates.
37. The method according to claim 26, wherein, For time-domain merge candidates, a fixed precision is assigned.
38. The method of claim 26, further comprising: Before encoding / decoding the current video region or generating the bitstream, at least one history-based motion vector prediction (HMVP) table is maintained, wherein the HMVP table includes one or more entries corresponding to motion information of one or more previously processed blocks, and The precision is stored in the HMVP table and inherited by the HMVP merge candidate.
39. The method according to claim 26, wherein, For the current video region encoded and decoded using the merge mode (MMVD) that updates motion information with motion vector differences, the accuracy of the basic merge candidate is inherited.
40. The method according to claim 26 or 27, wherein, Use the accuracy that has already been determined to be inherited to predict the accuracy of subsequent video regions.
41. The method according to claim 26, wherein, The motion vectors of the video blocks encoded and decoded according to the merge mode are rounded to the precision of the current video region.
42. The method according to claim 1, further comprising, before performing the conversion based on the determination: The encoding / decoding mode of the current video block of the video is determined to be the merge mode of motion vector difference (MMVD); as well as For the current video block, a distance table is determined based on the motion precision information of the basic merge candidates associated with the current video block, which establishes the relationship between a specified distance index and a predefined offset. The step of performing the conversion based on the determination also includes: The distance table is used to perform the conversion between the current video block and the bitstream of the video.
43. The method according to claim 42, wherein, When the basic merge candidate has 1 / N pixel precision, a distance with 1 / N pixel precision or a coarser precision than 1 / N pixels is used, where N is a positive integer.
44. The method according to claim 42, wherein, N is 2, and the distance is 1 / 2 pixel, 1 pixel, or 2 pixels.
45. The method according to claim 42, wherein, Based on the motion accuracy information of the basic merge candidates, the distance table including fractional distances is modified.
46. The method according to claim 45, wherein, A signaling notification flag is used to indicate whether the fractional distance is allowed for MMVD mode.
47. The method according to claim 45, wherein, If the basic merge candidate has 1 / N pixel accuracy and the finest motion vector difference accuracy of the distance table has 1 / M pixel accuracy, the distance in the distance table is multiplied by M / N, where N and M are positive integers.
48. The method of claim 1, further comprising, before performing the conversion based on the determination: A first determination is made: the motion vector difference used in the decoder-side motion vector refinement (DMVR) calculation for the first video region has a finer resolution than X pixel resolution, and the motion vector difference is determined using an alternative half-pixel interpolation filter, where X is an integer fraction; and Based on the first determination, a second determination is made: information on the candidate half-pixel interpolation filters associated with the first video region will not be stored, or the information will be made unusable by the second video region to be processed subsequently. The step of performing the conversion based on the determination also includes: Based on the first determination and the second determination, a conversion is performed between the video including the first video region and the second video region and the bitstream of the video.
49. The method according to claim 48, wherein, Video regions correspond to video blocks, encoding / decoding units, or prediction units.
50. The method according to claim 48, wherein, The second determination is made for each sub-block of the video when DMVR is performed at the sub-block level.
51. The method according to claim 48, wherein, For at least N sub-blocks or blocks of the video, the motion vector difference has a finer resolution than X pixel resolution.
52. The method according to claim 1, further comprising, before performing the conversion based on the determination: Determine whether the candidate half-pixel interpolation filter is available for the current video block of the video according to the rules; The rule specifies that, when the current video block is encoded or decoded in the bitstream as a merge block or a skipped encoding / decoding block, the determination is independent of whether the alternative half-pixel interpolation filter is used to process a previous block that was encoded or decoded before the current video block.
53. The method according to claim 52, wherein, The parameters for whether to use the alternative half-pixel interpolation filter for the transformation are not stored for any block of the video.
54. The method according to claim 52, wherein, Whether to inherit or discard the alternative half-pixel interpolation filter depends on the codec information and / or the codec tools enabled or disabled in the current video block.
55. The method of claim 1, further comprising, before performing the conversion based on the determination: According to the rules, determine whether the candidate half-pixel interpolation filter is available for the current video block of the video; The rule specifies the applicability of the alternative half-pixel interpolation filter, which differs from the default interpolation filter, based on the encoding and decoding information of the current video block.
56. The method according to claim 55, wherein, The current video block is encoded or decoded in the bitstream as a merge block or a skipped encoding / decoding block.
57. The method of claim 55, wherein, In the case that the bitstream includes a merge index equal to K, the alternative half-pixel interpolation filter is applied in the current video block encoded and decoded according to merge or skip mode, where K is a positive integer.
58. The method according to claim 55, wherein, In the case that the bitstream includes a merge index Idx satisfying Idx%S equal to K, the alternative half-pixel interpolation filter is applied in the current video block encoded and decoded according to merge or skip mode, where S and K are positive integers.
59. The method according to claim 55, wherein, The alternative half-pixel interpolation filter is applied to a specific merge candidate.
60. The method according to claim 59, wherein, The specific merge candidate corresponds to the pair merge candidate.
61. The method according to claim 1, further comprising, before performing the conversion based on the determination: The coefficients of the candidate half-pixel interpolation filters are determined for the current video block of the video according to the rules; The rule specifies the relationship between the candidate half-pixel interpolation filters and the half-pixel interpolation filters used in a certain encoding / decoding mode.
62. The method according to claim 61, wherein, The rule specifies that the candidate half-pixel interpolation filter is consistent with the half-pixel interpolation filter used in affine inter-frame prediction.
63. The method according to claim 62, wherein, The alternative half-pixel interpolation filter has a coefficient set of [3, -11, 40, 40, -11, 3].
64. The method according to claim 62, wherein, The alternative half-pixel interpolation filter has a coefficient set of [3,9,20,20,9,3].
65. The method according to claim 61, wherein, The rule specifies that the alternative half-pixel interpolation filter is the same as the half-pixel interpolation filter used in chroma inter-frame prediction.
66. The method according to claim 65, wherein, The alternative half-pixel interpolation filter has a coefficient set of [-4, 36, 36, -4].
67. The method of claim 1, further comprising, before performing the conversion based on the determination: For the conversion between the current video block and the bitstream of the video, an alternative half-pixel interpolation filter is applied according to the rules; The rule specifies that an alternative interpolation filter is applied to the position at pixel X, where X is different from 1 / 2.
68. The method according to claim 67, wherein, The use of the alternative interpolation filters is indicated explicitly or implicitly.
69. The method according to claim 68, wherein, The alternative interpolation filter has at least one coefficient that is different from the coefficient of the default interpolation filter already assigned to the position at pixel X.
70. The method of claim 67, wherein, The application of the alternative half-pixel interpolation filter is performed after decoder-side motion vector refinement (DMVR), or merge mode with motion vector difference (MMVD), or symmetric motion vector difference (SMVD), or other decoder-side motion derivation processes.
71. The method according to claim 67, wherein, The alternative half-pixel interpolation filter is applied when using reference image resampling (RPR) and the reference image has a resolution different from that of the current image including the current video block.
72. The method according to any one of claims 67 to 71, wherein, When bidirectional prediction is used for the current video block, the alternative half-pixel interpolation filter is applied to only one prediction direction.
73. The method of claim 1, further comprising, before performing the conversion based on the determination: Make a first determination: select an alternative interpolation filter for the first motion vector component with an accuracy of X pixels from the first video block of the video; A second determination is made based on the first determination: another alternative interpolation filter is used by the second video block of the video for a second motion vector component with a precision different from the X pixel precision, where X is an integer fraction; The step of performing the conversion based on the determination also includes: Perform a conversion between the video, which includes the first video block and the second video block, and the bitstream of the video.
74. The method according to claim 73, wherein, The alternative interpolation filter and the other alternative interpolation filter correspond to Gaussian interpolation filters.
75. The method according to claim 73, wherein, The alternative interpolation filter and the other alternative interpolation filter correspond to the Flattop interpolation filter.
76. The method according to claim 73, wherein, The second motion vector component has a finer precision than the X pixel precision.
77. The method according to claim 73, wherein, The second motion vector component has a coarser precision than the X pixel precision.
78. The method according to any one of claims 12 to 16, 21 to 39, 41 to 71, and 73 to 77, wherein, The conversion process includes generating the bitstream from the current video block.
79. The method according to any one of claims 12 to 16, 21 to 39, 41 to 71, and 73 to 77, wherein, The conversion process includes generating the current video block from the bitstream.
80. An apparatus for processing video data, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to: For the conversion between the current video block of the current image and the bitstream of the video, the use of a first half-pixel interpolation filter is determined based on whether a reference image resampling tool is enabled; and The conversion is performed based on the determination; in, If the reference image resampling tool is enabled, the resolution of the current image will be different from the resolution of the reference image of the current video block.
81. The apparatus according to claim 80, wherein, Based on the activation of the reference image resampling tool, the first half-pixel interpolation filter is not applied to the current video block.
82. The apparatus according to claim 80, wherein, Based on the activation of the reference image resampling tool, a second half-pixel interpolation filter, different from the first half-pixel interpolation filter, is applied to the current video block.
83. The apparatus according to claim 82, wherein, The index of the first half-pixel interpolation filter is equal to 1, and the index of the second half-pixel interpolation filter is equal to 0.
84. The apparatus according to claim 82, wherein, The second half-pixel interpolation filter corresponds to an 8-tap filter with filter coefficients [-1 4-11 40 40-11 4-1].
85. The apparatus according to claim 80, wherein, The first half-pixel interpolation filter corresponds to a 6-tap filter with filter coefficients [3 9 20 20 9 3].
86. A non-transitory computer-readable storage medium storing instructions, said instructions causing a processor to: For the conversion between the current video block of the current image and the bitstream of the video, the use of a first half-pixel interpolation filter is determined based on whether a reference image resampling tool is enabled; and The conversion is performed based on the determination; in, If the reference image resampling tool is enabled, the resolution of the current image will be different from the resolution of the reference image of the current video block.
87. The non-transitory computer-readable storage medium according to claim 86, wherein, Based on the activation of the reference image resampling tool, the first half-pixel interpolation filter is not applied to the current video block.
88. A method for storing a bitstream of video, comprising: For the current video block of the current image in the video, determine whether to use the first half-pixel interpolation filter based on whether the reference image resampling tool is enabled; The bit stream is generated based on the determination; as well as The bitstream is stored in a non-transitory computer-readable recording medium; If the reference image resampling tool is enabled, the resolution of the current image will be different from the resolution of the reference image of the current video block.
89. An apparatus in a video system, comprising a processor and a non-transitory memory having instructions located thereon, wherein, When executed by the processor, the instructions cause the processor to perform the method according to any one of claims 2 to 79.
90. A non-transitory computer-readable medium storing instructions that, when executed by a processor, cause the processor to perform program code according to any one of claims 2 to 79.
Citation Information
Patent Citations
Intra BC and inter unification
CN106797476A
Method and apparatus of adaptive inter prediction in video coding
CN108028931A